About

Every few days someone drops a new leaderboard proving that Model A is 2.3% better than Model B at reasoning, coding, or being polite. None of those spreadsheets tell you what you actually want to know: if I ask this thing to build a landing page, will it look good?

WhichAI.dev asks a growing list of AI models to design five landing-page concepts from the same prompt and brief, preserves multiple attempts, and lays the results out so you can see the differences with your own eyes — no benchmark literacy required. It now serves 16,000+ monthly active users, and has been used on camera by Theo Browne (t3.gg) and other creators in videos with around a million combined views — see it in action in GPT-5.6: The Review.

What you’ll find

  • A gallery, not a spreadsheet — model runs grouped by condition (baseline, design-skill-enabled, taste experiments), each card showing five real rendered previews.
  • Compare anything — any two runs side by side in a single shareable URL. Baseline vs. design skill, Model X vs. Model Y, iteration 1 vs. iteration 5.
  • Rankings with actual notes — subjective, because design is subjective: what worked, what felt like a template, where a model showed real taste.
  • Lab Guess — look at a generation, guess the model, see if you’re right. Surprisingly educational.

Why it exists

Model selection is a design decision. Some models give you solid structure and weak taste; some produce one gorgeous screen and ignore the rest of the brief; some completely change character when you toggle a design skill. WhichAI.dev makes those differences obvious.