Design Arena Raises $7.9M to Teach AI Taste

The startup behind Design Arena just landed a $7.9 million seed round, and the raise says something bigger about where AI is headed. According to TechCrunch AI, the company, dubbed Intelligence, closed the round on Monday with Index Ventures leading and Conviction (Sarah Guo and Mike Vernal), A*, Valkyrie, and others joining in. The pitch is simple but sharp: AI models can generate images, websites, and games, but they can’t tell you what’s actually good. Humans still can.

Where This Started

Co-founder Grace Li tells TechCrunch AI the company began just weeks before her 2025 graduation, when she and a handful of college friends were trying to make an AI game engine work. The models could build functional games. None of them were fun. That raised the question that became the whole business: how do you measure whether something is fun? They decided there was no substitute for human judgment, then started figuring out how to collect honest human feedback at scale.

That pivot turned into Design Arena, now used by 5.3 million people worldwide. “It was the missing bottleneck for a lot of these models to make improvements in the design space,” Li says. “About a week later, we closed our first major deal with a frontier lab, and the rest is kind of history.”

How It Works

For everyday users, Design Arena feels like a smart model router. There’s a ChatGPT-style prompt window with dropdowns for websites, images, and a dozen other visual formats. You enter your request, style, and format, then get served a run of “A vs. B” choices until you’ve ranked the outputs from best to worst.

The real money sits on the enterprise side. Participating models treat the platform as a firehose of instant human feedback. Users mostly don’t know or care which model made what. They just pick the output they like. That indifference is the point, since it produces clean signal on what people actually want. Frontier labs will pay for that, and they are: the site is generating $60 million in ARR, per TechCrunch AI.

There’s a second layer of value. Because users log in to get their results, Intelligence can track how taste shifts across regions and over time. Li notes that web dashboards in Asia lean toward a more maximalist design style.

Why It Matters

What stands out here is the growing gap between automated benchmarks and human judgment. Benchmarks scale beautifully, but they can be gamed or manipulated. Last week’s Hugging Face breach showed how badly that can go. Human-ranked evaluation is slower and messier, but it’s much harder to fake, which makes it a valuable complement rather than a replacement.

That matters for anyone building or buying AI products. Model quality on design tasks is increasingly judged by real preference data, not just leaderboard scores. If you ship AI-generated visuals, expect the labs behind your models to lean harder on this kind of human signal.

The Catch

Human evaluation isn’t an automatic win. TechCrunch AI points to Yupp, which shut down earlier this year less than 12 months after launch, despite raising $33 million from a16z crypto’s Chris Dixon. Yupp had frontier labs as customers and over 1.3 million users, and still couldn’t build a durable business.

Others are thriving, though. LM Arena, which runs a similar model for text responses, raised $150 million in a Series A in January, just four months after launching its paid product. So the category is real, but survival depends on turning feedback into something labs can’t easily replicate in-house.

What To Watch

The money is flowing toward companies that can turn human taste into structured data. Intelligence has revenue, marquee customers, and a growing dataset that gets more valuable as it tracks preferences across continents and time. The open question is whether human evaluation stays a paid service or whether the big labs eventually build it themselves. For now, taste is a product, and the labs are buying. You can read the full report at TechCrunch AI.

Scroll to Top