Vals Raises $40M to Fix AI’s Broken Report Card

Picture a student who gets a copy of the final exam a month early, memorizes every answer, and then walks out with a perfect score. Nobody would call that a measure of intelligence. Yet that’s roughly how a lot of AI benchmarking works today, and it’s the problem a San Francisco startup called Vals says it’s here to fix. According to TechCrunch AI, Vals closed a $40 million Series A last month led by Andreessen Horowitz, less than two years after the company was founded in 2024.

The seed round came from 8VC and Bloomberg Beta last year. Now a16z is betting that independent, private AI evaluation becomes infrastructure the whole industry has to pay for.

What’s broken with benchmarks

Benchmarks are how AI labs prove their models are good. When the numbers land in their favor, they become the headline of the launch post. Good benchmark, good PR.

The catch, as TechCrunch AI reports, is that many legacy benchmarks are old, public, and weren’t built for frontier models. If the test questions are on the internet, they end up in the training data. The model isn’t reasoning its way to the answer. It’s remembering it.

Rayan Krishnan, the 25-year-old co-founder of Vals, put it plainly: “We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance.”

How Vals does it differently

The company’s approach rests on two ideas:

  • Keep the test private. Vals doesn’t publish its test materials, so labs can’t train against them. It’s the SAT model: the exam stays sealed until you sit down to take it.
  • Test real work, not trivia. Instead of asking whether a model knows enough to pass a bar exam, Vals checks whether it can do the actual job of a lawyer, an analyst, or a software engineer at human quality. Krishnan’s framing: “Can they do work that produces a product of the same quality as a human within every domain?”

The coverage keeps expanding. Beyond law, finance, and coding, Krishnan told TechCrunch AI the company now runs benchmarks on recursive self improvement, mental health, cybersecurity, biosecurity, and even the law of armed conflict, testing whether models can apply the Geneva Convention correctly. Part of the goal is measuring downside risk, or as Krishnan says, “if these models ran wild in the world, what the negative implications would be.”

The business model sounds odd until it doesn’t

AI companies pay Vals to test their own models. On the surface, that’s strange. Why pay to find out your model underperforms?

Krishnan compares it to a student paying the College Board for the SAT. A trusted score is worth money. It helps labs find weak spots, and it’s increasingly what enterprise buyers ask for before signing a contract. The numbers back that up: revenue is eight times what it was last year, the team went from 8 to 25 people since January, and 10 to 15 more hires are planned. Vals also launched a program to evaluate models for federal agencies.

Why this matters

What stands out here is timing. AI companies are heading to public markets. Krishnan points to SpaceX already being public, Anthropic slated for later this year, and OpenAI likely to follow. Once model capability claims show up in S-1 filings and investor decks, “we scored 92% on a public benchmark” stops being enough. Someone independent needs to sign off.

That’s the bet a16z is making: evaluations become the audit layer for the AI economy, the way accounting firms are for financial statements.

For practitioners, a few things follow:

  1. Expect vendors to start quoting private, domain-specific scores instead of leaderboard rankings. Ask which evaluator produced them.
  2. If you’re building on top of models, industry-specific evals are a better signal for your use case than general knowledge tests.
  3. Regulators and government buyers are moving in this direction too. Vals’ federal program is an early sign.

The risk is obvious: whoever controls the private test gains real power over which models win. Vals will need to prove it can stay neutral while getting paid by the companies it grades. The College Board analogy cuts both ways.

Full interview and office tour details are at TechCrunch AI.

Scroll to Top