Frontier AI Models Keep Recommending Themselves

Ask seven frontier AI models what the best coding agent is, and a funny thing happens: each one tends to name its own family’s product. That’s one of the standout findings from the new Latent Space Frontier AEO tracker, a large-scale study of what AI models actually recommend when users ask for the best tool in a category. Latent Space ran the experiment across 161 categories, from coding agents to AI podcasts to payroll software, and the results are a useful map for anyone trying to get their product named by an AI.

What they did

The methodology builds on AmplifyingAI’s approach. Latent Space fired 6 prompt variations at 7 models (all with web search turned on) across those 161 categories. Answer extraction was handled by Astra, then each product got a proprietary AEO score.

The scoring isn’t just a popularity count. It weights:

  • First choices most heavily
  • Alternative choices and passing mentions
  • Negative weights for mild and strong anti-recommendations (rare, but they happen)

Every prompt and answer pair is inspectable, which matters because a study like this invites questions about contamination. Latent Space says it checked for that and left the receipts open.

The self-bias is real

The most quotable finding is model favoritism. When asked to recommend coding agents, Fable and Opus lean toward Claude Code, Sol and Astra lean toward Codex, Grok favors Cursor, and so on down the line. Each lab’s model tilts toward its own house brand.

What stands out is the exception. Latent Space flags cases where GPT models recommend Claude anyway, calling it “a laudable nonbias.” So the tilt is a tendency, not an iron law.

Where the battlegrounds are

Only 28 of the 161 categories have a single primary choice that every surveyed model agrees on. That’s the headline number for marketers. The other 133 categories are up for grabs, full of “close contests” where recommendations split between models.

Those contested categories are the real AEO battlegrounds. If you’re competing in one, no single incumbent owns the AI’s answer, which means your content and positioning can still move the needle.

Model generations shift what gets picked

Latent Space also tracked what happens when a lab ships a new model. Comparing generations turned up “VERY consequential flips” in which products get recommended. One measurable change is how many sources each model reads before answering:

  • Sol: median of 9 sources
  • Astra: median of 5 sources
  • Opus: median of 11 sources
  • Fable: median of 15 sources

Astra reads fewer sources and behaves more “confidently,” the report notes. It’s far less likely to change its answer when you lightly reword the question. That consistency is significant for practitioners: as answer randomness drops, the payoff for ranking well rises. If a model won’t flip its choice on a paraphrase, earning that top spot is worth more.

What you can actually do

The practical takeaway is about being readable to machines. Latent Space validated that AEO practices measured by Ora and Vercel, like markdown content-negotiation, genuinely affect outcomes. When your content fails those checks, models are discouraged from reading it at all.

So the to-do list is straightforward: make sure AI crawlers can cleanly read your pages, target the contested categories where no product dominates, and remember that each model carries its own soft bias you’re working against.

The limits

Latent Space is upfront about the caveats. The sources analysis rests on a small sample, and it only reflects what can be scraped from attempted tool calls, not the underlying pretraining data. Some entries still need de-duplication. Treat the source-level conclusions as directional rather than definitive.

The bigger signal here is that AEO is becoming measurable. As new model generations reshuffle their recommendations, tracking those flips is turning into its own discipline. Full rankings, the flip reports, and the inspectable data are available at Latent Space.

Scroll to Top