Grok 4.7 costs half as much as Opus, so is it half as good?

New numbers: Grok 4.7 lands at $2 per million input tokens and $6 per million output, while the top frontier models charge roughly five times that. The question is whether the cheaper model actually holds up. I just watched a breakdown from the creator behind the Forward Future channel, and he did something I appreciate: he showed every benchmark, including the ones xAI conveniently left off the blog post. Here’s what I took away from it.

The stat that frames everything

On Cursorbench 4.0, Grok 4.7 at extra high thinking scores 46.3%. That puts it just behind Claude Opus 5 at max thinking, but at about half the cost to run the whole benchmark. The expert points out one caveat right away: xAI now owns Cursor, so take that specific benchmark with a grain of salt.

Still, the shape of the result repeats across other tests. Grok 4.7 sits one step off the absolute frontier, and it gets there for a fraction of the price. Elon Musk tweeted a week earlier that 4.7 should be “roughly on par with Opus 5, not 5.1.” Based on what the creator showed, that prediction mostly holds.

Insight breakdown: where it wins, where it doesn’t

The post’s author walked through a lot of charts. Here’s the skimmable version:

  • Cost per task is the real metric. The creator stresses that tokens per task and price per token both matter. A model that burns twice the tokens at half the price is a wash. On Cursorbench, Grok 4.7 uses about the same tokens per task as Opus 5, so the lower price flows straight into cheaper completed work.
  • Deep Suite (engineer sentiment benchmark): Grok 4.7 hits 71%. GPT 5.6 Soul gets 72.7%, Fable 5.1 gets 70%. The blog omitted GPT6 Astra, so the creator had Astra rebuild the table. Astra Max came in at 74.1%, the top score. He calls the omission “slightly disingenuous” and I agree.
  • Terminal Bench 4.0: This one hurts. Grok 4.7 scores 38%. Fable 5.1 gets 57.9% and Astra gets 58.2%. The expert calls this one of the most important benchmarks for agentic coding, and Grok is well behind here.
  • Legal work: Grok 4.7 leads the big names at 19.6%, with Soul, Fable, and Astra all under 7%. The surprise winner is Muse Spark 1.2 at 42%, an open-weights model.
  • GDPval (OpenAI’s knowledge work benchmark): Fable 5.1 sits at 1735 ELO, Grok 4.7 at 1695, and Astra in fourth at 1542. Grok is genuinely near the frontier on office-style tasks.
  • Artificial Analysis Intelligence Index: Grok 4.7 lands in fifth place at 46, behind Fable 5.1, Astra, Opus 5, and Muse Spark 1.3 Max. The AA team also notes that Grok 4.7’s gains come with higher token usage, which nudges cost per task back up.

One more thing the creator flagged: the context window is 500K tokens, while most frontier models offer a million. That only matters for specialized long-context work, but it matters.

Why it’s so cheap

The expert’s explanation here is the part I found most interesting. Two reasons. First, Grok isn’t the frontier, so xAI can’t charge frontier prices. Second, xAI overbuilt compute early and then couldn’t ship a model good enough to fill those GPUs with demand. Cheap pricing is how you fill idle hardware.

He ties this to a Gavin Baker tweet showing enterprise token usage trending toward open-weights models at 62% versus 38% for closed models. Plenty of industries don’t need the best possible answer. They need automation that works and doesn’t bankrupt them. That’s the market Grok 4.7 is priced for.

3 practical applications

  1. High-volume knowledge work automation. Email triage, document summarization, task extraction, report drafting. Grok 4.7’s GDPval and briefcase scores say it’s near the frontier here, and the price makes it viable at scale. The creator’s sponsor segment even sketched a Gmail to Grok to Asana pipeline via Zapier as an example.
  2. Legal and compliance drafting. Grok 4.7 beat every major closed model on the legal benchmark. If you’re building anything around contracts, policy summaries, or regulatory text, this is worth a serious test. Just check Muse Spark too, since it scored double.
  3. Budget-conscious coding in Cursor. For everyday coding where you’re paying per token, Grok 4.7 gives you Opus-adjacent Cursorbench scores at half the cost. Keep the frontier models for the hard terminal-heavy agentic runs.

Tips and pitfalls

  • 💡 Crank the thinking effort. The gap between low effort (33%) and extra high (46.3%) on Cursorbench is one of the biggest curves the creator has seen. Low effort is cheap but noticeably weaker.
  • 💡 Watch total token burn. Artificial Analysis says 4.7’s improvements cost more tokens. Measure your actual cost per completed task, not the list price.
  • 💡 Don’t trust it in the terminal yet. At 38% on Terminal Bench, it’s far from Fable and Astra. One demo the expert showed had Grok 4.7 losing badly to Kimi K3 on a coding task, though settings were unclear.
  • Skip it for anything needing more than 500K tokens of context.
  • Remember the benchmarks on the blog were selected. The creator’s advice is to look at third-party sources like Artificial Analysis before deciding.

My take

Elon said a month ago that Grok 4.7 would “exceed all current models.” It didn’t. But the creator’s honest read is that it’s a great model for the price, roughly Opus 5 level on many tasks, and more competition is good for all of us. I think the cost story is the real headline here, and for a huge chunk of real-world work, good enough and cheap beats perfect and expensive.

Go watch the full video for the charts, the Astra-rebuilt table, and the demos. It’s worth the fifteen minutes.

Scroll to Top