OpenAI just started fighting NVIDIA on its own turf

For years, the smart money said closed AI models would always crush the open ones. This week flipped that script, and it’s happening faster than anyone expected. I’ve been glued to a new video from Matt Wolfe, the creator behind Future Tools, who broke down the whole open-weight versus closed-weight battle, and I think it’s one of the most important shifts in AI right now.

Here’s the old way: everyone rents NVIDIA chips and runs closed models in someone else’s cloud. The new way looks very different, and the original poster laid out three moves that prove it.

🔦 What the creator spotted this week

  • OpenAI showed off its own inference chip, code-named Jalapeno. The author explains it hit up to 104x performance on public open-weight models, and OpenAI claims the lead widens on their own frontier models. Translation: they’re building chips to stop depending on NVIDIA for inference.
  • According to reporting the author cites, NVIDIA agreed to buy Hugging Face, the “GitHub for AI models.” Neither company has confirmed it publicly yet, so treat it as a rumor for now.
  • Apple dropped the M6 and M5 Ultra, built for local AI. The expert notes the M5 Ultra has 4.5x the AI GPU compute of his current M3 Ultra, with versions up to 512GB of unified memory.

The contrast that got me

The mind behind the video connects the dots really well. OpenAI, Meta, and Google are all building their own chips to own more of the stack. That could pull business away from NVIDIA. So NVIDIA is betting the other direction, on open models, and grabbing the place people actually run them. If open weights keep growing, NVIDIA becomes the default compute provider for everyone using them.

And the data backs the trend. The author points to a post from Vercel’s CEO showing open models jumped from 28% to 62% of token usage in two months. One caveat the creator is careful to flag: by number of requests, closed models still lead 62% to 38%. Open models just chew through way more tokens.

🧠 Two open models worth knowing

  • GLM-5.3-Flash: The industry pro says it scored 63.4 on a coding benchmark, beating Opus 4.8, while being far cheaper. He points out it made a whole cluster of pricier models look obsolete overnight.
  • Qwen3.8-Flash: A 125B parameter model, solid scores, and cheap to run. The author generated a test for just over a penny.

Quick hits the curator rounded up

  • Google’s Gemini Omni 1.1 Flash took the top spot in a blind text-to-video test.
  • Google also shipped Gemini 3.5 Transcribe, travel booking in AI Mode search, and ebook imports in Gemini Notebook.
  • Anthropic gave Claude memory that works across chat and Cowork, plus a browser rolling out soon.
  • OpenAI cut the price of its GPT-5.6 Sol model by 20% for three months.
  • Perplexity launched a local-first agent that runs models on your own machine.

My honest reaction

I was genuinely surprised watching this. The person who posted it admitted that three or four years ago he’d have bet against open weights ever catching up. Now the gap is tiny, and the hardware to run great models at home is almost here.

One more thing worth celebrating: the creator crossed a million subscribers this week, and you can feel how much he cares about keeping people looped in.

There’s way more detail packed into the full video, including his takes on AI music and that messy TIME100 AI list. Go watch it and see which shift matters most for you.

Scroll to Top