Google Bets on a Custom Chip to Run Gemini Leaner

Alphabet is designing a new server chip built to run its Gemini models more efficiently, according to TechCrunch AI. The chip, internally nicknamed “Frozen v2,” is slated for release sometime in 2028, based on a report from The Information that TechCrunch AI cited. The headline claim: it could be six to 10 times more efficient than Google’s current AI chips, measured by tokens generated per unit of power.

Google didn’t confirm the report. It didn’t deny it either. “Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers,” the company told TechCrunch AI, adding that its “full stack approach” of co-designing hardware and software is central to how it works.

Why this matters

Efficiency is the new battleground. For the past two years, the AI story was about raw capability and who could train the biggest model. That’s shifting. As worries about AI spending cool the earlier market euphoria, tokens-per-watt has become a real selling point. A chip that squeezes six to 10 times more output from the same power isn’t a lab curiosity. It’s a direct lever on the cost of running Gemini at scale.

What stands out here is the timing and the money behind it. Google has said it plans to spend between $180 billion and $190 billion this year on its AI push. Investors have been nervous about that number, and for good reason. Spending that much only works if the models get cheaper to operate over time. Frozen v2 is Google’s answer to that pressure, and the market noticed. After The Information’s report, Alphabet stock climbed roughly 3% on Monday morning, right before the company’s earnings report later this week.

The bigger pattern: escaping Nvidia

Google isn’t acting alone. AI companies are racing to build their own silicon for two reasons:

  • Efficiency. Custom chips tuned for a company’s specific models run those models faster and cheaper than general-purpose hardware.
  • Independence. Nvidia has dominated the AI chip market for years, and that dominance left the major AI makers dependent on one supplier for both hardware and pricing.

The moves are stacking up. In June, OpenAI announced its first custom chip, an inference processor called Jalapeño. Earlier this month, TechCrunch AI notes, Anthropic was reported to be discussing a new chipmaking partnership with Samsung. Google already has a head start here through its long-running TPU program, so Frozen v2 is less a first step than a next gear.

A quick note on the technical framing. “Tokens per unit of power” measures how much text a chip can generate for a given amount of electricity. Inference, the work of actually running a model to answer your prompt, is where the ongoing cost lives once a model is trained. Making inference cheaper is what turns an expensive research model into a product you can serve to billions of users without lighting money on fire.

What to expect next

A few things worth watching:

  1. 2028 is the target, not a promise. Google itself flagged that “not every project moves into production.” Treat the timeline as a direction, not a shipping date.
  2. Earnings will set the tone. With the report landing days before Alphabet’s results, expect chip efficiency and capital spending to come up on the call. The narrative Google wants is simple: heavy spend now, cheaper models later.
  3. The custom-silicon arms race accelerates. Every major lab building its own chips chips away at Nvidia’s grip and reshapes who controls AI’s cost curve.

My take: the interesting shift isn’t the chip itself, it’s what it signals. The industry is moving from “can we build it” to “can we afford to run it.” Efficiency is where the next phase of competition gets decided, and Google is telling investors it intends to win on that front.

More details are available in the original TechCrunch AI report.

Scroll to Top