cnbc.com web signal

Alphabet gains on report of Gemini-tuned 'Frozen v2' chip

TL;DR

  • Alphabet shares climbed about 3% on Monday after The Information reported Google is developing a next-generation AI server chip code-named Frozen v2.
  • Google engineers project the chip could generate six to 10 times more AI tokens per unit of power than the company's current TPUs.
  • Deployment is targeted for 2028, with Google framing Frozen v2 as a complementary chip family rather than a replacement for its TPU lineup.

Google is reportedly designing a server chip that bakes part of the Gemini model architecture directly into the silicon, and if the numbers hold, it could serve six to ten times more tokens per unit of power than the company's current TPUs. That was enough to lift Alphabet around 3% on Monday, as CNBC reported off the back of reporting from The Information.

The chip, code-named Frozen v2, is targeted for deployment in 2028. The interesting design choice is what gets fixed. Instead of a general-purpose accelerator that can run any model, parts of Gemini's architecture are embedded into the hardware itself, reducing the computational steps and data movement a more flexible chip has to do. New weights can still be loaded, according to the reporting, so it is the shape of the model that is frozen, not the trained parameters.

Why anyone outside Google cares: inference cost is the binding constraint on how aggressively any AI company can price and scale its products. A six-to-10x tokens-per-watt improvement, if it holds up in production, is not the sort of gain you get from routine generation-on-generation upgrades. It would also help ease the internal compute shortages that have reportedly limited what Google Cloud can offer some enterprise customers.

The honest caveats are worth naming. Those efficiency figures are Google engineers' projections, not independently verified benchmarks, and the 2028 date is a target rather than a shipping calendar. The reporting also frames this partly as a trial run: Google reportedly does not plan to produce Frozen v2 at the same scale as its TPUs, and the chip only pays off with future Gemini models if Google sticks with the same underlying architecture, which is a real bet on a specific direction. What the reporting does not give you is how much of the silicon is actually frozen versus reprogrammable, or how the design would cope if a future Gemini needs a different shape.

If it works, the immediate beneficiaries are Google Cloud and any customer whose bill is dominated by inference. The wider read is that hyperscaler custom silicon is starting to move beyond a TPU-or-two-on-the-side story and toward hardware co-designed with a specific model, which is a different competitive picture than one where everyone rents the same generic GPUs.

Shared on Bluesky by 2 AI experts