YYaaa News

Google bets on Frozen v2 chip | 6-10x tokens per watt for Gemini, targeting 2028

TL;DR

The Information reports Google is building a Gemini-specific inference chip codenamed Frozen v2 that etches the model architecture into silicon. Engineers project 6-10x tokens per watt versus current TPUs, targeting 2028. The project is driven by an internal compute shortage severe enough that Google Cloud is turning away outside customers.

The Information reported on July 20 that Google is building a new AI inference chip codenamed Frozen v2 for Gemini — with the model's underlying architecture etched directly into silicon. Engineers project 6 to 10 times the tokens per unit of power versus current TPUs, with a launch target as early as 2028. Alphabet shares rose on the news.

Frozen v2 is not the shelved "original Frozen" design that fused model weights into the chip. It is a flexible variant that etches only the architecture, so weight updates don't require a new chip, extending useful life. The project is driven by an internal compute shortage severe enough that Google Cloud is now turning away new large outside customers. Google is positioning Frozen v2 as a new branch separate from its TPU line.

The direction points at AI silicon sliding from "general compute" toward "model-family-specific inference." Against Gemini's scale, Nvidia GPU's general-purpose advantage is being pressed by Google's most extreme vertical integration — one chip, one model family. Google has not officially confirmed the project.

via Crypto Briefing / CNBC / The Information
Google 押注 Frozen v2 專用晶片|為 Gemini 硬編架構,目標 6 到 10 倍效率、2028 上線