Google Is Building a Chip That Has Gemini Baked Right Into the Silicon

Google is developing a new server chip that would incorporate elements of its Gemini model directly into the hardware, aiming to serve its AI models more efficiently to users, The Information reported Monday, citing people familiar with the matter. FinancialMediaGuide views the project as a sign that the biggest AI labs are no longer treating chip design and model design as separate disciplines, but are increasingly building the two together from the ground up.

The Alphabet-owned company expects the new chip, informally dubbed “Frozen v2,” to help address an AI computing capacity crunch that has fueled internal tensions and prompted Google Cloud to decline deals with outside customers, according to the report. Shares of Alphabet rose 3.3% in early trading following the news.

Google plans to deploy the chip as soon as 2028, though engineers are still finalizing its design and the amount of model information that will be hardwired into it, the report said. The chip could be six to 10 times more efficient than Google’s latest custom AI chips, measured by the number of AI tokens served per unit of power. FinancialMediaGuide notes that an efficiency gain of that magnitude, if it materializes, would represent one of the larger single-generation leaps reported anywhere in the AI hardware race so far.

“Our teams are constantly researching and experimenting with new innovations… By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized,” a Google Cloud spokesperson said, without directly confirming the chip’s codename or specifications.

The Frozen project is aimed at creating a new set of homegrown chips separate from Google’s existing tensor processing units, rather than replacing them, according to the report. Financial Media Guide points out that maintaining two parallel custom-chip programs, rather than consolidating into one, signals Google sees enough differentiated future demand, from general-purpose AI workloads on one side and Gemini-specific inference on the other, to justify running both lines simultaneously.

The chip news comes as Google works through a separate, well-documented capacity strain: the company delayed the launch of its latest Gemini AI model earlier this month after it fell short of internal goals, with teams working to improve its capabilities, particularly in coding. The capacity crunch cited in connection with Frozen v2 appears directly linked to the same computing bottlenecks that have complicated Google’s broader Gemini roadmap this year.

If Google can bring a Gemini-optimized chip to market by 2028 at the efficiency gains being described, it would give the company a meaningful edge in serving inference workloads more cheaply than rivals reliant on general-purpose accelerators. FinancialMediaGuide concludes that Frozen v2 is best understood not as a replacement for Google’s TPU strategy but as a parallel bet that model-specific silicon, rather than increasingly general-purpose chips, may be where the next efficiency gains in AI infrastructure actually come from.

Share This Article