Google is building a server chip that embeds the structural blueprint of its Gemini model directly into the circuitry, bypassing the general-purpose flexibility that makes chips like TPUs useful for many things but less efficient for a single model family. The project, known internally as “Frozen v2,” sent Alphabet’s stock climbing as much as 3.7% intraday on Monday before closing the session up 1.51%.

Engineers working on the project project the chip could deliver between six and ten times the token output per watt compared with Google’s latest generation of Tensor Processing Units. Google has not officially confirmed the project. The 2028 deployment target is a reported goal, and the efficiency figures are projections from engineers familiar with the design, not independently benchmarked results.

Architecture vs. Weights

The original “Frozen” initiative, led by Google DeepMind Chief Scientist Jeff Dean, proposed burning Gemini’s specific trained numerical parameters — its weights — directly into the chip’s circuitry. That approach was set aside because a chip tied to a single trained model version would have a commercial lifespan measured in months.

Frozen v2 hardwires the architecture rather than the weights: the structural blueprint (transformer blocks, attention heads, hidden dimensions, feed-forward layout, normalization approach) stays fixed in silicon while weights remain updateable. A future Gemini release can run on the same hardware as long as Google keeps the underlying architectural structure intact.

Compute Crisis Context

“We are compute constrained in the near term,” Alphabet CEO Sundar Pichai said on the company’s first-quarter 2026 earnings call. Google Cloud’s order backlog roughly doubled in a single quarter to approximately 920 million per month for access to roughly 110,000 Nvidia GPUs as interim bridge capacity.

Not a Replacement, and Not for Sale

Frozen v2 is not a successor to the TPU line. Google’s current generation consists of the TPU 8t (training) and TPU 8i (inference), announced at Google Cloud Next in April 2026. Frozen v2 is a parallel track built for running Gemini-family models during inference only. Because its architecture is hardwired for Gemini, it almost certainly will not become a commercial product available to Google Cloud customers. Production volumes are expected to fall well short of TPU levels; Google currently views the project as an exploratory exercise.

A Google spokesperson told CNBC that teams are “constantly researching and experimenting with new innovations” and that “not every lab project” makes it into production.

Alphabet reports Q2 2026 earnings on Wednesday, July 22, and may confirm or provide further details about the project at that time.

Primary: https://www.techtimes.com/articles/321152/20260721/googles-frozen-v2-chip-hardwires-gemini-architecture-tenfold-inference-efficiency.htm