Running a large AI model at scale has become less a scientific feat than a brutal math problem about electricity bills and datacenter rent. Every token Gemini generates costs Google something real—silicon cycles, memory bandwidth, and watts drawn from the wall. As query volume grows, fractions of a cent add up to sums that can swallow margins whole. That quiet urgency is behind Frozen v2, an internal server chip project now taking shape inside Google. Rather than refining its general-purpose Tensor Processing Units for another generation, the company is attempting something far more radical: casting the skeleton of the Gemini model directly into the silicon itself.
From Flexible Accelerators to Model-Specific Silicon
Google’s TPUs have been the workhorses of its infrastructure for nearly a decade. They train models, power search ranking algorithms, and even rent out by the hour to cloud customers including Meta and others looking for an alternative to Nvidia’s GPUs. That versatility is exactly what makes a TPU a TPU. It speaks a general vocabulary of matrix multiplication and memory movement, usable by almost any neural network you can describe in software.
Frozen v2 deliberately trades that flexibility away. The chip is being designed as a domain-specific accelerator whose circuits physically mirror portions of Gemini’s own architecture. Where a TPU fetches instructions and interprets them as software operations, Frozen v2 would burn the model’s structural blueprint—the arrangement of its layers and data paths—directly into the chip’s layout. Google expects this tight marriage of model and metal to make the chip six to ten times more efficient than current TPUs at serving AI responses. Fewer compute steps per query means less time waiting for a token to appear, and far less energy spent generating it.
This is not merely a faster version of the same idea. It is a different category of chip, one that trades generality for devotion to a single model family.
Why the First “Frozen” Melted
This approach has roots in an earlier concept attributed to Jeff Dean, Chief Scientist at Google DeepMind. The original “Frozen” proposal suggested pushing specialization even further by hardcoding not just the architecture, but the actual model weights—the billions of tuned parameters that constitute Gemini’s learned behavior—directly into the chip itself.
The logic was sound. If you know exactly which numbers the model will use, why bother fetching them from external memory? You could etch them into the transistors and eliminate entire categories of delay.
The problem was permanence. AI models do not stand still. Google updates Gemini continuously, retraining on new data, adjusting parameters, releasing improved versions. A chip with weights frozen in silicon would become a paperweight the moment a new model revision shipped. That lack of flexibility killed the original concept.
Architecture Without the Anchor
Frozen v2 solves the obsolescence trap by hardcoding the architecture while leaving the weights free to change. Think of it as pouring a custom racetrack instead of welding the car to the road. The shape of the circuit stays fixed, optimized for Gemini’s specific patterns of computation, but the contents flowing through those circuits can be refreshed by loading new weights from memory.
This distinction matters in practice. When engineers train a new Gemini checkpoint, they can deploy it to Frozen v2 hardware without fabricating a new chip. The exact degree of hardcoding is still an open question inside Google; teams must decide precisely which structural elements deserve silicon immortality and which should remain configurable. But the principle is settled. By freezing the shape and fluidly swapping the parameters, Google keeps the efficiency upside without sacrificing the ability to iterate.
The Economics of Keeping It In-House
There is another reason you will not see Frozen v2 listed on Google Cloud’s pricing sheet. Because the chip is so intimately shaped around Gemini’s internals, it would make little sense to outside developers running PyTorch or custom Transformer variants. Google has no plans to sell it as a general-purpose product. It will remain an internal tool, aimed squarely at the crushing demand for inference capacity inside Google’s own datacenters.
Lựa chọn đó phản ánh một thực tế kinh tế phũ phàng. Trong thị trường AI tạo sinh hiện nay, khả năng của các mô hình đang hội tụ rất nhanh. Khoảng cách giữa các đối thủ cạnh tranh thường nằm ở việc ai có khả năng vận hành mô hình lớn nhất với chi phí trên mỗi token thấp nhất. Suy luận (inference) không còn là một yếu tố phụ sau quá trình huấn luyện; đối với một sản phẩm được sử dụng rộng rãi như Gemini, đây là khoản chi phí chiếm ưu thế. Nếu Frozen v2 cắt giảm được khoản chi phí đó gấp sáu lần hoặc hơn, Google sẽ có được dư địa mà các đối thủ không dễ dàng theo kịp. Google có thể giữ lại khoản tiết kiệm đó dưới dạng lợi nhuận biên hoặc chuyển sang giảm giá cho người dùng API và các tích hợp sản phẩm, từ đó gây áp lực lớn lên OpenAI, Anthropic và các bên khác.
Điều này báo hiệu gì cho ngành công nghiệp
Bước đi của Google cũng gợi mở về hướng đi của chiến lược phần cứng rộng lớn hơn. Trong nhiều năm, kịch bản tiêu chuẩn là xây dựng bộ tăng tốc linh hoạt nhất có thể và để phần mềm đảm nhận việc chuyên biệt hóa. Các GPU của Nvidia thống trị vì chúng có thể chạy mọi thứ, từ động lực học phân tử đến trò chơi điện tử cho đến các mô hình ngôn ngữ lớn. Các TPU của chính Google cũng được hình thành với tinh thần tiện ích rộng rãi tương tự.
Frozen v2 phá vỡ truyền thống đó. Đó là một sự thừa nhận rằng khi một dòng mô hình duy nhất tạo ra đủ khối lượng truy vấn, các chip silicon tùy chỉnh được thiết kế riêng cho mô hình đó có thể mang lại lợi nhuận gấp nhiều lần chi phí đầu tư. Các nhà cung cấp dịch vụ đám mây quy mô lớn (hyperscalers) khác cũng theo đuổi logic tương tự—ví dụ như chip Trainium và Inferentia của Amazon—nhưng cách tiếp cận của Google đi sâu hơn bằng cách đồng thiết kế phần cứng xoay quanh một kiến trúc mô hình cụ thể thay vì một lớp mạng tổng quát.
Tất nhiên, rủi ro nằm ở sự cứng nhắc. Nếu kiến trúc của Gemini tiến hóa theo một hướng mà các mạch được lập trình cứng không thể đáp ứng, Google có thể rơi vào tình cảnh sở hữu những con chip đắt đỏ nhưng không thể chạy được các ý tưởng mới nhất của mình. Đó chính là lý do tại sao sự thỏa hiệp chỉ tập trung vào kiến trúc lại quan trọng. Nó mang lại một con đường trung dung: đủ sự chuyên biệt hóa để đạt được những bước tiến vượt bậc về hiệu quả, nhưng cũng đủ linh hoạt để tránh đẩy công ty vào thế bí.
Bài học thực sự rút ra
Frozen v2 nên được hiểu không phải là một thông báo về chip, mà là một canh bạc chiến lược về hình thái cạnh tranh AI trong tương lai. Google đang đặt cược rằng những người chiến thắng sẽ không chỉ xây dựng được các mô hình tốt nhất, mà sẽ sở hữu toàn bộ hệ thống (stack)—từ bản thiết kế mô hình cho đến các electron chạy qua bóng bán dẫn. Nếu dự án thành công, thành quả sẽ không xuất hiện trong các điểm số benchmark. Nó sẽ xuất hiện trong cột chi phí của báo cáo thu nhập hàng quý, nơi mà việc tiết kiệm được vài cent trên mỗi triệu token có thể vẽ lại ranh giới của những gì khả thi về mặt thương mại trong lĩnh vực AI tạo sinh.
