Biphoo.eu - Guest Posting Services

collapse
Home / Daily News Analysis / Google is working on a new AI chip designed to make Gemini more efficient

Google is working on a new AI chip designed to make Gemini more efficient

Jul 24, 2026  Twila Rosenbaum  8 views
Google is working on a new AI chip designed to make Gemini more efficient

Alphabet, Google's parent company, is developing a new server chip designed to enhance the efficiency of its in-house Gemini artificial intelligence models. The chip, internally referred to as Frozen v2, is expected to launch sometime in 2028, according to a report from The Information, which cited anonymous sources familiar with the matter. The report claims that the new chip could be between six and ten times more efficient than Google's existing AI chips, measured by the number of tokens generated per unit of power.

Google did not directly confirm or deny the report when contacted by TechCrunch. In a statement, a company spokesperson said, "Our teams are constantly researching and experimenting with new innovations to deliver maximum performance and efficiency for our users and customers. While not every project moves into production, this rigorous exploration is central to our full stack approach. By co-designing our hardware and software from the ground up, we ensure our systems are integrated and highly optimized for real-world workloads."

Background on Google's AI Chip Strategy

Google has a long history of designing custom silicon for AI workloads. The company's Tensor Processing Units (TPUs) were first introduced in 2015 specifically for internal machine learning tasks and later made available to cloud customers via Google Cloud. The TPU series has evolved through multiple generations, with the latest TPU v5p and v5e chips offering significant improvements in training and inference performance. However, the Frozen v2 project appears to be a separate initiative focused specifically on inference—the process of running a trained model to generate outputs—rather than training.

The drive for greater efficiency comes as AI companies increasingly grapple with the immense computational costs of running large language models like Gemini. Google's Gemini models, which power everything from search enhancements to the Gemini chatbot and enterprise AI services, require massive amounts of computing power. Optimizing inference efficiency can directly reduce operational costs and energy consumption, making AI more sustainable and accessible.

Industry Context: The Race for Custom AI Chips

Google is not alone in pursuing custom chips for AI. The broader industry has seen a surge in chip development as companies seek to optimize performance and reduce dependence on Nvidia, which currently dominates the AI chip market with its H100 and upcoming B200 GPUs. Nvidia's hardware is widely used for both training and inference, but the company's pricing and supply constraints have pushed major tech firms to explore alternative architectures.

In June, OpenAI announced its first custom chip, an inference processor codenamed Jalapeño. Earlier this month, reports emerged that Anthropic was in discussions with Samsung about a new chipmaking partnership. These moves reflect a strategic shift: instead of relying solely on off-the-shelf chips from Nvidia or AMD, AI companies are designing custom silicon that can be tailored to their specific software stacks and model architectures.

The benefits of custom chips extend beyond cost savings. By integrating hardware and software, companies can achieve higher performance per watt, faster time-to-market for new model features, and tighter security. For Google, the Frozen v2 chip likely aims to complement its existing TPU lineup, potentially focusing on latency-sensitive inference applications where Gemini must respond in real-time, such as in search, voice assistants, and code generation.

Financial Implications and Investor Sentiment

News of the Frozen v2 chip appears to have resonated positively with investors, contributing to a 3% rise in Alphabet's stock on Monday morning ahead of the company's earnings report later in the week. The stock boost suggests that investors are encouraged by Alphabet's efforts to control costs and improve efficiency in its AI operations. Alphabet has previously announced plans to spend between $180 billion and $190 billion in capital expenditures over the coming years, much of it dedicated to AI infrastructure. With such massive investments at stake, demonstrating a clear path to better efficiency helps justify the spending and calms concerns about unsustainable AI costs.

The broader market euphoria around AI has cooled in recent months as questions about monetization and energy consumption have emerged. Companies like Google, Microsoft, and Amazon are under pressure to show that their huge investments will yield profitable returns. Custom chips that significantly improve efficiency—potentially reducing the number of GPUs or servers needed per query—can be a key differentiator.

Technical Details: What Makes a Chip Efficient?

Efficiency in AI chips is typically measured in terms of tokens generated per watt or per joule of energy. A token is a unit of text (roughly a word or subword) that a language model processes. Inference workloads are often memory-bandwidth bound, meaning the chip's ability to move data between memory and compute units is critical. Google's Frozen v2 is reportedly designed to address this bottleneck through novel memory architectures and optimized data paths, possibly leveraging advanced packaging techniques like 3D stacking or chiplets.

The reported 6x to 10x efficiency gain suggests a leap beyond incremental improvements. Achieving such a jump might involve a combination of architectural changes: specialized matrix multiplication units, reduced precision arithmetic (e.g., 4-bit or 8-bit floating point), sparse computation, and tighter integration with Google's software framework, including TensorFlow and the newly optimized Pathways system used by Gemini. The chip may also incorporate on-chip SRAM or near-memory computing to reduce latency and energy consumption.

Competitive Landscape: Google vs. Nvidia vs. Startups

Nvidia remains the undisputed leader in AI hardware, with its CUDA ecosystem and its GPUs powering most large-scale AI deployments. However, custom chips from Google (TPU, now Frozen v2), Amazon (Trainium and Inferentia), Microsoft (Azure Maia), and others are nibbling at the market. Startups like Groq, Cerebras, and SambaNova have also introduced specialized inference chips that claim dramatic efficiency improvements, but they lack the scale and software integration of the hyperscalers.

Google's advantage lies in its vertical integration: it controls both the models (Gemini) and the hardware (TPU and potentially Frozen v2). This allows for co-design where the chip is built specifically to run Gemini workloads efficiently. For example, if Gemini's architecture uses a particular size of matrix or attention mechanism, the chip can be optimized accordingly. In contrast, Nvidia's products must cater to a broad range of models and frameworks, which can dilute specialization.

The report of Frozen v2's 2028 timeline also suggests that Google is planning for next-generation models that may require vastly different compute characteristics. By 2028, AI models are expected to be far larger and more complex, possibly requiring new memory and interconnect technologies. Google's early investment in custom silicon positions it to have hardware ready when those models emerge.

Challenges and Risks

Developing a new chip is a multi-year, multi-billion-dollar endeavor. The project may face delays, cost overruns, or design flaws that prevent it from reaching production. Google's statement that "not every project moves into production" is a tacit acknowledgment of the risks. Moreover, even if Frozen v2 is successfully manufactured, it must be integrated into Google's vast data center infrastructure and seamlessly work with existing systems. Software compatibility and performance validation are nontrivial tasks.

Another risk is that by the time Frozen v2 arrives in 2028, Nvidia or other competitors may have already made similar leaps in efficiency. The AI chip market evolves rapidly, with architecture cycles accelerating. Google must ensure that its custom silicon remains competitive at launch and can be iterated upon quickly.

Broader Implications for the AI Industry

The trend toward custom AI chips has significant implications for the entire AI supply chain. Hardware diversity can reduce single-supplier risks and foster innovation. It may also drive down costs for AI inference, enabling new applications in areas like edge computing, autonomous vehicles, and healthcare. For Google, efficiency gains translate directly to lower operational expenses and improved profit margins on AI services such as Google Cloud AI and Google Workspace integrations.

Moreover, custom chips can be designed with specific privacy and security features, such as encrypted memory regions or hardware-backed attestation, which are important for enterprise customers handling sensitive data. As AI becomes more embedded in critical infrastructure, such features will become increasingly important.

Alphabet's stock movement following the Frozen v2 leak indicates that investors are watching these developments closely. With the earnings report approaching, the company will likely provide more color on its AI hardware roadmap and capital expenditure plans. The efficiency narrative could be a key theme in Alphabet's messaging to shareholders, highlighting how hardware innovation supports long-term margin improvement.

The Frozen v2 chip represents a strategic bet on vertical integration and efficiency. If successful, it could give Google a competitive edge in delivering cost-effective, high-performance AI at scale. The next few years will reveal whether this bet pays off, but the early signs are promising: a 3% stock bump is a small but meaningful vote of confidence in Google's ability to innovate beyond software.


Source: TechCrunch News


Share:

Your experience on this site will be improved by allowing cookies Cookie Policy