Breaking News
Loading latest updates...

Why Global AI Giants Are Racing to Build Custom Inference Chips

Why Global AI Giants Are Racing to Build Custom Inference Chips


Controlling the Stack: Why Leading AI Firms Are Building Their Own Silicon

A definitive pattern is emerging across the artificial intelligence landscape: the world's leading model developers, regardless of their open-source or closed-source philosophies, are simultaneously entering the capital-intensive semiconductor business. From Silicon Valley to Beijing, the race is no longer just about training smarter software—it is about controlling the physical silicon underneath it.

Recent market intelligence highlights the scale of this shift:

  • DeepSeek's Secret Project: On July 7, reports revealed that DeepSeek has spent the last year quietly expanding an in-house chip design team. The company is actively engaging with chip design firms, foundries, and memory manufacturers to develop a proprietary chip tailored specifically for large model inference.

  • Zhipu AI Evaluating Custom Silicon: The same day, news broke that Zhipu AI is evaluating the feasibility of custom AI chip development to sustain the growing enterprise demand for its GLM series models.

  • OpenAI's "Jalapeño": On June 24, OpenAI publicly unveiled its first custom inference chip, codenamed Jalapeño. Co-designed with Broadcom, samples are already running in the lab, with deployment planned by the end of the year.

  • Anthropic's Talent Poaching: In April, Anthropic was reported to be exploring in-house chip design, a move signaled clearly in June when they hired a key lead engineer away from OpenAI's chip project.

The Straightforward Economics of Everyday Inference

While training frontier models demands massive upfront capital, inference costs accumulate continuously. Every user interaction, every running autonomous Agent, and every token generated incurs a compounding infrastructure cost.

By designing in-house chips, firms like DeepSeek and Zhipu AI are prioritizing inference hardware over training hardware. Custom silicon allows them to optimize directly for their specific Mixture of Experts (MoE) architectures, key-value (KV) cache management, and low-precision computation paths.

The industry-wide bifurcation of hardware is accelerating to meet this demand. For example, Google split its eighth-generation TPU into the 8t (built for massive parallel training) and the 8i (optimized for low-latency inference). Huawei followed a similar playbook by dividing its Ascend 950 series.

Operational Control vs. Extreme Financial Risk

For Chinese AI firms, self-developed chips address a distinct strategic vulnerability. With tightening US export controls restricting the supply of advanced NVIDIA GPUs, domestic and in-house silicon has transformed from a cost-saving metric into a necessity for operational survival.

However, venturing into hardware carries immense risk. Industry estimates suggest that designing a single advanced AI chip costs approximately $500 million, and engineering success is far from guaranteed.

Because a full-stack hardware approach like Google's TPU network is incredibly difficult to execute, many model companies are choosing a hybrid path. By defining the internal architecture themselves while partnering with established semiconductor firms like Broadcom to handle the heavy lifting—a strategy exemplified by OpenAI’s Jalapeño—AI firms can mitigate risk.

Ultimately, these moves reflect a broader industry realization: to survive at the competitive frontier, companies must control the entire technology stack from the silicon up to the model.