On June 24, 2026, OpenAI announced Jalapeño — a custom inference chip built with Broadcom, taped out in just nine months. One week later, on July 7, Reuters cited three sources reporting that DeepSeek is developing its own AI inference chip, launched roughly a year ago and still in early stages. On the same day, The Information reported that Zhipu AI is also evaluating custom silicon. Meanwhile, Alibaba's T-Head has already shipped 560,000+ Zhenwu chips with billion-yuan annual revenue. This is not a China story or a geopolitics story — it is a unit economics story. This article covers the DeepSeek rumor evidence chain, Liang Wenfeng's actual public statements, Alibaba T-Head's eight-year roadmap, the global custom chip wave, five structural drivers, and why inference — not training — is where custom ASICs win.
- Bottom line: DeepSeek chip is credible but early-stage; Alibaba is already shipping at scale; OpenAI just deployed Jalapeño — custom silicon is now a mainstream infrastructure choice.
- Primary driver: Economics. Inference is AI's permanent monthly rent. Custom ASICs can cut TCO by 30–65% at scale versus general-purpose GPUs.
- Liang Wenfeng did not announce a chip program. Reuters reported company actions (hiring, foundry outreach) — not a founder declaration. Distinguish carefully.
- Alibaba T-Head is not a rumor. Jack Ma named it in 2018; eight years later the Zhenwu 810E is in mass production. DeepSeek is where Alibaba was in 2019.
- Disclaimer: DeepSeek has not officially confirmed the chip project as of this writing (July 10, 2026).
SECTION 01 What Reuters Actually Reported (And What DeepSeek Hasn't Confirmed)
The July 7, 2026 Reuters exclusive — followed by coverage from Bloomberg, WSJ, and others — reports a consistent set of facts. DeepSeek has not issued any press release or social media confirmation as of this writing.
- Target use case: Inference (serving), not training. This matches the broader industry trend.
- Timeline: Project reportedly started approximately one year before the report — around mid-2025.
- Status: Early stage. Contacts with chip designers, foundries, and memory suppliers underway.
- Hiring: Chip design engineers recruited privately, not through public job boards.
- Strategic implication: Reducing dual dependency on both Nvidia and Huawei Ascend — notable because DeepSeek V4 already runs on Ascend hardware.
| Evidence type | Assessment |
|---|---|
| Source quality | High. Reuters "three people familiar with the matter" is the standard phrasing for verified, cross-checked enterprise intelligence. |
| Official confirmation | None. DeepSeek has not issued any statement as of July 10, 2026. |
| Funding disclosure | Strong. June 2026 funding round of ~¥51B (~$7.4B) disclosed use of funds including "in-house AI chip" and "domestic compute expansion." |
| Model architecture signal | Moderate. DeepSeek's UE8M0 FP8 format and MLA architecture are interpreted by industry observers as hardware co-design signals for non-Nvidia targets. |
| Conflicting narrative | Some analysts argue DeepSeek's deep Ascend partnership suggests cooperation is prioritized over self-development. Most accurate view: both tracks run in parallel. |
| Date | Event |
|---|---|
| 2023–2024 | Liang Wenfeng in interviews: export controls are the biggest challenge; compute hunger is permanent |
| Jan 2025 | DeepSeek R1 released; trained on Nvidia H800 (banned from export since late 2023) |
| Mid-2025 | Chip project reportedly initiated |
| Apr 2026 | DeepSeek V4 adapted for Huawei Ascend; V4-Flash partial training on Ascend |
| Jun 2026 | First external funding round ~$7.4B; stated use includes in-house chip |
| Jul 7, 2026 | Reuters exclusive: DeepSeek developing custom inference chip, hiring engineers, contacting foundries |
| Jul 7, 2026 | The Information: Zhipu AI also evaluating custom chip (same day) |
Accurate framing: "According to Reuters and other outlets, DeepSeek has initiated a custom inference chip project." Inaccurate: "Liang Wenfeng announced DeepSeek will build chips." The distinction matters for credibility.
SECTION 02 What DeepSeek CEO Liang Wenfeng Has Said About Chips and Compute
Liang Wenfeng rarely gives interviews. The most authoritative source is two in-depth conversations with Chinese media outlet Angang Waves in May 2023 and July 2024. He has never publicly announced a chip development program. What he did establish is the strategic rationale:
- "The real challenge isn't money — it's export controls on advanced chips." — July 2024, Angang Waves. This is the single most-cited quote in the context of DeepSeek's hardware constraints.
- The 4x compute gap: "China's best level versus overseas has roughly a 1x training efficiency gap and a 1x data efficiency gap — combined, about 4x the compute needed to achieve the same result." This arithmetic made self-sufficiency a strategic imperative, not just a national goal.
- "China must stand at the frontier." "Many domestic chips fail because they lack a technical community — only second-hand information. China must have people at the frontier." This is a statement about ecosystem, not a chip announcement.
- Permanent compute hunger: "For researchers, the desire for compute is endless. We will consciously deploy as much compute as possible." A strategic intent, not a product roadmap.
Reuters reported on company actions (hiring, foundry contacts) — not a founder declaration. Liang's statements establish the motivation; they are not the announcement. This distinction matters when writing accurately about DeepSeek.
SECTION 03 Alibaba's T-Head Is Already Shipping — Jack Ma's 2018 Bet Pays Off in 2026
Alibaba's chip story is the clearest counterpoint to "DeepSeek rumor." It is not a rumor. It is an eight-year strategic investment now in mass production.
| Person | Role | Key statement on chips |
|---|---|---|
| Jack Ma | 2018 strategic decision-maker | Named T-Head (Pingtouge) at 2018 Yunqi Conference; elevated chips to group strategy. Stepped down as chairman in 2019. |
| Joe Tsai | Current chairman | 2024 podcast: US chip export limits "clearly impact" Alibaba Cloud; "long-term belief China will develop advanced domestic semiconductors"; export controls were a reason Alibaba Cloud spin-off was paused. |
| Eddie Wu (Wu Yongming) | Current CEO | FY2026 earnings call: T-Head cumulative shipments 560K+, billion-yuan annual revenue; T-Head IPO not ruled out. |
| Model | Release | Key specs and status |
|---|---|---|
| Hanguang 800 | 2019 | Early AI inference chip; established T-Head's commercial path |
| Zhenwu 810E | Jan 2026 | Training + inference; 96GB HBM2e; performance between Nvidia A800 and H20; in mass production |
| Zhenwu M890 | 2026 | 144GB VRAM; 800GB/s chip-to-chip; ~3× 810E performance |
| Zhenwu V900 | Planned Q3 2027 | 216GB VRAM; 1,200GB/s interconnect |
| Zhenwu J900 | Planned Q3 2028 | In-house parallel compute architecture iteration |
- Shipments: 560,000+ cumulative (H1 2026)
- Revenue: Annual run-rate in the billions of yuan
- Customers: Alibaba Cloud internally, China Unicom, 400+ enterprise customers
- CUDA compatibility: WSJ reports Zhenwu chips are designed to be compatible with Nvidia CUDA ecosystem, lowering migration friction — a major strategic difference from Huawei's Ascend approach
- Manufacturing: Moved from TSMC to domestic foundries (SMIC 7nm and mature nodes) to navigate US restrictions on TSMC serving advanced AI chips to China
- Capex: Alibaba announced ¥380B ($52B) investment over three years in cloud and AI infrastructure
SECTION 04 This Isn't Just China: OpenAI's Jalapeño and the Global Custom Chip Wave
Framing this story as a China-only narrative misses 80% of the picture. TrendForce data for 2026: hyperscaler custom AI chip shipment growth is 44.6% versus 16.1% for general-purpose GPUs — custom silicon is now outgrowing the GPU market for the first time.
| Company | Chip project | Stage | Key data point |
|---|---|---|---|
| DeepSeek | Custom inference ASIC (unnamed) | Early R&D | $7.4B funding; private hiring; not confirmed |
| Alibaba (T-Head) | Zhenwu 810E / M890 | Mass production | 560K+ shipped; billion-yuan annual revenue |
| Huawei | Ascend 950+ | Mass production | DeepSeek V4 adapted; surging orders |
| OpenAI | Jalapeño (with Broadcom) | Tape-out complete, deployment late 2026 | 9 months design to tape-out; ChatGPT inference target |
| TPU v6 / v7 | Large-scale commercial use | End-to-end Gemini workloads on TPU | |
| Amazon | Trainium3 / Inferentia | Commercial | Anthropic large-scale Trainium usage |
| Microsoft | Maia 100 | Deploying | Azure / OpenAI inference workloads |
| Anthropic | Samsung custom chip talks | Exploratory | The Information, July 2026 |
| Zhipu AI | Evaluating custom chip | Early | The Information, July 7, 2026 — same day as DeepSeek report |
The official OpenAI Jalapeño announcement is one of the most significant custom chip disclosures of 2026.
OpenAI — Jalapeño Inference Chip Official Announcement
SECTION 05 Why Tech Giants Build Custom AI Chips: Cost, Control, and the "Nvidia Tax"
One-sentence answer: AI competition has extended from "who has the best model" to "who has the cheapest, most controllable compute." Here are the five structural drivers, ranked by impact:
-
Economics: inference is AI's permanent monthly rent
Training = the down payment (one-time, concentrated). Inference = monthly rent (ongoing, scales linearly with users). When a ChatGPT-class product has hundreds of millions of daily active users, inference spend exceeds training spend. Morgan Stanley estimated: a 24,000-GPU Blackwell cluster costs ~$852M in hardware; an equivalent Google TPU cluster costs ~$99M. SemiAnalysis and Bernstein estimate custom ASICs can achieve 40–65% TCO advantage over GPUs in large-scale, multi-year inference deployments, with 30–40% lower cost per token. Nvidia data center GPU gross margins exceed 70%. Every H200 purchased means most of the margin goes to Nvidia. Custom chips convert that permanent "GPU tax" into a one-time R&D investment. -
Supply chain security and geopolitics
US export controls on H100/H800/H20 and subsequent generations hit Chinese AI labs repeatedly. Chinese regulators encourage domestic compute procurement. Even US companies face Nvidia GPU allocation constraints. Supply chain security means predictability: not being dependent on a single vendor or a single nation's export policy. -
Hardware-software co-design
DeepSeek's UE8M0 FP8 and MLA architecture are optimized for specific hardware characteristics. OpenAI Jalapeño is designed around actual ChatGPT serving patterns (KV cache, batching, latency budgets). Google TPU is tightly coupled with TensorFlow and JAX. General-purpose GPUs sacrifice efficiency for flexibility; custom chips sacrifice flexibility for known-workload efficiency. -
Competitive moat and negotiating leverage
Even without fully replacing Nvidia, in-house chips add leverage in procurement negotiations, create differentiated compute story for cloud customers, and enable full-stack narratives ("model + cloud + chip"). -
Energy and sustainability
Inference chips optimize performance per watt. In megawatt-to-gigawatt data centers, power and cooling costs are comparable to chip acquisition costs. ASICs eliminate the large portion of GPU die area dedicated to general-purpose circuitry, significantly cutting power draw.
SECTION 06 Inference Chips vs Training GPUs: Why the Industry Is Splitting
Most custom silicon efforts target inference, not training. This is not coincidence — it is workload economics:
| Dimension | Training | Inference |
|---|---|---|
| Workload nature | Dynamic, experimental, architecture changes frequently | Static, fixed model, request patterns predictable |
| Software ecosystem | Deep CUDA moat (cuDNN, NCCL, Nsight) | Custom kernels feasible for fixed models; lower ecosystem dependency |
| Hardware requirement | Peak compute + flexible programmability | Throughput, latency, cost-per-token optimization |
| Economic scale | Large one-time cluster investment; finite runs | 24/7 ongoing; scales linearly with product user growth |
| Custom ASIC fit | Low (high flexibility requirement) | High (workload precisely predictable and tunable) |
| Representative chips | Nvidia H100/B200 (dominant, entrenched) | TPU, Trainium, Maia, Jalapeño, Zhenwu, DeepSeek (rumored) |
Training remains Nvidia's home turf. Inference is where custom ASICs are winning. Every new custom chip announced in 2026 — OpenAI, Alibaba, DeepSeek — targets inference.
This compute economics shift has direct implications for how AI workloads are deployed in production. Shared virtual machines introduce hypervisor scheduling overhead that offsets the per-token gains from custom silicon. For teams running high-density inference, 24/7 AI Agent workflows, or macOS-native workloads that need zero-overhead native compute, MACNOX dedicated physical Mac Mini M4 nodes provide bare-metal performance with root access and no virtualization tax — deployable by the day, week, or month. See AI Agent production hardening and Mac Mini M4 rent vs buy cost comparison, or go to the pricing page.
Reference sources for data cited in this article (verify links after publication):
WSJ — Alibaba AI Chip to Fill Nvidia Void (Zhenwu analysis)
Caixin Global — Analysis: Alibaba's New Processor Shows Applications Are Key to AI Chip Success
OpenAI — Jalapeño Inference Chip Official Announcement
Last updated: July 10, 2026. Disclaimer: This article is based on public media reports and cited sources. DeepSeek's chip project has not been officially confirmed as of this writing. Information remains in early stages — verify against the latest news before drawing conclusions.
SECTION 07 FAQ
Is DeepSeek really building its own AI chip?
According to a July 7, 2026 Reuters report citing three sources, DeepSeek is in the early stages of developing a custom chip for AI inference. DeepSeek has not officially confirmed the project. The $7.4B funding round disclosing "in-house AI chip" development is the strongest indirect evidence available.
Did DeepSeek CEO Liang Wenfeng announce a chip program?
No public announcement. In 2024 interviews he said export controls on advanced chips were DeepSeek's main challenge, not funding. Reuters reported on company actions — hiring and foundry outreach — not a founder declaration. These are meaningfully different claims.
How far along is Alibaba T-Head?
T-Head's Zhenwu 810E entered mass production in January 2026. By H1 2026, cumulative shipments exceeded 560,000 units with annual revenue in the billions of yuan. The roadmap extends through Zhenwu M890 (2026), V900 (2027), and J900 (2028). DeepSeek is roughly where Alibaba was in 2019.
Why are companies building inference chips first, not training chips?
Inference workloads are repetitive, predictable, and run 24/7 — ideal for custom ASIC optimization. Training still requires the deep CUDA ecosystem and extreme programmability where Nvidia dominates. Inference is also the ongoing "monthly rent" of AI products, making the economics of custom silicon more compelling at scale.
Is this about national security or saving money?
Both, but economics is the primary driver. Custom ASICs can reduce TCO by 30–65% versus GPUs at large scale, converting permanent Nvidia margins into one-time R&D. Export controls and supply chain risk accelerate adoption but are not the sole cause — even US hyperscalers building custom chips face no such constraints.