Google’s Ironwood TPU Expansion: 40% Growth, 3.2M Units, and the Nvidia Challenge

Google’s Ironwood TPU Expansion: 40% Growth, 3.2M Units, and the Nvidia Challenge

Google’s TPU expansion just crossed a threshold that threatens Nvidia’s AI chip dominance for the first time in a decade. The company is ramping Ironwood production to an estimated 3.2 million units by 2026, representing over 40% year-over-year growth from 2025’s 2.5 million unit baseline. This isn’t just capacity expansion—it’s a calculated assault on Nvidia’s inference dominance, backed by TSMC’s 3nm process, Samsung’s HBM3E supply, and a potential Meta partnership worth billions.

The stakes are clear: Google Cloud executives are positioning TPUs to capture 10% of Nvidia’s annual revenue, a multi-billion-dollar opportunity that hinges on production execution and ecosystem adoption. With Nvidia’s $20 billion Groq acquisition demonstrating how seriously the GPU giant takes alternative architectures, Ironwood’s success or failure will reshape inference economics for every hyperscaler over the next 18 months.

Production Metrics: The 3.2M Unit Build-Out

Google’s TPU expansion isn’t speculative—it’s happening at wafer scale. TrendForce projects that Google TPU shipments will exceed 40% annual growth in 2026, pushing total volume to approximately 3.2 million units. This builds on the 2.5 million units shipped in 2025, when TPU v7 (Ironwood) began ramping in Q2.

The manufacturing reality behind these numbers: Samsung Electronics supplies over 60% of Google’s HBM3E through Broadcom, with Samsung’s HBM4 recently passing validation tests to double supply volume in 2026. TSMC’s 3nm process provides the silicon foundation, with MediaTek requesting a 7-fold increase in CoWoS capacity for TPU orders. Revenue projections follow: Morgan Stanley forecasts the expanded production could inject $13 billion in new revenue and add $0.40 to earnings per share.

Ironwood Performance Specifications

Ironwood’s technical specs reveal why Google thinks it can compete on inference. Each TPU v7 chip delivers 4,614 FP8 TFLOPS with 192 GB of HBM3E memory across eight stacks, achieving 7.4 TB/s peak HBM bandwidth. The architecture splits each chip into two TensorCores and four SparseCores, optimized specifically for sparse model inference—where most LLM computation actually happens.

Metric TPU v7 (Ironwood) TPU v6e (Trillium) TPU v5p
Compute (FP8 TFLOPS) 4,614 ~1,150 ~460
Memory per chip 192 GB HBM3E 32 GB HBM2e 16 GB HBM2
HBM bandwidth 7.4 TB/s ~3.2 TB/s ~1.6 TB/s
Max pod size 9,216 chips 4,096 chips 4,096 chips
Total pod memory 1.77 PB 131 TB 65 TB

Google scales Ironwood to 9,216 liquid-cooled chips linked via Inter-Chip Interconnect at 9.6 Tb/s per link, delivering 42.5 Exaflops of compute power and 1.77 Petabytes of shared HBM. That’s more than 13x the aggregate memory of a TPU v6e pod, critical for fitting trillion-parameter models in memory without constant weight streaming. Performance multipliers: 10x peak performance over TPU v5p and 4x better than TPU v6e, according to Google’s official announcement.

Strategic Positioning vs. Nvidia

Google’s targeting Nvidia’s inference moat, not its training dominance. The economics matter: inference cost analysis shows GPU-based serving burning $6.32 billion annually on inefficiencies that TPU’s sparse architecture sidesteps. Tom’s Hardware notes that TPUs now offer “real performance parity” with H200 for inference-specific workloads, particularly for models pre-optimized in JAX or PyTorch-XLA.

The caveat is ecosystem lock-in: CUDA’s software moat remains formidable. Google’s counter-strategy is external adoption—pitching direct sales and deployment inside customer data centers rather than GCP-only infrastructure. The $20 billion Groq acquisition signals Nvidia’s taking alternative architectures seriously. Meanwhile, hardware breakthroughs in chip architecture from academia keep pushing efficiency frontiers that both companies must chase.

Meta Partnership: The External Customer Proof Point

Meta’s reported TPU discussions with Google represent the inflection point for external adoption. According to The Information’s reporting, Meta is negotiating to rent Google Cloud TPU capacity in 2026, then transition to direct purchases for internal data center deployment in 2027—a multi-billion-dollar deal that would make Meta the first major hyperscaler to diversify away from Nvidia GPUs at scale.

Production capacity comparison: Google TPU 2026 expansion vs Nvidia GPU market share

The technical collaboration runs deeper than hardware sales. Google is working directly with Meta to optimize PyTorch-on-TPU performance, addressing Meta’s strategic goal of reducing inference costs without forcing migration to JAX. That’s crucial—most AI companies are locked into PyTorch ecosystems, and Google’s willingness to invest engineering effort into PyTorch compatibility signals genuine commitment to external adoption rather than ecosystem colonization.

Meta’s motivations are economic and strategic. Nvidia GPU shortages persist despite capacity expansion, and inference costs at Meta’s scale make even 20-30% efficiency gains worth infrastructure complexity. The 2027 timeline gives Google two years to prove production stability before Meta commits serious capital. If the Meta partnership succeeds, every other hyperscaler will evaluate TPU alternatives—which is exactly the ecosystem flywheel Google needs to justify 3.2 million unit production runs.

Market Impact: Revenue Projections and Ecosystem Growth

Morgan Stanley’s math is straightforward: Google TPU expansion could capture 10% of Nvidia’s annual revenue, translating to $13 billion in new Google revenue. That assumes successful external sales beyond Google Cloud’s internal consumption—a big assumption given TPUs have never meaningfully sold outside Google’s ecosystem before.

Anthropic’s million-chip commitment signals enterprise confidence. Anthropic will deploy over one million Ironwood chips beginning in 2026 for Claude model training and inference—the largest external TPU adoption announced to date. That’s validation from an independent AI company with options, not just Google Cloud’s internal consumption.

Ecosystem challenges remain. JAX and PyTorch-XLA are not TensorFlow replacements for most AI teams, and CUDA’s software moat extends far beyond the framework level. Google’s investing heavily in PyTorch compatibility for Meta, but every additional customer requires similar engineering effort—effort that scales linearly rather than amortizing like Nvidia’s ecosystem investments.

The Real Test: Software Ecosystem Maturity

Ironwood’s architectural decisions reveal Google’s inference-first strategy. The SparseCores—four per chip—specifically target sparse matrix operations where most LLM inference computation occurs. Standard GPUs process zeros and non-zeros equally, burning watts on mathematically unnecessary operations. Sparse architectures skip the zeros, delivering better performance-per-watt for models that leverage sparsity.

The memory hierarchy matters more than raw compute at scale. Ironwood’s 1.77 PB of pod-level shared HBM enables fitting trillion-parameter models entirely in fast memory, eliminating DRAM streaming bottlenecks that cripple GPU inference. The ICI network’s 9.6 Tb/s links create genuinely shared memory space rather than NCCL-over-networking approximations.

But hardware wins battles; software wins wars. Google’s PyTorch investment for Meta demonstrates understanding of this reality, but every additional customer requires similar framework optimization. Nvidia’s decade-long CUDA investment created moats Google can’t overcome in 18 months—the question is whether “good enough” PyTorch compatibility suffices for inference workloads where training stays on Nvidia.

Forward Outlook: 2026-2027 Trajectory

The next 18 months determine whether Ironwood is a Nvidia competitor or a Google Cloud footnote. Meta’s 2026 rental and 2027 purchase timeline creates clear milestones: if Meta actually deploys TPUs at scale, every hyperscaler will evaluate alternatives. If Meta quietly returns to Nvidia exclusivity, the external adoption thesis collapses regardless of production capacity.

TPU v8 is already in development for 2026 deployment, with Samsung’s HBM4 supply chain positioning and MediaTek’s v8e design work publicly reported. The roadmap suggests Google is committed to annual cadence matching Nvidia’s—no longer lagging but keeping architectural pace. TSMC’s 3nm capacity increases, Samsung’s HBM production ramps, and MediaTek’s 7-fold CoWoS request all indicate 2026-2027 production will exceed 3.2 million units if demand materializes.

Watch the Anthropic deployment. One million Ironwood chips in production, running Claude at scale, represents the highest-stakes public validation of TPU architecture for external customers. If Anthropic succeeds and publicly discusses cost/performance advantages, enterprise adoption accelerates. If they struggle or stay quiet about results, it signals challenges Google hasn’t publicly acknowledged. Either way, we’ll know by mid-2026.

The headline “Google Challenges Nvidia” is partially true but incomplete. More accurately: “Google builds infrastructure for its own AI ambitions and discovers it can sell the surplus.” The era of Nvidia’s unchallengeable dominance is ending—not because Nvidia loses market share, but because its share of new capacity growth will compress from near-100% to 75-80%. Companies will increasingly optimize for the cheapest, fastest path to their specific goal. For Google and Meta, that’s Ironwood. For PyTorch teams, that’s still Nvidia. For inference, both will compete with Groq and others. The market grows, the landscape fragments, and the winners are the teams that can optimize for hardware diversity.

Get the Daily Pulse

Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.

Get the Daily Pulse

Sharp AI analysis, daily. Two minutes, every morning.

Get the Daily PulseTwo minutes, every morning