At NVIDIA GTC 2026 on March 16, Jensen Huang told a packed San Jose audience that NVIDIA expects “at least $1 trillion” in orders and revenue from Blackwell and Vera Rubin through 2027 โ double the $500 billion figure he cited a year ago. The Vera Rubin platform, seven chips and five rack configurations deep, is the vehicle for that projection.
But the headline number obscures the real story. NVIDIA isn’t just shipping faster GPUs. It’s dismantling the boundary between GPU and inference accelerator, open-sourcing the software that orchestrates both, and betting $20 billion that this architectural split will lock in the next decade of AI infrastructure.
The $1 trillion number in context
That $1 trillion figure, reported by CNBC, is a demand projection โ Huang framed it as anticipated orders and revenue from Blackwell and Vera Rubin systems combined through 2027, with different outlets characterizing it as orders (CNBC, AP) or revenue (Axios). For context, global data center capital expenditure is projected at roughly $500 billion by 2027. NVIDIA is claiming 30-35% of every dollar spent building data centers worldwide.
At Intel’s peak dominance in 2000, server CPUs captured roughly 8-10% of total data center spend. NVIDIA is projecting three times that share โ from a position that didn’t exist five years ago. The cloud partners appear to agree: AWS committed to deploying more than one million NVIDIA GPUs across its regions beginning in 2026, and Microsoft Azure became the first hyperscaler to power on a Vera Rubin NVL72 system.
NVIDIA stock rose 4.8% intraday on March 16 before settling to close up 1.65%. Goldman Sachs maintained its Buy rating, noting that Huang’s remarks addressed two core investor concerns about AI growth sustainability. The stock reaction was measured โ investors have heard big numbers before. The mechanism delivering those numbers is what matters.
NVIDIA GTC 2026 Vera Rubin GPU: what a 5x inference leap looks like
The Rubin GPU (R200/VR200) packs 336 billion transistors onto a TSMC 3nm multi-chip module โ 1.6x Blackwell’s 208 billion. Each GPU carries 288 GB of HBM4 memory with 22 TB/s bandwidth, a 2.8x jump over Blackwell’s 8 TB/s. The per-GPU inference performance: 50 petaflops in NVFP4, a clean 5x over its predecessor.
Scale that to a rack and the numbers get absurd. The NVL72 configuration โ 72 Rubin GPUs and 36 Vera CPUs in a single liquid-cooled enclosure โ delivers 3.6 exaflops of FP4 inference with 260 TB/s of NVLink 6 bandwidth. Estimated rack price: $3.5-4.0 million, roughly a 25% premium over Blackwell’s ~$3.35 million. NVIDIA claims 10x inference throughput per watt and one-tenth cost per token versus the outgoing generation.
The power envelope tells its own story: 1,800-2,300W per GPU, requiring 100% liquid cooling. Every customer buying Rubin is also buying a plumbing upgrade. Azure was the first hyperscaler to power on and validate a Vera Rubin NVL72 system; broad availability from AWS, Google Cloud, Azure, and system manufacturers arrives in H2 2026. But these specs โ impressive as they are โ only deliver their full potential paired with NVIDIA’s other new chip. That’s where the inference hardware diversification story gets interesting.

Groq 3 LPU: the $20 billion bet on disaggregated inference
The Groq 3 LP30 is the first product from NVIDIA’s $20 billion Groq asset purchase and licensing deal, which closed in December 2025. Nine months from deal close to shipping product โ a pace that should make Intel wince. (Intel’s $16.7 billion Altera acquisition in 2015 took nine years of fumbled integration before Intel spun Altera back out in 2024.)
The LP30 is a deterministic dataflow processor with 500 MB of on-chip SRAM per chip delivering 150 TB/s bandwidth โ 7x faster than Rubin’s HBM4. Pack 256 of them into an LPX rack and you get 315 petaflops of FP8 compute with 40 PB/s aggregate bandwidth. Samsung manufactures the LP30 exclusively, with LPX racks shipping Q3 2026.
Here’s the architectural insight that makes this more than a spec upgrade: Rubin GPUs handle prefill and attention โ the computationally dense work of processing input tokens. Groq LPUs handle decode โ the latency-sensitive generation of output tokens, especially feed-forward and mixture-of-experts execution. Neither chip replaces the other. Together, orchestrated by software, they’re more efficient than either alone. This mirrors the CPU-GPU disaggregation of the early 2000s that created NVIDIA’s entire business. Now NVIDIA is splitting inference the same way it once split rendering.
NVIDIA claims the combined GPU+LPU system delivers 35x inference throughput per megawatt versus the prior generation. That number awaits independent validation โ and it’s the single most important benchmark to watch when LPX racks reach customer data centers. As AMD’s $60 billion chip deal with Meta showed, hyperscalers aren’t short on alternatives. NVIDIA’s disaggregated pitch needs to clear a high bar.
Dynamo 1.0 and the open-source lock-in
The hardware alone doesn’t explain NVIDIA’s strategy. The software does. Dynamo 1.0 launched on March 16 as open-source inference orchestration software โ the layer that splits inference traffic across GPUs and LPUs, manages memory, and routes prefill to Rubin while sending decode to Groq. On existing Blackwell systems, Dynamo boosts token generation from 700 to 5,000 tokens per second โ a 7x improvement, available now, for free.
Then there’s NemoClaw, an open-source enterprise agent platform built on OpenClaw with built-in security and privacy guardrails. Jensen Huang compared it to HTTP, Linux, and Kubernetes, declaring that “every company needs an OpenClaw strategy.” NemoClaw is hardware-agnostic โ it runs on non-NVIDIA chips. That’s not charity. That’s the strategy.
Huang’s own framing was disarmingly honest: “vertically integrated but horizontally open.” NVIDIA builds the most tightly coupled hardware stack in AI โ seven chips, five racks, one orchestration layer โ while open-sourcing the software that sits on top. Developers adopt Dynamo and NemoClaw because they’re free, useful, and run anywhere. Once workflows depend on that software, switching costs compound. The hardware-agnostic door lets everyone in; the Rubin+Groq performance advantage keeps them there. Dynamo is available on GitHub as of March 16. NemoClaw is early-stage alpha โ adoption timelines differ meaningfully.
What this changes for the competitive field
Every inference-focused startup โ Cerebras, SambaNova, pre-acquisition Groq โ pitched the same thesis: GPUs aren’t optimized for inference workloads, so customers need specialized silicon. NVIDIA absorbed that argument wholesale by purchasing the most prominent inference chip company’s assets and technology and integrating it at the rack level. The competitive lane that startups were racing in got annexed.
The implications extend to NVIDIA’s largest competitors too. AMD’s custom silicon pitch to Meta loses urgency when NVIDIA can offer GPU+LPU in a single orchestrated rack. The roadmap reinforces the pressure: Rubin Ultra ships in 2027 with an NVL576 configuration, and the Feynman generation follows in 2028 on TSMC 1.6nm. NVIDIA’s one-year cadence is locked in โ each generation layers on the software moat.
What to watch next
The $1 trillion demand projection and the 35x throughput-per-megawatt claim both rest on a single untested assumption: that Rubin GPUs and Groq LPUs work as advertised together, at scale, in production environments. No one outside NVIDIA has validated that yet. Groq 3 LPX racks ship in Q3 2026, Vera Rubin NVL72 goes broadly available from cloud partners in H2 2026 โ and those deployments will be the first independent scorecards.
Jensen Huang didn’t announce better hardware on March 16 โ he announced that the argument for buying anyone else’s inference chips is now NVIDIA’s argument to lose. The open-source generosity, the hardware-agnostic agent platform, the physical AI ambitions โ they all point to a company competing on ecosystem gravity, not specs. If CoreWeave or Together AI publish real-world throughput numbers approaching that 35x claim, NVIDIA’s trillion-dollar projection starts looking conservative. If they don’t, every competitor gets a second wind.
Get the Daily Pulse
Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.



