The inference era just got its price tag. Jensen Huang walked into GTC 2026 and announced $1 trillion in projected orders and revenue through 2027—double the $500 billion visibility he cited in late 2025. We ran the numbers on the Vera Rubin platform: 336 billion transistors per GPU (1.6x Blackwell’s 208 billion), 288GB of HBM4 running at 22 TB/s bandwidth (2.8x over Blackwell’s 8 TB/s), and 50 petaflops of FP4 inference per chip—a 5x jump. Each NVL72 rack packs 72 GPUs and 36 CPUs at an estimated $3.5–4 million per unit. But the real architecture shift is disaggregation: Rubin handles prefill, Groq’s LP30 chip—a $20 billion asset deal closed in December 2025—handles decode with 500MB of on-chip SRAM and 150 TB/s bandwidth, and the new open-source Dynamo 1.0 orchestrator ties them together. NVIDIA claims 35x inference throughput per megawatt. The supply chain is already locked in: Micron began mass production of 36GB 12-high HBM4 in Q1 2026 with 2.8 TB/s bandwidth and a 20%-plus power efficiency gain, Samsung confirmed as Groq 3 LPU foundry shipping H2 2026, and Intel Xeon 6 is the host CPU for DGX Rubin. When one company claims 30–35% of every dollar spent building data centers worldwide, the trillion-dollar number stops sounding like a forecast and starts sounding like a floor.
The same week AI infrastructure gets a trillion-dollar commitment, the U.S. government is labeling an AI company a supply chain risk. 30-plus employees from OpenAI and Google DeepMind—including Google chief scientist Jeff Dean—filed an amicus brief backing Anthropic after the Pentagon applied the “supply chain risk” designation to a U.S. company for the first time. Anthropic declined two military conditions: mass surveillance integration and autonomous weapons deployment. The preliminary injunction hearing is set for March 24 in San Francisco federal court. AWS committed over 1 million NVIDIA GPUs in 2026 and Azure will be first to power on Vera Rubin NVL72 racks—yet the government building those very data center contracts is simultaneously blacklisting the company that refused to remove its safety guardrails. The gap between infrastructure spending and governance clarity is not just widening—the two are moving in opposite directions.
And while Jensen pitches 50-petaflop inference chips, the models running on today’s hardware are already generating real harm. Three Tennessee minors filed suit against xAI on March 16 after Grok generated sexually explicit images from real school photos—homecoming dances, yearbook pages. One arrest followed in December 2025, with images of multiple girls from the same school. xAI had promoted Grok’s “spicy mode” before a partial rollback in January 2026. Dynamo 1.0 boosted token generation from 700 to 5,000 tokens per second on existing Blackwell hardware—a 7x jump—and NVIDIA stock rose 1.65% on March 16 after touching 4.8% intraday. Faster, cheaper inference is the thesis behind every deal announced at GTC this week. The question no one on stage answered is what happens when models that already lack adequate guardrails get 35x more throughput per megawatt.
Latest from PulseMark
![]() |
NVIDIA GTC 2026: Inside the $1 Trillion Vera Rubin Bet
The $20B Groq acquisition, GPU-LPU disaggregation architecture, why Dynamo 1.0 is open-source, and the per-rack economics that make $3.5–4M look like a bargain at 35x throughput per megawatt. |
That’s Tuesday. Five stories, zero fluff.
— The PulseMark Team
Get the Daily Pulse
Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.

