The code breaks itself

Alibaba's AI coding models regress 75% of the time, while Meta plans to cut 16,000 jobs to fund $135 billion in AI investments.

The code breaks itself. Alibaba tested 18 AI coding models across 100 real Python codebases—each evolved over 233 days of maintenance cycles—and found that 75% of models introduce regressions in more than three out of four maintenance iterations. We ran the numbers on the SWE-CI benchmark: Claude Opus 4.6 led with a 0.76 zero-regression rate, meaning it still broke working code roughly one in four times, while GPT-5.2 scored 0.23—breaking code 77% of the time. The test consumed 10 billion tokens across all models. That matters because the AI coding tools market hit $7–8 billion in 2025 and maintenance accounts for 60–80% of total software lifecycle costs. Replit just tripled its valuation from $3 billion to $9 billion in six months on the strength of “vibe coding”—accepting all AI diffs without reading them. The gap between what these tools cost and what they reliably deliver is widening, not closing.

The companies spending the most on AI infrastructure are also cutting the most people to pay for it. Meta is reportedly planning to eliminate 20% of its 79,000-person workforce—approximately 16,000 jobs—to offset $115–135 billion in 2026 AI capex, nearly doubling last year’s $72.2 billion. We dug into the 30x gap between executives who cut jobs anticipating AI efficiencies (60%) and those who cut because AI actually replaced the work (2%). Meta isn’t alone: over 40,000 tech jobs have been cut in Q1 2026, including 4,000 at Block (40% of its workforce), 1,600 at Atlassian (10%), and roughly 16,000 at Amazon. Block’s stock surged over 20% after its layoff announcement. The playbook is now standardized—cut headcount, cite AI, watch the stock price respond.

The hardware that makes those AI bets credible is getting dramatically cheaper. Jensen Huang takes the stage at GTC 2026 this afternoon in front of 30,000 attendees from 190 countries to showcase the Vera Rubin platform—first announced at CES in January and now entering its rollout phase—promising a 10x drop in inference token cost over Blackwell and 4x fewer GPUs needed for mixture-of-experts training. Each VR200 NVL72 rack packs 72 GPUs and 36 CPUs with HBM4 memory, and NVIDIA is now measuring infrastructure scale in gigawatts. That 10x inference cost drop is exactly the tailwind behind Replit’s $400 million Series D and the $20 billion valuation gap between Cursor ($29.3 billion) and Replit ($9 billion)—we broke down how Georgian Partners invested in the $3 billion Series C and then led the $9 billion Series D, and what Agent 4’s code-aware reasoning changes. Cursor doubled its ARR to $2 billion in three months with $878 million in total funding now flowing into Replit alone. Cheaper inference makes every AI coding bet look rational on a spreadsheet—even when the benchmarks say the code breaks itself 75% of the time.

Latest from PulseMark



SWE-CI: AI Coding Agents Break Working Code 75% of the Time

Alibaba’s 10-billion-token test across 18 models, why maintenance is the real AI coding frontier, and the per-model regression scores no one is talking about.

Read the analysis →



Meta’s 20% Layoffs Are AI-Washing at Scale — The Math Proves It

The $690 billion Big Tech AI infrastructure sprint, the 60%-vs-2% executive survey gap, and why Block’s 20%+ stock surge after cutting 4,000 people created the template.

Read the analysis →



Replit Raises $400M at $9B Valuation and Launches Agent 4

Georgian Partners doubled down twice, Cursor’s $2B ARR sets the ceiling, and what the $20 billion valuation gap says about where vibe coding goes next.

Read the analysis →

That’s Monday. Seven stories, zero fluff.

— The PulseMark Team

Get the Daily Pulse

Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.

Get the Daily Pulse

Sharp AI analysis, daily. Two minutes, every morning.

Get the Daily PulseTwo minutes, every morning