Gemini 3.1 Pro Tops AI Rankings at Half the Price of Claude

Gemini 3.1 Pro scored 77.1% on ARC-AGI-2 โ€” more than double Gemini 3 Pro’s 31.1% โ€” and Google didn’t raise the price by a cent. The model launched in preview on February 19, 2026, and within hours, Artificial Analysis ranked it first on their Intelligence Index, four points ahead of Claude Opus 4.6 at less than half the cost. The benchmark win is real. But the pricing decision โ€” holding at $2/$12 per million tokens while jumping to first place โ€” is the more interesting story.

Google is running the Android playbook on AI: give away the best model at cost to lock in the platform. The question for developers and CTOs isn’t just whether Gemini 3.1 Pro is the best model. It’s whether that strategy means the model pricing war is already over.

The ARC-AGI-2 Doubling โ€” What Gemini 3.1 Pro’s Number Actually Means

A 148% improvement on the hardest reasoning benchmark in AI doesn’t happen through incremental tuning. On the ARC-AGI-2 leaderboard, Gemini 3.1 Pro now sits at 77.1%, surpassing Claude Opus 4.6 (68.8%) by 8.3 points and GPT-5.2 (52.9%) by 24.2 points. That’s not a margin โ€” it’s a gap. The original ARC-AGI-1 benchmark is now essentially solved across frontier models, confirming that ARC-AGI-2 is the measuring stick that matters.

Context helps calibrate the achievement. Gemini 3 Deep Think’s 84.6% ARC-AGI-2 score โ€” set on February 12, 2026 โ€” still holds the overall record, but Deep Think is a specialized reasoning model with high latency. Gemini 3.1 Pro delivers 91% of Deep Think’s reasoning at standard production pricing. The meaningful comparison isn’t whether it beat the specialized model, but how close it got while remaining deployable.

The Artificial Analysis Intelligence Index tells a broader story. Gemini 3.1 Pro scored 57 points to Claude Opus 4.6’s 53, leading 6 of 10 evaluations including GPQA Diamond (94.3%), Humanity’s Last Exam without tools (44.4%), and CritPt (research-level physics). It also hit 80.6% on SWE-Bench Verified, neck-and-neck with Claude Opus 4.6’s 80.8%. As Demis Hassabis wrote on X: “Major improvements across the board including in core reasoning and problem solving.”

Three Thinking Levels: Why This Is Really a ‘Deep Think Mini’

The mechanism behind these gains is an adjustable reasoning system. Gemini 3.1 Pro introduces three thinking levels โ€” Low (speed-optimized), Medium (balanced), and High (inspired by Deep Think’s framework) โ€” where its predecessor only offered Low and High. That Medium tier is the sweet spot developers have been asking for: enough reasoning depth to handle complex tasks, without the latency and cost penalty of full Deep Think-style computation.

VentureBeat described it as a “Deep Think Mini” after testing showed the High setting solved an International Math Olympiad problem in 8 minutes versus 17+ minutes for full Deep Think. The practical implication: developers can now use a single model endpoint and adjust reasoning depth per request, rather than routing different tasks to different specialized models.

Early practitioner feedback confirms the architecture change is noticeable. Tech journalist Max Weinbach observed on X that the model was “reasoning 2-3x more than previous models” and delivering “deeper and more nuanced answers.” Developer @Lentils80 noted a tradeoff: 3.1 Pro produced 700 lines of code where 3 Pro produced 1,000 โ€” fewer lines, but better outputs. The “lazy but efficient” characterization captures the real tradeoff developers will navigate.

Illustration: Gemini 3.1 Pro pricing and benchmark performance

Gemini 3.1 Pro Pricing: The Android Playbook Applied to AI

The pricing tells the strategic story. Google held API pricing flat at $2 per million input tokens and $12 per million output tokens (up to 200K context) despite jumping to first on the Intelligence Index. Context caching drops input costs further โ€” to $0.20 per million for standard prompts. Compare that to Claude Opus 4.6 at $5/$25 and GPT-5.2 at $1.75/$14. On input tokens, Gemini is 2.5x cheaper than Claude โ€” though GPT-5.2 is priced comparably on input at $1.75 per million.

At enterprise scale, these ratios compound fast. A workload consuming 10 billion output tokens monthly costs $120K on Gemini versus $250K on Claude Opus โ€” a $1.56 million annual gap on output alone. Our AI API pricing comparison already showed Google leading on cost. Now it’s leading on performance too.

This isn’t accidental. Google doesn’t need to monetize the model itself โ€” it monetizes Vertex AI enterprise contracts, AI Studio developer adoption, the Apple/Siri partnership (announced January 12, 2026), and the downstream ad ecosystem. Commoditizing model intelligence makes Gemini the default infrastructure layer, exactly as Android commoditized mobile OS to monetize search and ads. Google’s announcement listed seven integration points from Gemini CLI to GitHub Copilot โ€” breadth of distribution that reinforces the platform strategy.

Where Claude and OpenAI Still Win โ€” and Why That Gap Matters

Here’s where the “Google wins everything” narrative breaks down. On GDPval-AA โ€” the evaluation where human experts judge response quality on complex real-world tasks โ€” Claude Opus 4.6 scores 1,606 Elo to Gemini’s 1,317. That 289-point gap is enormous by Elo standards. When the judges are domain experts evaluating nuanced, multi-step reasoning, they still prefer Claude by a wide margin. The automated benchmarks and the human evaluations are telling different stories.

ML author Andriy Burkov offered a sharper critique on X: “They have finetuned 3.1 Pro specifically to improve on this single benchmark.” His argument โ€” that the ARC-AGI-2 doubling reflects benchmark-specific optimization rather than general reasoning gains โ€” finds support in the GDPval-AA gap. If the reasoning improvements were truly general, you’d expect the human expert evaluations to move in the same direction. They didn’t.

OpenAI occupies an awkward middle position. GPT-5.2 is priced comparably to Gemini on input ($1.75 vs $2 per million) but costs nearly 17% more on output ($14 vs $12 per million) โ€” and doesn’t lead either automated benchmarks (52.9% ARC-AGI-2) or human-evaluated quality (behind Claude on GDPval-AA). GPT-5.3-Codex holds a narrow coding wedge at 75.1% on Terminal-Bench 2.0, ahead of Gemini 3.1 Pro’s 68.5%, but a specialized coding advantage doesn’t rescue the mid-tier positioning.

Gemini 3.1 Pro did show improvement on AA-Omniscience โ€” Artificial Analysis flagged the hallucination and knowledge evaluation as one of its six index-leading categories. That’s real progress on a dimension that matters for production deployments.

Then there’s the preview caveat. No general availability date has been announced, which means no production SLAs on Vertex AI. Simon Willison reported on launch day that a simple “hi” prompt took 104 seconds to process, alongside “high demand” error messages. These are likely teething problems, but they illustrate the gap between benchmark performance and production readiness. Enterprise CTOs writing procurement contracts today cannot deploy this model at production scale โ€” not yet.

The frontier model market now has three clear tiers: Gemini 3.1 Pro ($2/$12) for best price-performance, GPT-5.2 ($1.75/$14) as a coding specialist, and Claude Opus 4.6 ($5/$25) for best human-evaluated expert quality. Each has a defensible position โ€” for now.

The Real Question Isn’t Who Leads the Benchmarks

If Google’s platform strategy succeeds โ€” if Gemini becomes the default infrastructure layer the way Android became the default mobile OS โ€” does it matter that Claude still wins when expert humans judge quality? The GDPval-AA gap suggests there are tasks where paying 2.5x more on input โ€” and 2x more on output โ€” is justified. The question is whether that premium market is large enough to sustain Anthropic’s current pricing, or whether Google’s distribution advantage eventually collapses it.

Google didn’t just release a better model โ€” it released a better model at the same price. That combination is far more dangerous to competitors than any benchmark score alone. Android didn’t win the smartphone war by being the best OS. It won by being good enough and free, while the premium option consolidated into a profitable but shrinking share. If Gemini follows that trajectory, the AI battle won’t be decided by ARC-AGI-2 scores. It’ll be decided by procurement departments.

Gemini 3.1 Pro’s general availability date โ€” whenever Google announces it โ€” will be the moment enterprises can actually sign those contracts. Until then, every “Google retakes the AI crown” headline is technically describing a preview that no enterprise CTO can put in production.

Get the Daily Pulse

Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.

Get the Daily Pulse

Sharp AI analysis, daily. Two minutes, every morning.

Get the Daily PulseTwo minutes, every morning