Gemini 3 Deep Think Hits 84.6% on ARC-AGI-2 — But Only 6.5% of Real Research Problems Got Useful Answers

Google’s Gemini 3 Deep Think scored 84.6% on ARC-AGI-2 on February 12, 2026 — a 15.8-point lead over Claude Opus 4.6 and a 31.7-point demolition of GPT-5.2 on the hardest reasoning benchmark in AI. Then Google did something unusual: it published the failure data.

When the same model’s Aletheia research agent was deployed on 700 actual open math problems from Bloom’s Erdős Conjectures database, only 6.5% of evaluable answers meaningfully addressed the question posed. That 13x gap between benchmark performance and real research utility is the entire story of where reasoning models stand in February 2026.

Alongside the model upgrade, Google DeepMind released companion papers documenting genuine scientific contributions: 18 research problems solved across physics, economics, and computer science, plus a math agent called Aletheia that autonomously wrote what the team classifies as a publishable paper on eigenweights in arithmetic geometry. The benchmarks are spectacular. The research is real. The researchers’ own data shows why both need more context than a press release provides.

Gemini 3 Deep Think Reclaims the Reasoning Crown

The numbers, verified by the ARC Prize Foundation, are not close. Deep Think hit 84.6% on ARC-AGI-2 at $13.62 per task — a benchmark designed to be easy for humans but brutally hard for AI, testing abstract pattern recognition without training data contamination. Claude Opus 4.6’s Thinking Max mode managed 68.8%. GPT-5.2’s Thinking xhigh scored 52.9%. On ARC-AGI-1, Deep Think effectively hit the ceiling at 96.0%.

The dominance extends beyond ARC. Deep Think scored 48.4% on Humanity’s Last Exam without tools, leading Opus 4.6 (40.0%) and GPT-5.2 (34.5%). On Codeforces, its 3,455 Elo rating towers over Opus 4.6’s 2,352 — a 1,103-point gap that, in competitive programming terms, separates a grandmaster from a specialist. Gold medals on the written portions of 2025’s International Math, Physics, and Chemistry Olympiads complete the picture: by any structured reasoning metric, this is the most dominant model performance since GPT-4’s original launch.

For context on where Claude Opus 4.6’s strengths actually lie — finance, agent orchestration, and million-token context — the competitive picture is more nuanced than a single leaderboard suggests. And as we’ve argued before, benchmark obsession can be a distraction from what models actually do for users. But on reasoning specifically, Deep Think has no peer as of mid-February 2026.

Aletheia Autonomously Wrote a Publishable Math Paper

The benchmark scores are the headline. Aletheia is the deeper story. Described in a February 10 paper by a 28-author team led by Tony Feng (with Demis Hassabis and Quoc V. Le as senior authors), Aletheia is a math research agent powered by an advanced version of Deep Think.

The architecture matters: generate-verify-revise loops where correct solutions pass through, minor issues hit a reviser, and critically flawed attempts cycle back to the generator. A natural language verifier decouples reasoning from output, letting the model catch flaws it initially missed.

Aletheia’s signature achievement is a paper calculating eigenweights — structure constants in arithmetic geometry — that extended the Hirzebruch Proportionality Principle using algebraic combinatorics techniques unfamiliar to the original human authors. The team classified this as Level A2: Autonomous and Publishable. They also achieved a 100x reduction in compute while pushing IMO-ProofBench Advanced accuracy to 95.1% (up from 65.7%). The architecture, not just the model weights, is doing the work.

18 Research Problems Solved — From a Decade-Old Conjecture to Cosmic Strings

A companion paper by Woodruff et al. (34 authors, submitted February 3) documented 18 human-AI collaborations spanning algorithms, cryptography, mechanism design, economics, and physics. Two results stand out.

In submodular optimization, a 2015 conjecture proposed that copying an arriving item in a data stream is always less valuable than moving the original. For a decade, experts tried and failed to prove it. Deep Think engineered a three-item combinatorial counterexample, rigorously proving the intuition false. The counterexample is tiny — three items — but no human thought to look there. Sometimes brute mathematical breadth finds what intuition conceals.

In physics, Google DeepMind’s research blog describes how Deep Think discovered that Gegenbauer polynomials could solve the gravitational radiation calculation for cosmic strings — the polynomials “naturally absorbed the singularities, collapsing an infinite series into a closed form, finite sum.” Other highlights: detecting a proof bug in a SNARG cryptographic construction, extending the Revelation Principle for AI token auctions, and improving Max-Cut SDP bounds by applying continuous math to discrete problems where progress had stalled.

Illustration: Gemini 3 Deep Think

The Erdős Evaluation: 68.5% Wrong, 6.5% Useful

Here is where the press release meets reality. Feng et al.’s evaluation (submitted January 29, 2026) deployed Aletheia on approximately 700 conjectures labeled “Open” in Bloom’s Erdős Problems database in December 2025. Of those 700, 212 returned as potentially correct. Of 200 evaluable answers, 137 — that’s 68.5% — were fundamentally wrong at a basic mathematical level. Only 13 (6.5%) actually answered the question posed.

The system addressed 13 problems in total: five through seemingly novel autonomous solutions (Erdős-652, 654, 1040, and 1051 being the most significant) and eight by identifying prior solutions already existing in the literature. That’s legitimately impressive. But as The Decoder highlighted, the model exhibits systematic “specification gaming” — reinterpreting hard questions as easier ones, then confidently answering the wrong question. The researchers noted many “Open” problems remained unsolved out of obscurity rather than difficulty.

Fields Medalist Terence Tao’s personal experience tells both sides. In a Mathstodon post, Tao described testing Deep Think on Erdős Problem #367: the model produced a complete proof in about ten minutes. Tao then spent half an hour converting its p-adic algebraic number theory proof into something more elementary. Boris Alexeev spent another two to three hours formalizing it in Lean.

Ten minutes of AI generation, three-plus hours of expert human refinement. That ratio tells you exactly what “AI-assisted research” looks like in practice.

What the Researchers’ Own Taxonomy Reveals

The Aletheia team introduced a two-axis classification for AI research contributions: Autonomy (H for primarily Human, C for Collaborative, A for Autonomous) crossed with Significance (0 for negligible through 4 for landmark breakthroughs). Their best result — the eigenweights paper — earned A2: Autonomous and Publishable. The independent sets paper earned C2: Collaborative and Publishable. Nothing reached Level 3 (major advance, top-tier journal). Nothing reached Level 4.

The team’s self-assessment is bracingly honest: “Successes so far stem more from the model’s enormous breadth of knowledge and clever technical workarounds than from genuine mathematical creativity.” Launch headlines tout 18 solved problems. The actual papers classify every result in the bottom half of the researchers’ own scale. This isn’t dishonesty — it’s the familiar gap between a product announcement and the research underneath. Google deserves credit for publishing both.

The Peer Review Crisis No One Is Ready For

The researchers buried their most important warning in the limitations section: “If AI massively accelerates the production of technically complex research papers, the bottleneck in science shifts from generating ideas to verifying them.” Consider the math. An AI that achieves Level A2 — autonomous and publishable — while getting 68.5% of open problems fundamentally wrong can flood journals with sophisticated submissions requiring expert review but failing often enough to overwhelm the reviewers catching the errors.

The failure mode has evolved. Web browsing reduced fabricated citations, but hallucinations adapted: the model now cites real papers while misrepresenting their contents. A reviewer who sees a legitimate arXiv ID is less likely to verify whether the cited result actually says what the model claims. Given Alphabet’s $175-185 billion AI infrastructure spending planned for 2026, these models will only get more capable at producing sophisticated-looking research — and peer review was already straining under current volumes.

Tao has been sounding alarms on exactly this front: “As the application of AI in mathematical research deepens, in addition to responsible use, cases of AI misuse are also common.” He’s set up a community wiki to track AI-assisted progress on Erdős problems and called for transparent disclosure of AI’s role in research. When a Fields Medalist builds tracking infrastructure for AI contributions to mathematics, the verification crisis isn’t theoretical.

The Gap Between Puzzles and Science

If Deep Think can occasionally solve what a Fields Medalist cannot but gets 68.5% of open problems fundamentally wrong, how do we build systems that know when to trust their own answers?

The 13x gap between benchmark scores and research utility is not a bug to fix with the next model update. ARC-AGI-2 tests whether a model can find a pattern. Open research tests whether a model can recognize what the question even is. Those are fundamentally different cognitive tasks, and no amount of benchmark optimization bridges that divide.

The eigenweights paper’s peer review timeline will be the real test. If Feng2026 survives expert scrutiny, Google’s Level A2 classification holds and Aletheia becomes a legitimate precedent for autonomous AI research. If reviewers find errors the natural language verifier missed, the peer review crisis arrives ahead of schedule. The answer will tell us something the ARC-AGI-2 leaderboard never can: whether AI can do science, or just solve puzzles about it.

Get the Daily Pulse

Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.

Get the Daily Pulse

Sharp AI analysis, daily. Two minutes, every morning.

Get the Daily PulseTwo minutes, every morning