AI Transparency

Gemini 3 Deep Think Hits 84.6% on ARC-AGI-2 — But Only 6.5% of Real Research Problems Got Useful Answers

Gemini 3 Deep Think Hits 84.6% on ARC-AGI-2 — But Only 6.5% of Real Research Problems Got Useful Answers - Featured Image

Google’s Gemini 3 Deep Think scored 84.6% on ARC-AGI-2 on February 12, 2026 — a 15.8-point lead over Claude Opus 4.6 and a 31.7-point demolition of GPT-5.2 on the hardest reasoning benchmark in AI. Then Google did something unusual: it…

Read MoreGemini 3 Deep Think Hits 84.6% on ARC-AGI-2 — But Only 6.5% of Real Research Problems Got Useful Answers

Get the Daily Pulse

Sharp AI analysis, daily. Two minutes, every morning.

Get the Daily PulseTwo minutes, every morning