Benchmarks lie, safety labs arm, PS6 waits

Google's Gemini 3 Deep Think leads benchmarks but fails on real math problems. Anthropic's Claude is now used by the Pentagon.

The benchmarks are lying to you. Google’s Gemini 3 Deep Think just hit 84.6% on ARC-AGI-2—a 15.8-point lead over Claude Opus 4.6 and a 31.7-point demolition of GPT-5.2 on the hardest reasoning benchmark in AI. Gold-medal-level scores on math, physics, and chemistry Olympiad problems. A 3,455 Elo on Codeforces. By every structured test, it’s the most dominant model since GPT-4’s original launch. Then Google did something unusual: it published the failure data. We dug into what happens when that same model hits real research problems—and 68.5% of its answers on 700 open math conjectures were fundamentally wrong. Only 6.5% meaningfully addressed the question. That’s a 13x gap between benchmark performance and research utility.

The gap between capability and reality showed up in a different way this week. Anthropic’s Claude is now deployed inside the Pentagon for military planning, intelligence analysis, and operational logistics. We broke down how the company that built AI’s most sophisticated safety infrastructure ended up enabling exactly the deployment it was designed to prevent. Meanwhile, the agent race accelerated: OpenAI hired Peter Steinberger, the creator of OpenClaw—an open-source AI assistant with nearly 200,000 GitHub stars and 1.5 million user-created agents. Steinberger chose OpenAI over building his own company, saying he wants to “change the world, not build a large company.” Gartner already rates OpenClaw an “unacceptable cybersecurity risk.” Now it has OpenAI’s backing.

And for the story that writes itself: a KPMG Australia partner was fined A$10,000 for using AI to cheat on an internal training exam about responsible AI use. Twenty-eight employees got caught. The firm didn’t self-report until regulators came asking. On the hardware side, Sony is pushing the PS6 back to 2028 or 2029 because AI chip demand has made RAM too expensive for consumer hardware. The AI industry is now literally delaying your next PlayStation.

Latest from PulseMark



Gemini 3 Deep Think Hits 84.6% on ARC-AGI-2 — But Only 6.5% of Real Research Problems Got Useful Answers

Google’s Aletheia agent autonomously wrote a publishable math paper—then got 68.5% of open problems fundamentally wrong. The researchers’ own data tells the real story.

Read the deep dive →



Anthropic Built the Gun: Claude Pentagon Military AI Deployment

The safety lab that built Constitutional AI now powers military intelligence analysis. How Anthropic’s own infrastructure made the Pentagon deployment inevitable.

Read the analysis →

That’s Monday. Six stories, zero fluff.

— The PulseMark Team

Get the Daily Pulse

Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.

Get the Daily Pulse

Sharp AI analysis, daily. Two minutes, every morning.

Get the Daily PulseTwo minutes, every morning