Anthropic measured the wrong thing

Anthropic's latest evaluation shows an 18% sabotage success rate, but zero staff believe Claude can replace entry-level researchers.

They published 53 pages of honesty—and missed the point. Anthropic just dropped its latest sabotage risk evaluation for Claude Opus 4.6, testing whether the model could covertly undermine human oversight. The result: an 18% success rate on covert side-task evaluations. We broke down why the report measures the wrong thing—and why the 0-out-of-16 researcher replaceability score matters more than the sabotage numbers. Zero of 16 Anthropic staff think Claude could replace an entry-level researcher within three months. Those same staff report up to 700% productivity gains. Both can’t be true. Unless the threat isn’t replacement—it’s amplification.

That amplification problem just got personal. The people who should be building these safety evaluations are walking out the door. OpenAI, Anthropic, and xAI have all seen senior safety researchers leave for startups and competing labs, creating a talent vacuum at the exact moment these evaluations need to scale. And a new Harvard Business Review study from UC Berkeley shows why speed matters: AI isn’t reducing workloads—it’s intensifying them. Eight months of embedded research at one tech company found task expansion, blurred boundaries, and burnout as “doing more” became the default. We called it Jevons’ Paradox for knowledge work. Berkeley just proved it.

Lawmakers are starting to notice. New York’s FAIR News Act would require disclaimers on AI-generated news content and mandate human editorial review before publication—backed by the WGA, SAG-AFTRA, and the NewsGuild. The bill won’t fix the sabotage problem or the burnout crisis. But it signals something: the gap between what AI can do and what we’ve prepared for is now too wide to ignore.

Latest from PulseMark



Anthropic’s Sabotage Risk Report Measures the Wrong Thing

Eight sabotage pathways, a 42% autonomy threshold breach, and the contradiction hiding in plain sight.

Read the deep dive →



The AI Safety Talent Exodus Nobody’s Talking About

Senior safety researchers are leaving OpenAI, Anthropic, and xAI faster than they can be replaced.

Read the analysis →



AI Is Jevons’ Paradox for Knowledge Work

The productivity tool that makes you work more, not less. Berkeley’s data tells the story.

Read the deep dive →



ChatGPT Ads Are Here. OpenAI’s Business Model Just Changed.

OpenAI is testing ads in ChatGPT Free. What it means for the AI business model war.

Read the analysis →

That’s Wednesday. Six stories, zero fluff.

— The PulseMark Team

Get the Daily Pulse

Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.

Get the Daily Pulse

Sharp AI analysis, daily. Two minutes, every morning.

Get the Daily PulseTwo minutes, every morning