The machines crossed the line. OpenAI’s GPT-5.4 scored 75.0% on OSWorld-Verified—surpassing the 72.4% human baseline on real-world computer use tasks. Not toy demos. Not cherry-picked screenshots. Full desktop workflows: navigating browsers, writing terminal commands, operating spreadsheets. We dug into what the benchmark actually measures and where GPT-5.4 pulls ahead of Claude—investment banking tasks hit 87.3% in Thinking mode versus GPT-5.2’s 68.4%, Terminal-Bench coding reached 75.1% against Claude’s 65.4%, and BrowseComp hit 82.7% overall, up 17 points from GPT-5.2. At $2.50 per million input tokens—half of Claude Opus 4.6’s price—the model that just surpassed human-level computer use is also the cheaper option. The human baseline isn’t the ceiling anymore; it’s the floor.
And it’s not just computer use where AI is outpacing humans—it’s finding the bugs humans miss. Anthropic pointed Claude Opus 4.6 at nearly 6,000 C++ files in the Firefox codebase, and in two weeks the model submitted 112 unique reports that led to 22 CVEs—14 of them rated high-severity. That 14 alone represents almost a fifth of all high-severity Firefox remediations in 2025. The first Use After Free vulnerability took 20 minutes to surface. Anthropic spent $4,000 in API credits attempting exploitation and landed 2 successful exploits, with fixes shipping to hundreds of millions of users in Firefox 148.0. When a model finds more critical vulnerabilities in 14 days than human auditors flag in two months, the economics of security research change permanently.
The logical next step is cutting humans out of the loop entirely—and Cursor just shipped the architecture for it. Its new Automations feature triggers coding agents from Slack messages, PagerDuty incidents, code changes, and timers—no human prompt required. Cursor is already running hundreds of automations per hour internally, with agents handling incident response, log queries, and weekly codebase summaries. The company’s annualized revenue run rate doubled to over $2 billion in three months. Between GPT-5.4 beating humans at desktop tasks, Claude outfinding humans on security bugs, and Cursor removing the human trigger from agent workflows, the pattern is clear: AI isn’t just matching human performance—it’s replacing human presence in the loop.
Latest from PulseMark
![]() |
GPT-5.4 Tops Humans at Computer Use and Rivals Microsoft Copilot
The 43% input price hike over GPT-5.2, 47% token savings on MCP Atlas tasks, and what Copilot should be worried about as OpenAI moves into enterprise desktop automation. |
That’s Friday. Three stories, zero fluff.
— The PulseMark Team
Get the Daily Pulse
Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.

