OpenAI launched GPT-5.4 on March 5, 2026 with a headline capability that clears a milestone no general-purpose model has reached: computer use that beats humans. GPT-5.4 scores 75.0% on OSWorld-Verified, above both the 72.4% human baseline and Claude Opus 4.6’s 72.7%. But the more strategic move is a beta Excel add-in called ChatGPT for Excel, loaded with Moody’s, Dow Jones, and MSCI data, that positions ChatGPT as a direct competitor to Microsoft Copilot โ on Microsoft’s own platform.
The timing matters. Six days after ChatGPT uninstalls surged 295% following OpenAI’s Pentagon deal, GPT-5.4 arrives with an enterprise-first feature set โ financial data partners, investment banking benchmarks, and pricing that undercuts Claude Opus 4.6 by half on input and 40% on output. This is a technically impressive release. It’s also a trust-rebuilding exercise aimed at the high-value users who were most alienated.
GPT-5.4 Is the First General-Purpose Model to Beat Humans at Computer Use
Anthropic pioneered AI computer use with Claude’s beta in October 2024. But GPT-5.4 is the first model to surpass human performance on the standard benchmark for it. On OSWorld-Verified, GPT-5.4 scores 75.0% โ above both the 72.4% human baseline and Claude Opus 4.6’s 72.7%. For context, GPT-5.2 scored 47.3% on the same test. That’s not incremental improvement; it’s a generational leap.
The implementation runs in two modes. Code mode writes Python with Playwright to navigate web pages, click buttons, and fill forms. Screenshot mode issues raw mouse and keyboard commands from visual input alone. Both feed into a build-run-verify-fix loop โ the model confirms task completion before declaring success, which is the kind of reliability feature that matters when you’re building agents that run without human oversight.
OpenAI calls it “the strongest model currently available for developers building agents.” The critical distinction: this is the first time computer use is native to a general-purpose model, not a separate specialized capability bolted on after the fact. Coding, reasoning, and desktop automation live in the same model, which simplifies agent architecture considerably.
The Excel Add-In That Competes With Microsoft
ChatGPT for Excel is a sidebar add-in โ currently in beta for Business, Enterprise, Edu, Pro, and Plus users in the US, Canada, and Australia. Users prompt in natural language; calculations execute in Excel, so every formula is visible and verifiable. The distinction from Copilot is subtle but important: ChatGPT for Excel treats the spreadsheet as the execution environment, not as a black box that returns answers.
The financial data integrations are where it gets strategic. Launch partners include Moody’s, Dow Jones Factiva, MSCI, Third Bridge, and MT Newswire. TechFundingNews reported that FactSet, S&P Global, LSEG, and Daloopa are coming soon. That list maps almost exactly to Bloomberg Terminal data sources โ and a Bloomberg Terminal now costs roughly $32,000 per year per seat (up from ~$24,000 five years ago). ChatGPT Plus costs $240.
The investment banking benchmark progression tells the story in numbers. GPT-5.2 hit 68.4% on tasks like building three-statement financial models with proper formatting and citations. GPT-5.4 Thinking reached 87.3%. That’s not a model getting marginally better at finance โ it’s crossing the threshold from “can’t do the job” to “can do the job.”
The strategic tension is obvious: OpenAI is building a product that competes with Microsoft Copilot for Excel while Microsoft remains OpenAI’s largest investor, primary distribution channel, and the company that powers Copilot with OpenAI’s own models. Google Sheets integration is planned, which means the Office integration strategy isn’t even Microsoft-exclusive.

What GPT-5.4 Actually Costs โ and How It Stacks Up
GPT-5.4’s standard tier runs $2.50 per million input tokens and $15 per million output tokens, with cached input at $0.25 per million. That’s half the input cost of Claude Opus 4.6 ($5/$25) and slightly above Gemini 3.1 Pro ($2/$12). It also represents a 43% input price increase over GPT-5.2’s $1.75/$14 โ though OpenAI argues improved token efficiency offsets the higher per-token cost.
At scale, the gap compounds. Processing 100 million output tokens costs $1,500 with GPT-5.4 versus $2,500 with Claude Opus 4.6 โ a $1,000 difference that matters to any team running agents in production. Gemini 3.1 Pro is still cheapest at $1,200 for the same volume, but it doesn’t offer native computer use.
The Pro tier is a different product entirely: $30 per million input, $180 per million output โ 12x the standard tier. For that premium, you get ARC-AGI-2 at 83.3% and BrowseComp at 89.3%. No single model wins every benchmark. DigitalApplied’s comparison shows GPT-5.4 leading on GDPval professional tasks (83.0% vs. Opus 78.0%) and Terminal-Bench coding (75.1% vs. 65.4%), while Claude Opus 4.6 holds SWE-Bench Verified and multi-agent coordination, and Gemini 3.1 Pro leads on GPQA Diamond scientific reasoning (94.3%) and price.
Tool Search: The Developer Feature That Changes Agent Architecture
If you’ve built agents with more than a handful of MCP servers, you know the problem: tool definitions eat context windows before reasoning even starts. With 36 MCP servers enabled, the model spends its token budget describing tools rather than using them. GPT-5.4’s Tool Search feature addresses this directly โ the model receives a lightweight tool list and retrieves full definitions on demand.
On 250 tasks from Scale’s MCP Atlas benchmark with 36 MCP servers, Tool Search cut total token usage by 47% while maintaining the same accuracy. Separately, GPT-5.4 scores 82.7% on BrowseComp โ up 17 points from GPT-5.2 โ reflecting the base model’s stronger web search and synthesis capabilities. As developer @rohanpaul_ai noted on X: “This model running on medium reasoning now matches the quality of the older generation running on extra-high.”
The real unlock isn’t the token savings โ it’s the ceiling removal. Agents can now connect to hundreds of integrations without context-window degradation, which was the practical bottleneck preventing complex multi-tool workflows from scaling. Two days before GPT-5.4, OpenAI shipped GPT-5.3 Instant โ an aggressive release cadence that shows the company is iterating faster than any frontier lab in the field.
The Pentagon Shadow Over a Technical Win
OpenAI’s Pentagon deal, announced on February 27, triggered the most visible consumer backlash in the company’s history. TechCrunch reported 295% day-over-day uninstall spikes, roughly 1.5 million users signed a “Quit GPT” campaign, and Claude downloads jumped 37-51% across two days. That 1.5 million represents just 0.17% of ChatGPT’s ~900 million weekly users โ not existential, but disproportionately weighted toward developers and technical professionals who influence enterprise procurement.
The safety card adds a genuine tension. GPT-5.4 controls only 0.3% of 10,000-character reasoning traces โ it essentially cannot hide its chain of thought, which is the best empirical evidence yet that CoT monitoring works as a safety tool. But it also received the first-ever “High Capability” cybersecurity rating for a general reasoning model, meaning it can automate end-to-end attacks on protected targets. The model can’t obscure what it’s thinking. What it’s thinking is now more dangerous.
What GPT-5.4 Means for OpenAI’s Business
The question nobody can answer yet: how does Microsoft respond? OpenAI’s Excel add-in is live on the same platform where Microsoft sells Copilot, powered by the same OpenAI models that Microsoft licenses. At what point does Microsoft decide it’s cheaper to acquire a competing model provider than to keep subsidizing the one that’s eating its Office revenue?
OpenAI didn’t launch a spreadsheet tool. It launched a distribution channel that bypasses its most important business partner, stocked with data sources that cost roughly $32,000 a year per seat to access through existing terminals. Every data provider on OpenAI’s launch list and coming-soon list โ Moody’s, Dow Jones, MSCI, FactSet, S&P Global, LSEG โ already sells to the same financial institutions that pay for Bloomberg. The channel is different. The data is the same. And the next time Microsoft renegotiates its OpenAI investment terms, the Excel add-in will be on the table.
Get the Daily Pulse
Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.



