The race broke open. TechCrunch reports OpenAI released GPT-5.5 on April 23, and the model scores 60 on Artificial Analysis’s Intelligence Index—three points clear of Claude Opus 4.7 and Gemini 3.1 Pro, both sitting at 57. On Tuesday we flagged that those three models were within half a point of each other, a gap Artificial Analysis itself described as effectively a tie. Forty-eight hours later, OpenAI reports 82.7% on Terminal-Bench 2.0 (up from 75.1% for GPT-5.4), 73.1% on Expert-SWE (up from 68.5%), 84.9% on GDPval across 44 knowledge-work occupations, 78.7% on OSWorld-Verified for computer use, and 98.0% on Tau2-bench Telecom with no prompt tuning. API pricing is $5 input and $30 output per million tokens with a 1M-token context window, with access rolling out to ChatGPT Plus, Pro, Business, and Enterprise tiers. A tie at the top of a benchmark only holds until someone retrains the foundation underneath it.
The tie broke from the other end on the same 48-hour window. Simon Willison covered DeepSeek’s V4 preview release on April 24, two MIT-licensed open-weight models posted on Hugging Face for download and modification. V4-Pro is a 1.6-trillion-parameter mixture-of-experts model with 49 billion active parameters and a 1M-token context; V4-Flash is 284 billion total with 13 billion active. API pricing lands at $1.74 / $3.48 per million tokens for Pro and $0.14 / $0.28 for Flash, which puts GPT-5.5 output at roughly 8.6x the per-token cost of V4-Pro output. DeepSeek’s own benchmarks place V4-Pro roughly 3–6 months behind GPT-5.4 and Gemini 3.1 Pro, and the company attributes the efficiency to a technique it calls Hybrid Attention Architecture—V4-Pro uses 27% of the single-token FLOPs of V3.2 at 1M-token contexts, and V4-Flash uses 10%. Frontier performance and frontier pricing are now decoupling on different curves.
The third move reframed the competition entirely. Google announced the Gemini Enterprise Agent Platform on April 22, an evolution of Vertex AI into a single build/scale/govern/optimize layer for agents, backed by a $750 million partner fund targeting Google Cloud’s 120,000-member partner ecosystem. The Model Garden exposes more than 200 models—Gemini 3.1 Pro, Gemma 4, and third-party Claude Opus, Sonnet, and Haiku—and the Agent Runtime claims sub-second cold starts and multi-day workflow support. Launch partners include Adobe, Atlassian, Deloitte, Oracle, Palo Alto Networks, Replit, S&P Global, Salesforce, ServiceNow, and Workday, and the platform ships with Agent Identity (cryptographic IDs and audit trails) plus Agent Anomaly Detection. While OpenAI and DeepSeek trade blows on base-model intelligence and per-token price, Google is arguing that the enterprise contest is no longer about which chat model scores highest—it’s about which platform runs the agents on top of them.
That’s Friday. Three stories, zero fluff.
— The PulseMark Team
Get the Daily Pulse
Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.
