OpenAI killed a product this week. The company is shutting down Sora—the app, the API, and sora.com—dissolving a 3-year, $1 billion Disney partnership before any money changed hands. Sora hit #1 on the App Store after its September 2025 launch, and OpenAI is now valued at roughly $730 billion after a $110 billion funding round. Sam Altman says resources are shifting to coding, reasoning, and a text-generation model called Spud that will “really accelerate the economy”—no further details disclosed. A $730 billion company killing a hit consumer product to double down on enterprise AI tells you exactly where the margins are: not in generating movie clips, but in generating code. And the coding tool market OpenAI is pivoting into is already crowded—Cursor, Windsurf, Claude Code, and Codex are all competing on pricing from $15 to $200 per month, and recent research from Sun Yat-sen University and Alibaba found AI coding agents break working code 75% of the time when making changes. The gap between what these tools promise and what they reliably deliver is the real story OpenAI is walking into.
While OpenAI cuts products, Anthropic is shipping new ones. Claude can now control your Mac remotely through Dispatch—assign a task from your phone, come back to find completed work on your desktop. The agent controls mouse, keyboard, and screen autonomously, prioritizing existing integrations like Slack and Google Calendar before falling back to screen control. It’s available to Pro ($20/month) and Max ($100/month) subscribers, macOS only, as a research preview. Safety constraints require explicit per-app permission and prohibit stock trading or sensitive data access. On the other end of the platform spectrum, Apple is planning a standalone Siri chatbot app for iOS 27, code-named Campo and powered by Google’s Gemini models, with a WWDC 2026 reveal targeted for June 8. The app will handle text and voice conversations, access messages and emails, and execute tasks within apps—but the project has faced multiple scheduling delays. Two very different agent strategies: Anthropic ships a desktop controller in research preview while Apple builds on a competitor’s models and targets June.
And none of these agents may be as capable as their demos suggest. ARC-AGI-3 launched this week with $2 million in total prizes—135 interactive game-based environments with no instructions, no descriptions, and no stated win conditions. Every frontier model scored below 1%: Gemini 3.1 Pro Preview at 0.37%, GPT-5.4 at 0.26%, Opus 4.6 at 0.25%, and Grok-4.20 at 0.00%. Untrained humans solve these same tasks at 100%. The benchmark uses RHAE scoring—(human actions / AI actions)²—which penalizes brute-force approaches and measures genuine agentic intelligence in unfamiliar environments. Submissions close November 2, 2026. The models that can write code, control desktops, and pass increasingly sophisticated benchmarks still can’t figure out a puzzle that requires no prior knowledge—just the ability to reason about something new.
Latest from PulseMark
![]() |
Which AI Coding Assistant Fits Your Workflow in 2026?
SWE-bench Pro scores (23–46% across all models), monthly pricing from $15 to $200, Sun Yat-sen/Alibaba research showing agents break working code 75% of the time, and why GPT-5.4’s 81.8% Terminal-Bench lead doesn’t tell the full story. |
That’s Thursday. Four stories, zero fluff.
— The PulseMark Team
Get the Daily Pulse
Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.

