GPT-5.2-Codex: OpenAI’s Context Compaction Breakthrough Changes Agentic Coding

OpenAI released GPT-5.2-Codex on December 18, 2025—exactly one week after launching GPT-5.2 base and just eleven days after CEO Sam Altman’s “code red” memo in response to Gemini 3’s benchmark dominance. This isn’t just a coding model. It’s OpenAI’s answer to the context problem that has plagued every AI coding assistant since the beginning: what happens when your project exceeds the model’s memory?

The headline number: 64% on Terminal-Bench 2.0, beating Claude Opus 4.5’s 59.3%. But the real story is context compaction—a new capability that lets GPT-5.2-Codex work coherently across millions of tokens in a single coding session. That means 24+ hour tasks without losing track of what it was doing. For developers, this changes everything about how agentic coding actually works.

The Benchmark Battle: Where GPT-5.2-Codex Wins (and Loses)

Let’s cut through the marketing and look at the numbers. GPT-5.2-Codex claims state-of-the-art performance on agentic coding benchmarks, but the competitive landscape is nuanced.

BenchmarkGPT-5.2-CodexClaude Opus 4.5Gemini 3 Pro
SWE-Bench Verified80.0%80.9%76.2%
SWE-Bench Pro56.4%~54%~52%
Terminal-Bench 2.064.0%59.3%~55%
CVE-Bench (Security)87%~80%~75%

On SWE-Bench Verified—the standard benchmark for real-world software engineering tasks—Claude Opus 4.5 still edges out GPT-5.2-Codex by 0.9 percentage points. Not a decisive lead, but Anthropic maintains the crown they’ve held since we first compared Claude and GPT earlier this year.

Where OpenAI pulls ahead is Terminal-Bench 2.0, a newer benchmark focused on command-line and terminal operations. The 64% score beats Claude’s 59.3%—a meaningful gap for developers who spend their lives in the terminal. And CVE-Bench, measuring security vulnerability detection, shows GPT-5.2-Codex at 87%, suggesting OpenAI is serious about positioning this model for cybersecurity use cases.

One benchmark worth scrutinizing: OpenAI claims GPT-5.2 Thinking achieves 70.9% on their own “GDPval” benchmark, measuring AI performance on knowledge work tasks. They claim this ties or beats human experts. But GDPval is OpenAI’s internal benchmark—not independently validated—so treat that number with appropriate skepticism.

Context Compaction: The Technical Breakthrough That Actually Matters

Every developer who’s used AI coding assistants knows the frustration: start a complex refactoring task, and three hours in, the model has forgotten what you were trying to accomplish. Context windows have expanded—Claude offers 200K tokens, Gemini 3 pushes to 2M—but raw token capacity isn’t the same as coherent long-term task execution.

GPT-5.2-Codex introduces context compaction, a technique OpenAI describes as “loss-aware compression of conversation state.” In practice, the model intelligently summarizes and prioritizes information as the context grows, maintaining coherence across what OpenAI claims can be 24+ hour coding sessions.

The technical details from OpenAI’s system card are sparse, but the implications are significant. This addresses a real limitation: previous models would complete tasks successfully in short bursts but fail on complex, multi-day refactoring or feature implementations. If context compaction works as advertised, it’s the first genuine architectural solution to the “model amnesia” problem.

Benchmark comparison chart showing GPT-5.2-Codex performance versus Claude Opus 4.5 and Gemini 3 Pro on coding benchmarks

The Cybersecurity Pivot: Real Vulnerabilities, Controlled Access

OpenAI buried the lead on something potentially significant: a cybersecurity pilot program offering enhanced access to vetted security professionals. This isn’t just marketing—there’s already proof of concept.

A researcher using GPT-5.1-Codex-Max (the predecessor) discovered and responsibly disclosed three React vulnerabilities: CVE-2025-55183, CVE-2025-55184, and CVE-2025-67779. The 87% CVE-Bench score suggests GPT-5.2-Codex is even better at this.

The pilot program requires:

  • Demonstrated track record of responsible disclosure
  • Clear professional cybersecurity use case
  • Invitation-only access

This is a calculated move. OpenAI gets real-world testing on defensive security applications while maintaining control over potential dual-use concerns. The approach mirrors what Anthropic disclosed when they disrupted an AI-orchestrated cyberattack earlier this year—responsible capability building with guardrails.

Pricing and Availability: The API Gap

GPT-5.2-Codex is available now for ChatGPT Plus, Team, Pro, and Enterprise subscribers. The API? “Coming in the coming weeks.” That’s a problem for developers building production applications who can’t wait on an unspecified timeline.

When the API does arrive, expect pricing similar to GPT-5.2 base: $1.75 per million input tokens, $14 per million output tokens. That’s a 1.4x increase over GPT-5.1 pricing. For context, Gemini 3 Pro runs 6.7x cheaper than GPT-5.2 at equivalent quality levels. OpenAI is betting on capability differentiation, not price competition.

The cost calculus gets interesting with context compaction. If the feature genuinely enables longer, more coherent sessions, the total cost per completed task might actually decrease despite higher per-token pricing. That’s the efficiency trade-off OpenAI seems to be banking on.

The Competitive Context: Code Red Continues

This release makes more sense when you remember the timeline. On December 2, Sam Altman sent his “code red” memo after Gemini 3 dominated key benchmarks. December 11 brought GPT-5.2 base. December 18 brought GPT-5.2-Codex.

OpenAI is shipping at startup speed—three major releases in seventeen days. Whether that pace introduces quality concerns remains to be seen. The system card acknowledges ongoing safety work, and the delayed API rollout suggests OpenAI isn’t rushing everything.

The competitive picture:

  • Claude Opus 4.5 still leads on SWE-Bench Verified (80.9% vs 80.0%)
  • Gemini 3 Pro leads on LiveCodeBench Pro (2,439 vs ~2,300 Elo) and cost efficiency
  • GPT-5.2-Codex leads on Terminal-Bench 2.0, CVE-Bench, and abstract reasoning (54.2% vs Claude’s 37.6%)

No single model dominates across all dimensions. The “best” coding model depends on your specific use case—security analysis, terminal automation, or general software engineering.

What This Means for Developers

If you’re using ChatGPT for coding, the upgrade is automatic and worth testing. The context compaction feature addresses a genuine pain point, and the terminal and security improvements are meaningful for those workflows.

If you’re building applications on the API, you’re waiting. The “coming weeks” timeline is frustrating given that competitors have their latest models available now. When GPT-5.2-Codex API does arrive, the 1.4x price increase over 5.1 will factor into production cost calculations.

If you’re doing security research, the pilot program is worth applying for. Real vulnerability discoveries with responsible disclosure suggest this isn’t vaporware—the model genuinely performs on security-relevant tasks.

The bottom line: GPT-5.2-Codex is a genuine advancement, particularly for long-running agentic tasks. It doesn’t definitively beat the competition on every metric, but context compaction is the kind of capability that changes workflows. OpenAI’s rapid shipping pace shows they’re taking the Gemini 3 challenge seriously. Whether that urgency introduces hidden trade-offs, we’ll learn as developers put the model through real-world production loads.

Get the Daily Pulse

Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.

Get the Daily Pulse

Sharp AI analysis, daily. Two minutes, every morning.

Get the Daily PulseTwo minutes, every morning