Intelligence just got cheaper

The price moved first. Z.ai says GLM-5.3’s post-training jump came from the same base model as GLM-5.2, yet its in-house code benchmark improved 50% and CyberGym reached 84.5.

The price moved first. Z.ai says GLM-5.3’s post-training jump came from the same base model as GLM-5.2, yet its in-house code benchmark improved 50% and CyberGym reached 84.5. Google then shipped Gemini 3.7 Flash just 3 weeks after 3.6 at half the earlier introductory price, while DeepSeek scheduled new V4 rates for August 16. Three releases point the same way: the frontier is becoming less about one enormous pretraining run and more about extracting, packaging, and pricing capability faster.

Google’s workhorse reset makes the economics tangible. Gemini 3.7 Flash pricing starts at $0.75 per million input tokens and $3.75 for output through December 31, while FrontierCode rose to 43.6% from 34.4%. DeepSeek’s off-peak schedule cuts its own peak rate by 50%, and GLM-5.3’s promised weights arrive in roughly 2 weeks. Better models are landing on shorter clocks with sharper discounts—a combination that rewards teams able to route workloads, retest benchmarks, and switch providers without rebuilding the application around each release.

DeepSeek is turning that flexibility into a timetable. V4-Pro’s new rate card creates 2 peak windows, offers 3 reasoning-effort levels, and prices output at $1.98 off-peak versus $3.96 during peak hours. Gemini’s introductory discount expires after December 31, while Z.ai plans to release weights only after 2 weeks of additional safety hardening. The takeaway is operational: model selection now changes by hour, release channel, and training stage, so static annual procurement looks increasingly mismatched to a market repricing itself in days.

The Pulse

Anthropic finds agent swarms can coordinate, collude, and sabotage
Parallel intelligence still needs governance before emergent teamwork becomes an operational liability.

What open source projects taught GitHub about AI-era security
AI accelerates triage, but maintainers still own judgment, accountability, and every merge.

White House opens a path for private offensive cyber operations
Public-private hacking demands unusually clear authority, boundaries, and mechanisms for stopping escalation.

Latest from PulseMark



How to Detect Claude AI Content: Watermarks and C2PA

Separate Claude’s text signal from C2PA Content Credentials, then inspect supported files without treating either lane as universal proof.

Read the guide →



Grok 4.6 API Tutorial: Reasoning, Tools, and Costs

Call Grok 4.6 with reasoning controls and native tools, then route storage, caching, and spend deliberately.

Read the tutorial →



How to Run Nemotron 3.5 Lightning Locally With llama.cpp

Pick a quant that fits nearly 24 GiB at Q4, build llama.cpp, and verify the model’s sparse local runtime.

Read the tutorial →

That’s the signal.

— The PulseMark Team

Get the Daily Pulse

Sharp analysis on what's actually moving in AI. No hype, no filler, no weekly digest.

Get the Daily Pulse

Sharp AI analysis, daily. Two minutes, every morning.

Get the Daily PulseTwo minutes, every morning