The headline: DeepSeek shipped V4-Flash production release on July 31. No keynote. No hype thread. Just an API that’s live, and benchmark scores that make the April preview look like a beta. V4-Pro is next.
The numbers that matter
DeepSWE: 7.3 → 54.4. That’s a 6.5× jump in software engineering capability. Same model architecture. Same parameter count. The only difference? Post-training.
Terminal Bench 2.1: 82.7. Two points behind Claude Opus 4.6 — a model with ~25× the active parameters. Three points ahead of the V4-Pro preview from April.
Price: ~$0.28/M input tokens. That’s 1/10 of GPT-4o. For a coding agent that calls the model 50 times per task, the math just changed.
What’s actually new
- Agent performance is the story. Terminal Bench, DeepSWE, Cybergym, Toolathlon — every agent benchmark jumped. This isn’t a better chatbot. It’s a better autonomous worker.
- Same architecture, better training. 284B total, 13B active, MoE, 1M-token context. The skeleton didn’t change. The brain did.
- MIT license. No custom license games. Deploy it anywhere.
- Native Responses API. Drop-in replacement for OpenAI’s agent-oriented API format.
What’s coming
V4-Pro production release — 1.6T parameters, 49B active — is expected in early August. If the post-training gains scale proportionally, the heavyweight division is about to get a new contender.
The immediate take
You don’t need to wait for Pro. V4-Flash is live, it’s MIT-licensed, it costs 1/10 of GPT-4o, and it scores within striking distance of Claude Opus on agent tasks. Add it to your evaluation pipeline today. For the engineering story behind those numbers, see our V4-Flash architecture deep dive, or the full launch analysis for the complete picture including what Pro changes.
V4-Flash is live. It’s already on MonkeyCode. Go.
No waitlist. No credit card. No catch.
👉 Start coding with V4-Flash now — 30M free tokens/day, zero setup.