DeepSeek pushes V4-Flash API into public beta at cut-rate pricing
DeepSeek released its V4-Flash API into public beta on July 31, 2026, scoring 82.7 on Terminal-Bench 2.1, near Anthropic's Claude Opus 4.8.
DeepSeek released its V4-Flash API into public beta on July 31, 2026, a cheaper, agent-tuned model the company says outscores its own larger V4-Pro-Preview on a leading coding benchmark.
On Terminal-Bench 2.1, a test of AI agents working in a command line, DeepSeek-V4-Flash scored 82.7, versus 72.1 for V4-Pro-Preview and 85.0 for Anthropic’s Claude Opus 4.8, according to figures on the model’s Hugging Face card. The scores are DeepSeek’s own and have not been independently verified.
V4-Flash is a Mixture-of-Experts model, an architecture that activates only part of the network per query, with roughly 13 billion active parameters, a 1-million-token context window and re-training for agent and coding work. API pricing is $0.14 per million input tokens and $0.28 per million output tokens, with cached input from $0.0028, undercutting most frontier rivals.
Other reported results show sharp movement on agent tasks: the DeepSWE score jumped to 54.4 from 7.3 in the July preview. The weights ship under an MIT license as DeepSeek-V4-Flash-0731; the V4-Pro API and DeepSeek’s consumer app and web models are unchanged.
The framing to watch is price versus parity. A smaller model beating a larger sibling on benchmarks says as much about post-training as raw scale, and benchmark leads rarely survive contact with real workloads. Still, at roughly a fifth of typical frontier output pricing, V4-Flash sharpens the cost pressure Chinese labs keep applying to U.S. model makers.
More news

Ornith AI open-sources Ornith-1.5, a self-improving LLM family it says rivals Claude Opus 4.8

Krea AI releases Krea-2 open-weight image models with 2-second 2K generation
Dmytro Spodarets·Jun 29, 2026
DeepReinforce releases Ornith-1.0, an open-weight coding model that writes its own RL scaffolds
Dmytro Spodarets·Jun 29, 2026