DeepSeek pushes V4-Flash API into public beta at cut-rate pricing
DeepSeek released its V4-Flash API into public beta on July 31, 2026, scoring 82.7 on Terminal-Bench 2.1, near Anthropic's Claude Opus 4.8.
DeepSeek released its V4-Flash API into public beta on July 31, 2026, a cheaper, agent-tuned model the company says outscores its own larger V4-Pro-Preview on a leading coding benchmark.
On Terminal-Bench 2.1, a test of AI agents working in a command line, DeepSeek-V4-Flash scored 82.7, versus 72.1 for V4-Pro-Preview and 85.0 for Anthropic’s Claude Opus 4.8, according to figures on the model’s Hugging Face card. The scores are DeepSeek’s own and have not been independently verified.
V4-Flash is a Mixture-of-Experts model, an architecture that activates only part of the network per query, with roughly 13 billion active parameters, a 1-million-token context window and re-training for agent and coding work. API pricing is $0.14 per million input tokens and $0.28 per million output tokens, with cached input from $0.0028, undercutting most frontier rivals.
Other reported results show sharp movement on agent tasks: the DeepSWE score jumped to 54.4 from 7.3 in the July preview. The weights ship under an MIT license as DeepSeek-V4-Flash-0731; the V4-Pro API and DeepSeek’s consumer app and web models are unchanged.
The framing to watch is price versus parity. A smaller model beating a larger sibling on benchmarks says as much about post-training as raw scale, and benchmark leads rarely survive contact with real workloads. Still, at roughly a fifth of typical frontier output pricing, V4-Flash sharpens the cost pressure Chinese labs keep applying to U.S. model makers.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
