DeepSeek launches V4.1 Flash with native multimodal support
DeepSeek released V4.1 Flash on September 10, bringing native multimodal visual understanding to its API under the deepseek-flash model name and setting out a transition from V4 Pro.
DeepSeek released DeepSeek-V4.1-Flash on September 10, 2026, bringing native multimodal visual understanding to its API under the new deepseek-flash model name, according to the company’s API changelog.
DeepSeek’s April 24 V4 API rollout added V4 Pro and the earlier V4 Flash model. The September 10 release is distinct because it introduces a new model version and API name, adds native multimodal support, retires two earlier Flash models and sets replacement routing for V4 Pro.
With the release, DeepSeek is retiring two earlier Flash-branded models: DeepSeek-V4-Flash, which entered public beta on July 31, and the experimental DeepSeek-V4-Flash-Vision-Exp, released August 21. The company said it will temporarily accept requests using the legacy deepseek-v4-flash and deepseek-v4-flash-vision-exp names, routing them to V4.1 Flash at V4.1 Flash pricing. It has not said when that support will end.
DeepSeek also set out a transition from its V4 Pro model. From 12:00 Beijing time on September 14, 2026, requests using the deepseek-v4-pro model name will route to V4.1 Flash and be billed at V4.1 Flash rates. That arrangement will remain in place until DeepSeek releases V4.1 Pro, for which the company has not announced a date.
“Today, we officially release the DeepSeek-V4.1-Flash model,” DeepSeek said in the changelog.
DeepSeek describes V4.1 Flash as having native multimodal visual understanding, and its pricing page lists Vision as a supported API capability. The earlier Vision-Exp release was a separate experimental model accessed through its own model name.
The page lists off-peak prices of $0.003 per million input tokens for a cache hit, $0.15 per million input tokens for a cache miss, and $0.60 per million output tokens. Peak-hour rates are exactly double. It also lists a one-million-token context window, a maximum output of 384,000 tokens, and an account-level concurrency limit of 2,500 requests.
DeepSeek said its own testing found V4.1 Flash outperformed V4 Pro on performance, cost, speed, and total completion time. The company reported scores of 90.9 on GPQA Diamond, a 3471 Codeforces rating, 90.6 on Terminal-Bench 2.1, and 74.2 on DeepSWE v1.1. The figures and broader V4 Pro comparison are company claims and were not independently verified for this article.
More news

AWS releases six open-source Hugging Face deployment skills for SageMaker

Google Research releases MilleMiglia logistics benchmark generator

AWS launches AgentCore Runtime V2 with elastic memory and snapshot starts
