DeepSeek announced on September 14 2026 that it will retire its V4-Pro large-language model (LLM) and route every API request to the newer V4.1-Flash version.
Why DeepSeek is pulling the plug on V4-Pro
V4-Pro has been DeepSeek’s top-tier offering since launch, marketed as a general-purpose LLM with broad knowledge. Over the past year the company measured V4.1-Flash against V4-Pro on the tasks its API customers care about: code generation, terminal command synthesis, and rapid response to short prompts. The data show the newer model runs faster, costs less per request, and scores higher on several task-specific benchmarks.
The shift reflects market pressure. Cloud-based AI services now charge per token or per compute unit, and developers are hunting ways to trim operating expenses without sacrificing speed. By moving all traffic to V4.1-Flash, DeepSeek can shut down the more expensive inference pipeline that powers V4-Pro and pass the savings to users.
Performance and cost at a glance
| Benchmark | V4.1-Flash | V4-Pro |
|---|---|---|
| Terminal-Bench 2.1 (coding & command generation) | 90.6 | 87.9 |
| Terminal-Bench 4.0 (complex multi-step tasks) | 31.2 | 12.4 |
| DeepSWE v1.1 (software-engineering reasoning) | 74.2 | 62.7 |
In each test V4.1-Flash scores higher, sometimes by a wide margin. The model also outperforms Claude Opus 5 and GPT-5.6 Sol on the same coding-focused benchmarks, according to DeepSeek’s internal comparisons.
Cost per typical task shows a similar pattern:
- GLM-5.3-Flash: $0.25
- DeepSeek V4.1-Flash: $0.27
- Moonshot Kimi K3: $2.00
Developers can achieve comparable intelligence for roughly one-eighth the expense when they choose Flash over a larger, pricier alternative.
The trade-off: less encyclopedic knowledge
Efficiency gains come with a dip in factual recall. On the SimpleQA-Verified fact-checking benchmark, V4.1-Flash scores 42.3 while V4-Pro posted 55.2. DeepSeek describes the newer model as “a better worker but a weaker encyclopedia.” Applications that rely heavily on up-to-date world knowledge—news summarisation, legal research—may need external knowledge bases or a fallback to a larger model.
How the industry is positioning itself
DeepSeek’s all-in move contrasts with other AI providers:
- Z.ai offers a spectrum of models, from compact to large, letting customers balance scale and efficiency per project.
- Moonshot doubles down on size with its Kimi K3 model, accepting higher costs for raw capacity.
- DeepSeek has placed every API request on Flash, betting the market will prioritise speed and price over sheer scale.
These divergent approaches show the market has not settled on a single path. Some developers still value the breadth of massive models; others trade that for faster, cheaper responses.
Practical advice for developers
- Don’t pick a model solely by its name or token price. A “Flash” model can outperform a “Pro” model on the tasks you care about.
- Build flexibility into your stack. Design your system to switch between providers or model versions without major rewrites.
- Validate with your own data. Public benchmarks give a useful baseline, but real-world performance depends on your workload.
Following these steps helps teams avoid lock-in and secure the best value as the LLM ecosystem evolves.
What to watch next
Retiring V4-Pro marks a clear pivot toward efficiency-focused LLMs. Whether this ends the “flagship” era or simply starts a new phase of a rapidly shifting market remains to be seen, but developers who stay agile will be best positioned to profit from the change.
