Newsbusiness5 min read
DeepSeek is retiring its flagship into its cheap model, and the cheap one has vision
Ahmad JSep 10, 2026Updated Sep 15, 2026

DeepSeek released V4.1-Flash on 10 September 2026 and, in the same set of notes, said it is retiring the model that until that morning was its most expensive. Not deprecating it in a year. Routing it away in four days.
The footnote on DeepSeek's own pricing page is the clearest statement of it:
"After extensive testing, V4.1 Flash has comprehensively surpassed V4 Pro in performance, cost, speed, and total time, so we plan to retire V4 Pro in an orderly manner. From 12:00 Beijing Time on September 14, 2026, and until V4.1 Pro is released in the future, requests to deepseek-v4-pro will all be routed to V4.1 Flash and billed at the V4.1 Flash price."
The release note gives the same instant in a different clock, "Starting at 04:00 UTC on Sept 14, 2026". Beijing runs eight hours ahead of UTC, so the two pages agree to the minute. That is worth checking rather than assuming, because two pages from one company on one day do not always agree.
What the prices actually do#
DeepSeek bills at a peak rate and an off-peak rate, and the pricing page states the split:
"Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday (all other hours are off-peak)."
That is 35 hours out of 168, so most usage pays the lower number.
Comparing like with like, at peak, cache-miss input, per million tokens:
Input and output each fall to a little under a quarter of what the outgoing flagship charged. The cached-input line falls much further than that, to about a seventh, and DeepSeek says why it built for that: "Cache-hit charges often account for a large share of agent costs. Compressing the cache cuts those costs significantly." An agent that re-reads the same long prompt on every turn spends most of its input budget on cache hits, so a seven-fold cut on that line is not a rounding difference in a real bill.
The concurrency limits move the same way. The table gives deepseek-flash 2500 and deepseek-v4-pro 500, five times as many simultaneous requests on the cheaper model.
The strange part is in the features table#
Two rows above the prices, the same table lists what each model can do.
So the model being retired is the one that cannot see, and the model absorbing its traffic is the one that can. DeepSeek describes V4.1-Flash in its release note as "the smallest model in our new architecture family, with native visual understanding". That is an unusual direction for a lineup to move: normally the cheap tier is the one that gives something up, and on context and output nothing is lost either.
What DeepSeek says it changed#
The architecture claims are DeepSeek's own and are not independently verified here.
Asymmetry is the point: reading is made cheaper than writing, rather than both running through the same active slice. The cache figures are the mechanism behind the cache-hit price above.
On quality it says two things and they are worth separating. The first is its own benchmark claim, that new pre-training and larger-scale reinforcement learning post-training "deliver benchmark results ahead of flagship models, including DeepSeek-V4-Pro". The second is an appeal to others:
"Tests by multiple parties put V4.1-Flash ahead of V4-Pro on performance, cost, speed & total runtime."
The parties are not named. Treat both as the vendor's position, which the retirement decision at least makes expensive to be wrong about.
Weights and a technical report are on Hugging Face, so the claims are checkable by anyone with the hardware to check them.
Why this matters if you are not a DeepSeek customer#
Three things generalise.
Our AI model comparison puts the current rows side by side for exactly that check, and the model picker narrows it by what you are actually doing.
If you are paying a monthly subscription rather than per token, none of this reaches you directly, and the arithmetic in which AI subscription you should actually pay for is unchanged. If you are billed per token, it is worth knowing what a token is and how you are charged for one before comparing any two of these numbers, because input, cached input and output are three different prices and only one of them is the one people quote.
Sources
- DeepSeek-V4.1-Flash Release, DeepSeek API Docs, 10 September 2026api-docs.deepseek.com
- Models & Pricing, DeepSeek API Docsapi-docs.deepseek.com



Discussion