DeepSeek cuts Flash prices and bills agents by time of day
New Flash rates took effect on 10 September 2026: off-peak, 0.02 yuan per million cached input tokens and 4 yuan per million output tokens. Peak hours cost twice as much.

DeepSeek rebuilt its price list alongside the launch of V4.1-Flash. From 10 September 2026, starting at 12:00 Beijing time, the Flash series carries these rates. Off-peak, 0.02 yuan per million input tokens that hit the cache, 1 yuan per million input tokens that miss it, and 4 yuan per million output tokens. At peak, every one of those figures doubles.
That is a real cut against the previous tariff. The outlet IT之家 compared deepseek-v4-flash rates before and after the change. The cache-hit price fell from 0.05 to 0.02 yuan per million tokens, down 60 percent. Input without a hit went from 1.5 to 1 yuan, a drop of one third, and output from 4.5 to 4 yuan.
Peak, valley and the weekend
DeepSeek introduced the peak-and-valley mechanism on 17 August 2026. On 23 August it sharpened the rules: weekends no longer distinguish peak from valley at all, so Saturdays and Sundays are billed at valley rates. On 19 September the vendor clarified the calendar further. Weekends given over to make-up work days for public holidays, and full Chinese statutory holidays, count as off-peak. At contexts in the hundreds of thousands of tokens, a doubled input price changes the budget of an entire deployment.
V4-Pro was meant to disappear, but stays
The launch came with an operation on the older model. On 9 September DeepSeek announced that once V4.1-Flash went live it would route V4-Pro requests to V4.1-Flash and bill them at V4.1-Flash rates, until V4.1-Pro arrived. Retirement of V4-Pro was set for 14 September 2026 at 12:00.
Two days after the launch the vendor changed its mind. On 11 September it said that in response to user demand the V4-Pro service would remain available after 14 September, with no change to how it is billed. V4-Pro pricing is clearly higher than Flash: a cache hit costs 0.15 yuan in the valley and 0.30 at peak, input without a hit 4.5 and 9 yuan, against Flash rates of 0.02/0.04 and 1/2 yuan respectively.
The vendor's official statement adds a second element: DeepSeek is opening a partnership with agent tools. WorkBuddy, together with CodeBuddy, and OpenCode are to support V4.1-Flash fully from launch day. For users this means the price cut lands at the same time as the harnesses in which the model actually runs.
Sources
4- 01DeepSeek 官宣明日 flash 系列 AI 模型降价 (IT之家)ZH
- 02DeepSeek 将继续提供 V4 Pro API 调用服务 (IT之家)ZH
- 03DeepSeek:调休上班的周末、法定节假日全天按空闲时段计费 (IT之家)ZH
- 04DeepSeek V4.1-Flash — cennik szczytowy i pozaszczytowyEN
All figures and quotations in this text come from the sources listed below.
Content prepared by the editorial team with AI assistance.
Comments
0- No comments yet — be the first.