DeepSeek · 2026-08-13 · major
DeepSeek raises V4 API prices — output costs more than double from August 16
DeepSeek raises API prices for DeepSeek-V4-Pro and DeepSeek-V4-Flash at 16:00 UTC on August 16, 2026, and splits billing into peak and off-peak bands. V4-Pro output goes from $0.87 to $1.98 off-peak and $3.96 at peak per million tokens.

DeepSeek ends its flat, ultra-cheap API rates and moves the V4 models to time-of-day pricing.
Quick facts
| Maker | DeepSeek |
|---|---|
| Effective | 16:00 UTC, August 16, 2026 |
| Peak hours (UTC) | 01:00–04:00 and 06:00–10:00 |
| Models affected | DeepSeek-V4-Pro, DeepSeek-V4-Flash |
| Peak vs off-peak | Peak is 2x off-peak |
| Reported increase | 50%–1,100% (Reuters) |
Pricing
| V4-Pro input (cache miss), off-peak | $0.66 / 1M tokens |
|---|---|
| V4-Pro output, off-peak | $1.98 / 1M tokens |
| V4-Pro output, peak | $3.96 / 1M tokens |
| V4-Flash input (cache miss), off-peak | $0.22 / 1M tokens |
| V4-Flash output, off-peak | $0.66 / 1M tokens |
| V4-Flash output, peak | $1.32 / 1M tokens |
What is it?
DeepSeek is replacing a single flat API rate with two price bands, peak and off-peak, for DeepSeek-V4-Pro and DeepSeek-V4-Flash. The new rates start at 16:00 UTC on August 16, 2026. Reuters puts the increase at 50% to 1,100% above current prices, depending on the model, token type, and hour of the day.
How does it work?
The change splits each day into two peak windows, with off-peak covering everything else, and sets peak rates at exactly double off-peak across both V4 models. Each band prices three token types separately — cached input, uncached input, and output — so a cache-heavy workload and an output-heavy one see very different increases.
Why does it matter?
Cheap inference was DeepSeek's main draw, and a V4 bill now depends on when a job runs. Batch work, evals, and background agent runs can shift out of the peak windows and pay half. Anything that must answer live during peak hours pays the full rate, so cost planning now needs a clock, not just a token count.
Who is it for?
teams running high-volume DeepSeek API workloads
Frequently asked questions
- How can I avoid DeepSeek's peak pricing?
- DeepSeek prices peak hours at exactly double the off-peak rate, so moving work outside the two peak windows halves the bill. Batch jobs, evals, and background agent runs are the easiest to shift. Requests that must answer live during peak hours pay the full rate, so latency-sensitive apps have less room to save.
- Which DeepSeek token type rises the most?
- Cached input tokens see the largest jump on DeepSeek-V4-Pro: $0.003625 per million today becomes $0.022 off-peak and $0.044 at peak, roughly twelve times the old rate. Uncached input rises from $0.435 to $0.66 off-peak, and output from $0.87 to $1.98 off-peak. Cache-heavy workloads are hit hardest.
- Does DeepSeek-V4-Flash get more expensive too?
- DeepSeek-V4-Flash rises alongside the Pro model. Uncached input goes from $0.14 to $0.22 off-peak and $0.44 at peak per million tokens, while output moves from $0.28 to $0.66 off-peak and $1.32 at peak. Flash stays the cheaper of the two DeepSeek V4 models after the change.
- Why did DeepSeek raise its API prices?
- DeepSeek announced the increase in the same post that took DeepSeek-V4-Pro to general availability on August 13, 2026. The company frames the split as a way to schedule flexible workloads, noting off-peak rates are 50% lower than peak. Reuters reports the new rates run 50% to 1,100% above current prices.
Try it
https://api-docs.deepseek.com/quick_start/pricing