DeepSeek API Pricing: What You Actually Pay Per Run

Most DeepSeek pricing pages quote a single input number. The official docs quote six, split by peak hour and cache hit. Here is what the rate card actually says, why the aggregators disagree with it, and why per-token price is the wrong thing to optimise anyway.

Jason Bao
deepseek-api-pricing-hero.png

Key takeaways

  • DeepSeek publishes six prices per model, not one: cache-hit input, cache-miss input, and output, each in two tiers.
  • Peak hours are 01:00-04:00 and 06:00-10:00 UTC, Monday to Friday. Off-peak rates are half of peak.
  • A cache hit on v4-flash costs about one thirty-first of a cache miss. It is the biggest lever on the card.
  • Figures below were checked on 7 September 2026. DeepSeek's pricing page carries no date stamp of its own.
  • Cost per completed task is the number worth tracking. Retries and re-read context multiply whatever rate you pay.

The rate card, as of 7 September 2026

All figures checked on 7 September 2026 against DeepSeek's official Models & Pricing docs, which is the only source used for DeepSeek prices in this article. Prices are USD per 1M tokens.

| Model | Tier | Input, cache hit | Input, cache miss | Output | |---|---|---|---|---| | deepseek-v4-flash | Peak | $0.014 | $0.44 | $1.32 | | deepseek-v4-flash | Off-peak | $0.007 | $0.22 | $0.66 | | deepseek-v4-pro | Peak | $0.044 | $1.32 | $3.96 | | deepseek-v4-pro | Off-peak | $0.022 | $0.66 | $1.98 | | deepseek-v4-flash-vision-exp | Peak | $0.014 | $0.44 | $1.32 | | deepseek-v4-flash-vision-exp | Off-peak | $0.007 | $0.22 | $0.66 |

All three models carry a 1M-token context window. Billing is stated simply on the same page: "The expense = number of tokens × price," deducted from your topped-up or granted balance, granted balance first.

Two caveats sit on that page and both matter more than they look. DeepSeek states that "Product prices may vary and DeepSeek reserves the right to adjust them," and recommends checking the page regularly. And the page has no publish date, no last-modified date, and no changelog, so there is no way to tell from the page itself when a figure last moved. That is why this article states its own check date, and why you should confirm against the source before putting any of these numbers in a budget.

On the reported price increase: outlets covered a significant DeepSeek API increase in August 2026, with Yahoo Finance reporting rises of as much as roughly 1,100% depending on model, token type, and peak status, and the SCMP reporting an earlier peak-hour surcharge. The official docs publish the current rates and nothing else. No effective date, no prior card to compare against. Treat the coverage as coverage and the table above as the figure you can verify.

How much does the DeepSeek API cost?

There is no single answer, and any page giving you one is wrong. On the official card, v4-flash runs $0.007 to $0.44 per million input tokens and $0.66 to $1.32 output. v4-pro runs $0.022 to $1.32 input and $1.98 to $3.96 output. Your effective rate depends on when the job runs and how much of the prompt was cached.

Why there is no single DeepSeek price

The card is four-dimensional: model, tier, cache state, and direction. Multiply those out and one model has six prices. Most cost models collapse this to "input price and output price," which is where the estimates start drifting from the invoice.

Peak versus off-peak

Off-peak rates are half of peak, and peak is a narrow, fixed window: 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. Everything else is off-peak, including all weekend.

For anything you control the schedule of, this is free money you have to remember to take. A nightly enrichment job, a batch classification pass, a weekly report build: put them outside those seven weekday hours and the same tokens cost half as much. For interactive traffic you do not control the clock, so model at peak rates and treat the off-peak share as upside rather than baseline. A team in London on business hours is buying peak tokens through the first part of every morning without ever having chosen to.

Cache hit versus cache miss

This is the largest multiplier on the card and the one that most cost spreadsheets ignore entirely. On v4-flash, off-peak cache-miss input is $0.22 per million and cache-hit input is $0.007. That is roughly a 31x difference on the same tokens. The ratio on v4-pro is about 30x.

What it rewards is prompt stability. If your system prompt, tool definitions, and retrieved context sit in the same order at the front of every request, the shared prefix is cheap on every call after the first. Reorder it per request, interpolate a timestamp into the top of the prompt, or shuffle retrieved documents, and you pay cache-miss rates on the whole thing forever. Prompt construction has become a billing decision.

Why third-party pricing pages disagree

Because the four dimensions give them six numbers to get wrong. Checked on 7 September 2026, deepseek.ai/pricing publishes v4-flash at $0.14 input and $0.28 output, and v4-pro at $0.435 and $0.87, under the claim that its figures are "independently verified against DeepSeek's official pricing page." None of those four match any figure on the official card. costgoat.com has since corrected its main table to the official rates, though its per-model cards still show input figures its own guide contradicts.

The practical rule: source DeepSeek prices from DeepSeek, and stamp the date you checked.

What it costs at your volume

Three scenarios. Every input number below is an assumption I am labelling as an assumption, not a benchmark. Swap in your own and the arithmetic still holds.

A low-volume internal feature. 2,000 calls a month, 3,000 input tokens per call with no cache reuse, 500 output tokens, v4-flash, scheduled off-peak. Input is 6M tokens at $0.22, so $1.32. Output is 1M tokens at $0.66, so $0.66. Total: $1.98 a month, $3.96 at peak. At this volume, pick on quality and stop reading pricing pages.

High-volume classification. 5M documents a month, 800 input tokens each of which 600 is a stable shared prefix, 20 output tokens, v4-flash, off-peak. Cache-hit input is 3,000M tokens at $0.007, or $21.00. Cache-miss input is 1,000M at $0.22, or $220.00. Output is 100M at $0.66, or $66.00. Total: $307.00 a month. Run the identical workload with an unstable prompt so nothing caches and the input line becomes 4,000M at $0.22, taking the bill to $946.00. Same tokens, same model, three times the cost, decided by prompt layout.

An agent loop that re-reads its context. 6,000 runs a month, six model calls per run, a 40,000-token context re-read on every call, 700 output tokens per call, v4-pro, off-peak, with the first call a cache miss and the other five hits. Cache-miss input is 240M at $0.66, or $158.40. Cache-hit input is 1,200M at $0.022, or $26.40. Output is 25.2M at $1.98, or $49.90. Total: $234.70 a month, which is $0.0391 per run.

Now the number that actually matters. That $0.0391 is cost per attempt, not cost per completed task. Assume 80% of runs land first time and the rest need one full retry: 6,000 completed tasks consume 7,200 attempts, so the bill is $281.64 and the cost per completed task is $0.047. Twenty percent more spend, from a factor that appears nowhere on the rate card.

That is the whole arithmetic of LLM spend, and it is worth being precise about the levers that actually move an LLM bill: tokens per attempt, rate per token, attempts per completed task. A model swap moves one of the three.

One more constraint that outranks price above a certain volume. Concurrency limits are per-model and account-level: 2,500 for v4-flash and the vision variant, 500 for v4-pro, with over-capacity requests returning HTTP 429, per DeepSeek's rate limit docs. If 500 concurrent requests is your ceiling, the cheapest per-token rate in the market does not help you.

DeepSeek versus the frontier rate cards

Every figure sourced to the vendor's own page, all per 1M tokens, all checked 7 September 2026. DeepSeek rows use off-peak, cache-miss input.

| Provider | Model | Input per 1M | Output per 1M | Source | |---|---|---|---|---| | DeepSeek | deepseek-v4-flash | $0.22 | $0.66 | api-docs.deepseek.com | | DeepSeek | deepseek-v4-pro | $0.66 | $1.98 | api-docs.deepseek.com | | Google | Gemini 3.8 Flash (standard) | $0.75 | $3.75 | ai.google.dev | | Anthropic | Sonnet 5 | $2.00 | $10.00 | claude.com/pricing | | OpenAI | gpt-5.6-terra | $2.00 | $12.00 | developers.openai.com |

The gaps are real, and narrower than the headlines suggest once you put v4-pro at peak against a mid-tier frontier model. A rate card comparison also tells you nothing about attempts per completed task, which is why comparing models on more than price and reading the cost gap honestly do more for a bill than one cheap default. Send the easy 90% of a workload to a small model, reserve the expensive one for hard cases: routing cheaper work to cheaper models beats any single rate card.

Where per-token pricing stops being the right question

Look again at the third scenario. Of the $234.70, the output tokens are $49.90 and the re-read context is $184.80. Most of that bill is the agent paying, over and over, to be told things it already established.

That pattern is what makes agent spend climb with usage. The same context gets re-read on every call. The same classification gets re-derived on the next run. The same routing decision gets re-made from scratch because nothing wrote it down. Switching providers changes the price of a unit of reasoning. It does not change how many units you buy, and in an agent loop the number of units is the term that grows.

So the honest reading of a cheap rate card is that it lowers the coefficient and leaves the curve alone. If your workload is a single high-quality inference per user action, a cheaper model is a straightforward win and you should take it. If your workload is an agent doing multi-step work on a schedule, the rate card is the smaller half of the problem and coordinating models and the app layer is the larger half.

What this article doesn't cover

Fine-tuning and batch-mode pricing, non-USD billing, reseller and inference-provider rates for open DeepSeek weights, and latency under load. The last one deserves its own treatment: at 500 concurrent requests on v4-pro, queueing behaviour will shape your architecture more than the price per million.

The Major take

The constraint is arithmetic, and no provider can price its way out of it for you. A workflow that re-reads the same context, re-derives the same classification, and re-decides the same routing on every run buys the same units of reasoning every time it executes. A cheaper card lowers the invoice. Spend still climbs with usage, because the shape of the curve is set by how many times the model is asked to work something out, not by what each answer costs. That is the ceiling on model substitution as a cost strategy, and it holds whoever wins the price war.

Major moves the other term. When an agent on Major works out how to handle a repeatable part of a task, it builds an app for that part and runs the app instead of reasoning through the step again, so that step stops appearing on any rate card at all. The app carries its own database, storage, and logs, which means the context the model used to re-read on every run persists in the app rather than being re-bought each time. The economics are front-loaded and then flat: figuring out the workflow costs real reasoning once, and repeated execution after that does not scale token spend with usage. The model keeps the judgment calls and keeps building the apps. It stops re-deriving the parts that never vary. Reason once, run forever, and the cheapest token turns out to be the one you never spend.

If you costed out that agent loop above and did not like the $184.80 of re-read context, the fix sits one layer above the rate card. Have the agent build an app for the deterministic steps, hold its state there, and keep the model for the judgment. Get started on Major and build an agent that turns its repeatable steps into apps instead of re-buying them every run.

Related articles

Frequently asked questions

How much does the DeepSeek API cost?
DeepSeek bills per million tokens across six prices per model, not one. Checked on 7 September 2026, v4-flash runs $0.007 to $0.44 per million input tokens and $0.66 to $1.32 output. v4-pro runs $0.022 to $1.32 input and $1.98 to $3.96 output. The spread comes from cache state and from peak versus off-peak hours.
How much is 1M tokens in DeepSeek?
For deepseek-v4-flash, one million input tokens costs $0.22 off-peak or $0.44 at peak on a cache miss, and $0.007 or $0.014 on a cache hit. Output is $0.66 off-peak, $1.32 at peak. v4-pro is three times those figures. Peak is 01:00-04:00 and 06:00-10:00 UTC on weekdays.
Why is the DeepSeek API cheap?
Three mechanics, all on the rate card. Context caching prices a repeated prompt prefix at roughly a thirtieth of a fresh one. Off-peak rates are half of peak, and peak is only seven weekday hours, so most of the week bills at the lower tier. And the flash-tier model is small enough to serve at high concurrency, which spreads serving cost across more requests.
Is DeepSeek free and unlimited?
No. The consumer chat product is free to use, but the API is paid: charges are deducted from a topped-up or granted balance as tokens are consumed. It is also not unlimited. Concurrency is capped per model at the account level, 2,500 for v4-flash and 500 for v4-pro, with over-capacity requests returning HTTP 429.
Is a cheaper model actually cheaper to run?
Only if it needs the same number of attempts. What a workload costs is tokens per attempt, times rate per token, times attempts per completed task. A model that is five times cheaper but needs two tries, longer prompts, or a fresh context read on every call can cost more in production than the one it replaced. Track cost per completed task.