The receipt
Three days' notice
On August 7, a Hacker News commenter pasted an itemized DeepSeek invoice into a thread about V4-Flash: 1,265,646,976 cache-hit input tokens, 18,208,088 cache-miss input tokens, 9,615,178 output tokens across 10,837 requests. Total cost, $8.79. A 98.582% cache hit rate.
Today, August 13, 2026, DeepSeek replaced the price table that invoice was billed under. Reprice the identical workload against the new one: $19.21 off-peak, $38.42 at peak. Same tokens, same cache hits, 2.19x the bill, or 4.37x if you're unlucky about the hour. It takes effect at 16:00 UTC on Sunday, August 16.
- Billed August 7
- $8.79
- 1.28B input tokens at a 98.582% cache hit rate
- Same job, off-peak
- $19.21
- 2.19x, for seventeen hours of the day
- Same job, peak
- $38.42
- 4.37x, for the other seven
A week of warning preceded it. On August 6 the pricing page swapped its footnote for a notice with the figures left out: a significant increase was coming, plan your usage accordingly. There is a date now, and it is three days out.
That is the headline. What it buries: for the first time since DeepSeek started selling tokens, buying your own hardware stopped being obviously stupid, and GPU prices had nothing to do with it.
The table
The cache line took the worst of it
Per million tokens, DeepSeek's current table against Sunday's:
V4-Flash gets off lightly, if you squint. Cache-miss input: $0.14 to $0.22 off-peak, 1.57x. Output: $0.28 to $0.66, 2.36x. Cache-hit input: $0.0028 to $0.007, 2.5x on a number small enough to still look like a rounding error.
For V4-Pro, the pattern is violent. Cache-miss input rises 1.52x. Output rises 2.28x. Cache-hit input goes from $0.003625 to $0.022, a 6.07x rise off-peak and 12.14x at peak. DeepSeek raised the price of remembering far more than the price of thinking.
That invoice is the shape of every coding agent: an enormous stable prefix (system prompt, tool definitions, file context) replayed on every turn, with a few hundred fresh tokens on the end.
DeepSeek's cache-hit discount was the most aggressive in the industry. On V4-Pro, a cache read cost one-hundred-twentieth of a cache miss. On Sunday it becomes one-thirtieth. It just stopped being the thing that made the arithmetic absurd.
The July 31 version of the docs, still in the Internet Archive, promised something gentler: peak hours at 2x regular prices, implying off-peak stayed where it was. What shipped raises off-peak too, by between 1.5x and 6x depending on the line. There is now no hour of any day at which you pay what you paid last week.
The clock
Peak hours are somebody else's business day
Peak is defined as 01:00–04:00 and 06:00–10:00 UTC. Seven hours out of twenty-four.
In Beijing, peak runs 09:00–12:00 and 14:00–18:00. That is the Chinese working day, minus lunch, drawn with surgical precision. In San Francisco it lands at 18:00–21:00 and 23:00–03:00.
Europe wakes up inside it. In Bucharest the second peak window runs 09:00–13:00, so our entire working morning sits in the expensive column. London's lands at 07:00–11:00, Bengaluru's at 11:30–15:30.
Your price increase depends on your longitude. That is a strange sentence to write about software.
Run around the clock (a fleet, an always-on reviewer) and you pay a blend: seven peak hours and seventeen off-peak works out to 1.29x the off-peak rate, or about 2.82x what you pay today.
The escape hatch
Buy a GPU, they said
DeepSeek-V4-Flash-0731 is downloadable: 48 safetensors shards, 166.9 GB on disk, FP8 block quantization, MIT licensed, roughly 304B total parameters, 1M context. The practitioner who benchmarked it in early August measured a 204.5 GB high-water mark against a single MI300X's 192 GiB. His numbers on vLLM and ROCm: 168.6 tok/s single-stream, 542 tok/s aggregate across 8 streams, 830 tok/s across 64.
That card rents for $2.39/hour on RunPod as of this week. At 500 tok/s sustained you produce 1.8 million output tokens an hour, worth about fifty cents at DeepSeek's current $0.28 list price. You are paying $2.39 to manufacture $0.50 of product. Breaking even on output alone at that rental rate needs about 2,370 tok/s, roughly triple what the benchmark sustained. A commenter in the thread eyeballed the same gap and put the bar at “~1500 tps.”
and at that point, you will have dinosaur machines to attempt to run 2032’s AI lineup. Which will probably just run on your phone anyways.
That 327-upvote comment from August 10 priced a $10,000 DGX Spark cluster at a six-year payback running flat out. Dax Raad of OpenCode ran it for a single interactive user that same week ($1.14/day of Flash tokens) and got a dual-DGX payback of 24 years. Someone who built a five-GPU workstation for GLM in July put his at “over 10 years to break even at today's prices,” having spent, by his own blackout-inducing estimate, something touching $80,000.
A May 2026 writeup of a $6,406 four-MI100 server serving 20.4M input and 1.32M output tokens a day put API-equivalent pricing at $3,701/year against $2,993 to run it locally in year one, including power and depreciation. First-year saving: $708.
The arithmetic
Where the break-even actually sits
On August 4, a commenter modelled a cache-heavy agentic workload (2M input, 1M output, 98.5M cached tokens) priced with his own assumptions of 8,000 tok/s prefill and 800 tok/s decode on a $2/hour card. He computed $0.8358 on DeepSeek's API against $0.8333 self-hosted, a dead heat to within a quarter of a cent. Several people replied that the margin was “within shooting distance of ‘at cost’”.
Repriced against Sunday's table, the API side becomes $1.7895 off-peak and $3.5790 at peak. The self-hosted side does not move, because electricity and rental rates are not indexed to DeepSeek's pricing page. Self-hosting goes from break-even to 2.15x cheaper off-peak and 4.29x cheaper at peak.
The real $8.79 invoice reduces to one number: required utilization. Producing that workload's 9.6M output tokens at the measured 542 tok/s, plus prefilling the 18.2M cache-miss tokens at the measured 7,000 tok/s, takes about 5.65 GPU-hours: $13.50 at $2.39/hour, assuming the box does nothing else, ever.
- Today: never$13.50 of GPU time against an $8.79 bill. The curve does not cross at any duty cycle
- Off-peak: 70%Seventeen hours of every twenty-four, generating tokens, to beat $19.21
- Peak: 35%Against $38.42, a box busy a third of the time already wins
Against today's $8.79 bill, that is a losing trade at any utilization: the API price already sits under your hardware cost at a 100% duty cycle. There is no break-even. The curve never crosses.
Against Sunday's $19.21 off-peak bill, you need about 70% sustained utilization. Against the $38.42 peak bill, 35%.
Seventy percent is demanding: your card generating tokens for seventeen hours of every twenty-four. But it is a number, and it is on the map. Last week it wasn't.
The moat
The line item you can't rent
The break-even moved while GPU prices sat still, because the thing DeepSeek repriced is the one thing self-hosting gives away free.
When you own the box, a cache hit costs you memory, not money. Keeping a full 1M-token context resident runs about 9.62 GiB, roughly 5% of a 192 GB card, or about twelve cents an hour of amortized rent. The commenter benchmarking two Sparks noted it in passing: cached input tokens on local inference are free.
Google bills Gemini cache storage by the hour: $1.00 per million tokens per hour on Flash tiers, $4.50 on Gemini 3.1 Pro. Holding a million-token context warm for a working day on Pro costs $36 in storage alone, before a single token is generated. Anthropic charges no storage fee but prices the write: a 1-hour cache write costs 2x base input, reads 0.1x. Every serious vendor meters the cache somewhere. DeepSeek was the one that didn't.
The docs are explicit: caching is on by default for everyone, there is no TTL to configure, and a cache is cleared “usually within a few hours to a few days” once idle. One commenter reported hitting a cache after more than 24 hours, against the five minutes most providers give you, and said so directly: there is no way to match DeepSeek's prices profitably if you're renting a GPU and reselling tokens.
That was the moat, and it was never the per-token price. Sunday's table puts a number on it.
The hostile read
The strongest case for doing nothing
Everything above is the bull case, and it deserves a hostile reading.
Utilization is where these calculations go to die. A 42-upvote reply demolishing a “34x first-year ROI” self-hosting post put it plainly: you cannot take aggregate tokens per second and divide by per-user tokens per second. With a 1M context you fit fewer than 100 concurrent users on the box, and when one finishes you have to swap that context out and load another. Real traffic is bursty, and an idle GPU bills exactly the same as a busy one.
The throughput numbers themselves are all over the place. The August MI300X benchmark reported 830 tok/s at high concurrency. A different practitioner, on the same class of hardware in June 2026, reported 2,699 tok/s after kernel work, while openly admitting he hadn't proven a win on tokens per second per dollar. That is a 3.25x swing, and every break-even I quoted above inherits it.
Cache hit rate is equally load-bearing and equally variable. A June 2026 paper measured prefix hit rates by workload: 90.7% for multi-turn conversation, 81.8% for QA, 6.3% for summarization, 0.1% for code completion. The 98.5% figure everyone is reasoning from is the best case. One commenter reported ~99% on a minimalist tool against ~79% through OpenRouter, because subagents and mid-context injection bust the prefix.
V4-Pro-0813 hit Hacker News on August 12 as an API product; its weights return a 401 on Hugging Face. Self-hosting the best DeepSeek means running the April Pro checkpoint at 864.7 GB, which doesn't fit on 8x H100 and needs 8x H200 at roughly $28.70/hour. For frontier quality there is no build option to price.
The steelman came from inside r/LocalLLaMA on August 10, 89 upvotes, and it says none of this is close: “unless you have a use case to run tokens 24/7 or need the privacy, there's no point in ever running a non-finetuned local model from a financial perspective.” It was written before Sunday's numbers existed. It is still mostly right. “Mostly” is the part that changed.
The playbook
What to do before Sunday
You have until 16:00 UTC Sunday. Three things, in order of how much money they save.
Measure your cache hit rate today, from your own billing data, not from a blog. It is the single number that decides your exposure: after Sunday the same pile of cache hits costs 2.5x more on Flash and 6x more on Pro. If you're running a coding agent through a harness that injects context mid-prompt, you're busting the prefix and paying cache-miss rates for the privilege.
Move anything that doesn't need a human waiting to off-peak. Peak is only seven hours, and in the Americas it falls in the evening and overnight. Batch evaluation and nightly reviews can run at 2.19x instead of 4.37x for the cost of a cron expression.
Then, and only then, price hardware. Write down your real duty cycle, not your best hour. If your GPU would sit idle two-thirds of the time, the API still wins. If you run tokens around the clock, at a hit rate above 90%, on Flash rather than Pro, the arithmetic now supports a conversation it couldn't a week ago.
For two years, “just call the API” was the correct answer to nearly every question. At 16:00 UTC on August 16, it stops being automatic. The answer becomes “it depends on your utilization,” which is a worse answer, and a much more interesting one.
The free cache was the best deal in software. Sunday is when we find out how many people had built a business on it.