All writing

A Billion Tokens for $6.36

The receipt

The receipt and the warning

A dashboard from this week. Cost: $6.36 USD. API requests: 9,147. Tokens: 1,015,102,000.

Cost
$6.36
API requests
9,147
Tokens
1,015,102,000

That is $6.27 per billion, $0.0007 per request, and about 111,000 tokens per request, the signature of an agent that re-reads a large context on every turn and does it all day.

Three days ago we ran “Same Score, 1/36th the Bill.” It quotes a page that now carries a banner:

We plan to raise the overall pricing for DeepSeek API services in the near future, with a significant increase expected. Please plan your usage accordingly. The specific pricing plan will be subject to official notice.

No number. No date. No reason.

So screenshot your receipt. The number on it was never the price of a token, and your dashboard will not tell you what it was the price of.

The arc

One cut, one retraction, one warning

DeepSeek V4 Pro launched on April 24, 2026 at $1.74 per million input tokens and $3.48 per million output.

On May 22 the company cut all three of its billing lines by 75%, to $0.435 input and $0.87 output, and said the cut was there to stay. The Hacker News thread ran to 621 points. Sanchit Vir Gogia of Greyhound Research, May 25: “This is why the price cut is permanent rather than promotional. It is not a discount. It is an efficiency gain being passed through.”

Seventy-three days later, the same company warned of a significant increase.

In between came a peak-hour surcharge, announced June 30: 2x on all V4 models, 09:00–12:00 and 14:00–18:00 Beijing time.

The efficiency gain being passed through lasted about ten weeks.

The swap

The fuse got swapped for a bigger one

Our August 3 piece quoted the surcharge off the pricing page: the API “will soon adopt a peak/off-peak pricing policy. During peak hours, prices will be 2x the regular prices, applicable to all billing items,” with the effective date “subject to the official announcement.”

The pricing page today: that language is gone. The rates are flat. The undated warning sits where the surcharge used to.

DeepSeek API pricing per 1M tokens · read from the pricing page, August 6, 2026
ModelCache hitCache missOutputConcurrency
deepseek-v4-flash$0.0028$0.14$0.282,500
deepseek-v4-pro$0.003625$0.435$0.87500

We were wrong about the fuse, and not in the direction that flatters anyone. A 2x surcharge on published hours is a thing you can model. You cannot model “significant.” It is an adjective.

A company that knows its new prices publishes them. A company still deciding tells you to plan your usage accordingly.

The real exposure

You do not have a token bill, you have a cache-hit rate

At V4 Flash's cache-miss rate, 1,015,102,000 tokens cost $142.11. At the output rate, $284.23. At the cache-hit rate, $2.84.

The actual bill was $6.36. Three price tiers and two unknowns do not solve uniquely, but they box the answer in hard: output tokens cannot be more than about 2% of that total, and roughly 97.5% of the traffic has to be cache hits.

Every token a cache hit
$2.84
the floor, at $0.0028 per million
What the dashboard actually billed
$6.36
which puts cache hits at roughly 97.5% of traffic
Every token a cache miss
$142.11
same tokens, same sticker price, a 50x line item

That receipt measures how well one harness caches, not how cheap DeepSeek is. The same billion tokens through a harness that reuses nothing bills at $142, and the sticker price never moved.

On Flash, a cache miss costs 50x what a cache hit costs. On Pro, 120x. One commenter on the May thread put DeepSeek's cache reads “on the scale of 100x cheaper” than the next cheapest provider. That one line item is what made agent economics work.

The cause

A chip story wearing a pricing costume

Not greed, and not the end of the subsidy era. Something more permanent: DeepSeek cannot get hardware.

On July 25 a leaked investor transcript surfaced. Founder Liang Wenfeng: “the biggest gap between us and the United States lies in resources.” As reported, he wants on the order of two hundred thousand cards and has roughly sixteen thousand in hand, called Huawei's capacity insufficient, and put the crunch at three years minimum. The transcript's own card arithmetic is ambiguous and it reached the public secondhand, so treat the figures as directional. Twelve days later came the notice.

The Council on Foreign Relations reported on April 29 that “DeepSeek itself admits that it currently cannot serve its v4 pro model to most customers because it lacks the chips,” and that the shortages “render the pricing moot for the time being.” That was launch week, so it describes the constraint rather than this month's decision.

The concurrency limits on the pricing page today: 2,500 for Flash, 500 for Pro. That is not a pricing decision. That is a capacity disclosure hiding in a table.

The hardware market underneath is moving too. NVIDIA raised the DGX Spark from $3,999 to $4,699 mid-cycle. The March 2026 thread put it down to RAM “likely being reflected in the price”; NVIDIA's own explanation, given separately, was memory supply constraints. Apple stopped producing the 512GB Mac Studio on March 5 and the 256GB on May 5. But when ADATA's chairman warned in July that the shortage would last another decade, the top comment was “Warns. CEO of company that needs AI to keep going.”

Memory is a real input. It is not the whole explanation, and anyone selling it as one is selling something.

The other side

The strongest case that nothing is ending

A week before DeepSeek's notice, on July 30, OpenAI cut GPT-5.6 Luna by 80% ($1.00/$6.00 down to $0.20/$1.20).

One comment in the 608-point thread: “After a year of ever-increasing prices it suddenly feels (between this, Kimi K3, GLM 5.2) that prices are falling again.”

The margin argument is better still. One Hacker News commenter: “third-party inference providers serve DeepSeek V4 Flash just as cheaply as DeepSeek themselves, if not even more so. This is very strong evidence that the low price of the model is not subsidized.”

If serving Flash at $0.14 is profitable engineering rather than a loss leader, then a large increase on Flash is competitive suicide. Artificial Analysis scores DeepSeek V4-Flash-0731 at 50 and GPT-5.6 Luna at 51; running its full index cost $72.02 on Flash against $174.06 on Luna. A moat of roughly 2.4x at parity intelligence does not survive a “significant” hike. Developer Michael Guo, quoted by SCMP: “DeepSeek choosing to raise prices at this time — isn't this just asking for trouble?”

Nobody is capacity-constrained on Flash, so nobody can raise it much. Pro is the model DeepSeek admits it cannot serve, and scarce things do not get cheaper.

Anthropic's Sonnet 5 reverts from $2/$10 to $3/$15 on September 1, 2026. Grok went from $1.25/$2.50 to $2.00/$6.00 between 4.3 and 4.5, and GLM from $1.0/$3.2 to $1.4/$4.4 between 5 and 5.2. The floor is moving in both directions at once, which is what a market looks like when it stops being a land grab.

The exit

Flash has an exit, Pro does not

V4 Flash has open weights, so the market routes around DeepSeek. On OpenRouter today, deepseek-v4-flash serves at $0.0882/$0.1764, 37% below DeepSeek's own list price.

Xiaomi's MiMo-V2.5 sits at $0.14/$0.28, an exact clone of DeepSeek's sheet. Qwen3.7-Flash undercuts everyone at $0.03/$0.13. If DeepSeek raises Flash pricing, you change a base URL.

V4 Pro is the trap. On OpenRouter it prices at exactly $0.435/$0.87, and nobody undercuts it. In the 621-point thread, Baidu came in at $1.521/$3.042 and Novita at $1.64/$3.38, three to four times DeepSeek's own rate for the same weights. The one you cannot leave is the one you cannot fully get.

  • Flash: change a base URLOpen weights, and OpenRouter already serves it at $0.0882/$0.1764 — 37% below DeepSeek's own list price
  • Pro: no exitNobody undercuts $0.435/$0.87. Baidu asks $1.521/$3.042 and Novita $1.64/$3.38 for the same weights
  • Self-host: control, not savingsA $6,406 box pays back in two years against Qwen prices and past fifteen against DeepSeek's

One developer ran 150 real coding tasks across cloud and a local 3090, cutting his bill from $85/month to $22 by routing 65% of work locally.

A $6,406 server on a real business workload, 20.4 million input and 1.32 million output tokens a day, burns $770 of electricity a year against $3,701 of API-equivalent tokens. Two-year payback. That equivalence is priced against Qwen3.6 27B at $0.29 input and $3.20 output. Run the identical workload through DeepSeek Flash's sheet and it costs about $1,177 a year. The payback on that box stretches past fifteen years.

Self-hosting does not lose on its merits. It loses because DeepSeek is cheap. Which means the price increase is what brings it back, and every “just self-host” thread from the last six months was written in the shadow of a price that is now moving.

As the top reply to one of those threads put it: “self hosting means you have full control over pricing. Prices can change.”

What to do

Plan your usage accordingly

Instrument your cache-hit rate this week. It is the only number on your invoice with 50x of leverage behind it, and nobody has a graph for it.

Then sort your workloads: Flash, where the exit is a base URL, and Pro, where there is none.

CNBC reported on July 7 that Chinese models hit a weekly peak of 46% of US enterprise token usage on OpenRouter, with DeepSeek the single largest vendor at 17.6%. Forbes reported on July 28 that Telnyx moved from $100,000 a day on Anthropic to roughly $100 per agent per day on Chinese models. Those two figures are denominated differently and are not a like-for-like comparison, but the direction is not subtle.

A very large amount of Western software has been priced against a Chinese company's spare capacity. That company has now said, in writing, that the spare capacity is over, and its founder was telling investors twelve days earlier that he cannot buy the chips to make more.

So print the receipt. Frame it. Engineers are going to talk about mid-2026 the way they talk about dollar-a-gigabyte storage or unmetered bandwidth: a brief, structurally weird window when the meter was off and one developer could burn a billion tokens for the price of a sandwich without ever checking what it cost.

The window did not close today. But somebody just put a hand on the blind.