The correction
Nobody shipped a model today
DeepSeek published benchmark scores on July 31, 2026 showing a 645% jump on a model whose weights have not changed since April. That is the entire event. DeepSeek-V4-Flash-0731 keeps, in the company's own words, “the same model architecture and size as DeepSeek-V4-Flash-Preview, and was only re-post-trained.” Twenty-four hours earlier, OpenAI had cut GPT-5.6 Luna by 80%. DeepSeek's answer to that cut was to ship no model at all.
Same parameter count. Same serving cost. Same price sheet, down to the fourth decimal place. What moved was the scoreboard: Terminal Bench 2.1 from 56.9 to 82.7, Toolathlon-Verified from 51.8 to 70.3, and DeepSWE — the one that should make you put your coffee down — from 7.3 to 54.4.
If those numbers survive independent replication, this is the most important thing that happened in AI this month, and it isn't close. Cheap models get good all the time. What makes this one matter is where the headroom came from. For three years the industry's answer to “how do we make it better” was “make it bigger, and charge more.” DeepSeek just published a counterexample where the answer was “post-train it again, and charge nothing extra.”
That's the press release. What it leaves out is that almost every naive way to act on these numbers is wrong: a verbosity tax that eats a chunk of the price advantage, a 2x surcharge sitting on the pricing page with no fire date, and a version string that quietly means something different today than it meant yesterday. Every number in this piece is a July 31, 2026 snapshot, and this month has been an object lesson in how long those last.
The price sheet
The cache line is the whole price war
The sticker prices are the part everyone screenshots and the part that explains the least.
DeepSeek's official pricing page lists V4-Flash at $0.14 per million input tokens on a cache miss, $0.28 per million output, with a 1M context window and 384K max output. OpenAI's freshly-cut Luna, after Thursday's 80% haircut, sits at $0.20 input and $1.20 output. On output tokens, the ones agentic workloads actually burn, DeepSeek is 4.3x cheaper than a price that was itself cut 80% the previous day.
- Output tokens
- 4.3x
- $0.28 against Luna's $1.20, one day after Luna was cut 80%
- Cache-hit input
- 7x
- $0.0028 against $0.02, on the line item agents actually burn
- DeepSWE, frozen weights
- +645%
- 7.3 to 54.4, self-reported on an unreleased harness
Widen the frame and it gets silly. At $0.28 per million output tokens, V4-Flash is roughly 89x cheaper than Claude Opus 4.8 and 179x cheaper than Fable 5 on output. Those are not the same product, and the frontier models win the evals they're built for. They are also increasingly pointed at the same ticket.
The number that actually matters, though, is the one nobody puts on a slide: cache-hit input at $0.0028 per million, against Luna's $0.02. That's a 7x gap on the single line item that dominates any real agent's bill. Agents re-read the same context on every turn. If your harness caches well, the cache line is your invoice, and a 7x gap there swamps a 4.3x gap on output.
Artificial Analysis puts the blended rate at about $0.06 per million, first on price out of 162 models. Check that number against the sticker prices before you quote it: no mix of $0.14 input and $0.28 output produces $0.06, because the floor of any such blend is $0.14. The figure only reconciles if the blend assumes cache-hit input throughout, which makes it a best-case number rather than a list price.
DeepSeek answered OpenAI's price cut without touching its own prices. It made the same cheap model materially smarter and let the existing price do the talking. That's a flex, and it's a more expensive flex to answer than a discount.
The asterisk
Every one of those numbers is self-reported
Now the part where we spoil our own lede. Every one of those benchmark numbers is self-reported, run on DeepSeek's own agentic harness, which has not been released. The official Terminal-Bench 2.1 leaderboard does not list V4-Flash-0731 at all. Its top entries as of today are Claude Code with Fable 5 at 83.8%, Codex with GPT-5.5 at 83.1%, and Codex with GPT-5.6 Terra at 78.4%. DeepSeek's claimed 82.7 would land near the top of that board if an independent harness reproduces it. “If” is carrying the entire sentence.
A 645% improvement on DeepSWE from post-training alone has two innocent explanations and one guilty one. A commenter on the Hacker News launch thread put the innocent pair on the table: “The either released a severely undercooked model or they made insane improvements to their post training pipeline.” Those aren't flattering in the same way. The first is a bugfix wearing a launch announcement: the April model was broken. The second is that DeepSeek found something in post-training the rest of the industry hasn't. The guilty explanation, which the thread was less polite about, is that the eval was fit rather than beaten.
Here's what keeps it from being pure vendor theater: Artificial Analysis ran it independently and scored it 50 on their Intelligence Index: a 10-point jump over the April build, one point behind GPT-5.6 Luna at 51, and six points above DeepSeek's own V4-Pro flagship. A cheap model outscoring its vendor's premium tier is an awkward result to publish, which is a decent signal it wasn't reverse-engineered from a marketing goal. The direction of the jump is corroborated. The magnitude on DeepSeek's own harness is not.
The verbosity tax
Cheap tokens, expensive tasks
The counter-narrative landed the same day, and it comes with a receipt. Artificial Analysis's own eval page shows V4-Flash-0731 burned 210 million output tokens to complete their Intelligence Index, against a 62 million median for comparable models. That's roughly 3.4x the tokens for the same work, and it cost $72.02 to run the index. One Hacker News commenter measured it needing “3.6x more tokens to finish the same work as Gemini Flash 3.6.”
Practitioners hit it within hours, independently, and got progressively less polite about it. On r/LocalLLaMA, one user reported it “over two times as verbose compared to the ‘high’ thinking variant of flash.” Another: “The verbosity is a turn off. It basically eats away your context window.” A third was blunter: “It's using a LOT more tokens, meaning it's so verbose it might be incoherent for real work due to context saturation. Feels like benchmaxxing.”
Do the arithmetic on the right line item. A 4.3x cheaper token that you need 3.4x more of is a 1.3x cheaper task. That is the number for your spreadsheet, and it is still the pessimistic read: run the same division on the cache line, where the money lives, and 7x over 3.4x lands nearer 2x. Two things keep even that honest. The 3.4x is measured against a peer-group median, not against Luna specifically, so it is a directional figure rather than a head-to-head one. And none of it prices context-window pressure, which is a correctness problem, not a line item.
There is a second way these comparisons go wrong, and it is quieter. Both models expose reasoning effort levels, and the level is a price multiplier. One commenter's read of the Artificial Analysis chart put DeepSeek Flash at max effort near $0.03 per task for its index score of 50, against Luna at $0.03 on high, $0.04 on xhigh, and $0.07 on max — where Luna finally reaches 51. Treat those figures as approximate; they were read off a chart, not a published table. The structural point survives the imprecision: a favourable DeepSeek benchmark is usually DeepSeek at max against Luna at medium. Same prompt, same task, two different amounts of thinking, and the cheaper column wins because it was allowed to work harder.
The most careful comparison in the HN thread controlled for reasoning effort level and landed at: “OpenAI Luna between 2x and 3x the price of Deepseek Flash, what you get is 2 to 5 times faster inference.” That's a real tradeoff with a real answer: money-bound, take DeepSeek; latency-bound, take Luna. It is not the 89x that fits in a tweet.
The rebuttal is fair too, and also from the thread: “Cost per task is what matters. Not all tokens are created equal.” Correct, and it cuts both ways. The sticker price is the wrong number to plan around in either direction.
Real invoices
What the bill actually looks like
Enough modeling. Here are invoices from people who pay them.
One engineer posted their last 30 days on DeepSeek: “Cost: $4.55USD, API requests: 3,467, Tokens: 323,183,886.” Three hundred and twenty-three million tokens for the price of a sandwich, on a harness they credit with roughly 99% cache-hit rates. Another described one-shotting a low-poly flight simulator: “almost 300M tokens (cached) and circa 1$ cost,” burning nearly two full 1M context windows in 45 minutes, with a result they called outstanding.
“It hallucinates plenty, about the same as Codex models… You must still review 100% of the code as if it's trying to sell you insurance.”
That is the actual price: your review capacity, spent at the rate a very cheap and very fast model can generate things to review. We wrote a whole article about that particular exhaustion yesterday, and a day later the arithmetic hasn't improved.
Self-hosting works, with an asterisk about effort. One team running QA analysis on voice transcriptions across two B300s reported operating “at 2-5% of the cost of running on Equiv Frontier,” then added the part that belongs in every self-hosting pitch: “It took us about a month to get the inference configured to achieve these numbers.”
The weights are MIT-licensed on HuggingFace: a sparse mixture-of-experts, roughly 13B active parameters, and a total parameter count that DeepSeek's own model card and OpenRouter cannot agree on. When the vendor and the aggregator disagree about how big the thing is, treat every exact figure with suspicion.
Launch-day reports cluster around 10-14 tokens/second on a single high-end GPU with 128GB of system RAM, and around 60 on a pair of DGX Sparks. Every one of those numbers came off a pre-release community quantization, and nobody has published benchmark scores at any quantization level. Plenty of people got it running. Not one of them can tell you what it cost in quality, and by their own admission they are eyeballing it.
Half-life
The 35-minute panic
At 07:48 on launch morning, a user on r/LocalLLaMA sounded an alarm: “Smth is fucked. It's showing a 10x cost increase compared to original deepseek v4 flash.”
Three minutes later, same user, having actually diagnosed it: “Its not the token use. It's that it does not ever cache read. It is only repeatedly cache writing.” Someone else pulled the breakdown showing cache-write cost per task going from about $0.01 to $0.18. At 08:23, thirty-five minutes after the first post, the same user closed the loop: “Has been fixed.” A wrong number on a benchmark dashboard, corrected before most of the Western hemisphere had opened a terminal.
Nothing about the model changed in those 35 minutes. What changed was a third-party dashboard's cost accounting, and in that window a real “DeepSeek just got 10x more expensive” narrative was born, propagated, and killed. Thirty-five minutes is the half-life of a cost number in this market. Build your model on someone else's dashboard and that is the shelf life you inherit.
The fine print
The fine print has a fuse in it
Three things in the documentation are not in anybody's cost model. Start with the one on a timer. DeepSeek's pricing page states that “during peak hours, prices will be 2x the regular prices, applicable to all billing items,” with peak defined as 09:00–12:00 and 14:00–18:00 Beijing time. The effective date is pending an official announcement.
Here is the peak sheet nobody has printed, since “all billing items” means all three lines double: cache-hit input goes from $0.0028 to $0.0056, cache-miss input from $0.14 to $0.28, and output from $0.28 to $0.56. Most of the advantage survives that. Output is still 2.1x cheaper than Luna instead of 4.3x, and the cache line is still 3.6x cheaper instead of 7x.
One line does not survive. At $0.28, peak cache-miss input costs 40% more than Luna's $0.20. A workload with a poor cache-hit rate, running in the Beijing afternoon, would pay DeepSeek more per input token than the model it supposedly undercut. Almost nobody in the launch-day threads noticed the surcharge at all, and every price-per-benchmark chart published today has an undated 2x sitting underneath it.
Second: V4-Flash is blind. No vision at all, in a model sold on agentic benchmarks, so the moment your agent has to look at a screenshot, a diagram, a PDF, or a UI, it phones a friend. It is the most-cited reason practitioners keep Luna or Gemini alongside it rather than instead of it.
The third is the sneakiest. DeepSeek kept the API model string deepseek-v4-flash and swapped the model underneath it. If you called that endpoint yesterday and call it today, you get a different model, silently, with different verbosity characteristics and a different bill. As one commenter put it: “now when someone talks about DeepSeek V4 Flash, in benchmarks, on inference providers, which version do they actually mean?” Every DeepSeek-versus-anything benchmark published before this morning is now ambiguous, and the SWE-bench and LiveCodeBench figures currently circulating for “V4 Flash” belong to the April build, not this one.
The playbook
What to do before the surcharge lands
None of this is a reason not to try it. It's a reason to try it with a stopwatch and a spreadsheet instead of a screenshot.
- Price the task, not the token3.4x verbosity turns a 4.3x cheaper token into a 1.3x cheaper task; measure completions
- Instrument the cache firstAt $0.0028 against $0.02 the cache line is the whole advantage; a bad harness throws it away
- Pin the version, model the surchargeThe API string changed meaning underneath you, and a 2x peak multiplier has no start date
Measure cost per completed task, not per million tokens. The 3.4x verbosity multiplier means the sticker price will lie to you by roughly the amount you'd care about. Instrument your cache-hit rate before you migrate anything, because at $0.0028 versus $0.02 the cache line is where the entire advantage lives, and a harness that caches badly throws the whole thesis away. Pin your model version explicitly rather than trusting a string that now means two different things. And put the 2x peak surcharge in the forecast now, at the share of your traffic that lands in the Beijing afternoon, rather than the week it switches on. The pricing page has already told you where the fuse is; the only missing number is the date.
Then keep the exit door open. What made this week remarkable runs past DeepSeek: OpenAI cut a mainstream tier by 80%, and inside of 24 hours a competitor made the cut irrelevant without shipping hardware or architecture, and without changing a single digit on its own price sheet. On OpenRouter's own numbers, DeepSeek's token share doubled from 9% to 18% between January and June 2026, before any of this.
July 2026 is the month the frontier stopped being a place you buy your way to. If a 645% benchmark jump really can come from re-running post-training on frozen weights, then the moat everyone spent three years and several hundred billion dollars digging is, at minimum, considerably narrower than the invoices implied. Verify the number before you bet the roadmap on it. But put a calendar reminder on the replication, because if it holds, all of that money bought a lead that one post-training run closed to a single index point.