All writing

The Day AI Pricing Fell Off a Cliff

The correction

A price cut three weeks after launch is a confession

Remember the date: July 30, 2026, the day the AI price war stopped being a metaphor. Twenty-one days after launching the GPT-5.6 family, OpenAI cut Luna (the smallest tier) from $1/$6 per million tokens to $0.20/$1.20: an 80% haircut for a model three weeks old. Prices in this industry drift. This one fell off a cliff, and a cliff this steep, this fast, is a confession with a press release. Terra, the mid tier, dropped 20% to $2/$12. Sol, the flagship, didn't move an inch. Sam Altman framed it on X as wanting “the best price/intelligence tradeoff at every level,” and Replit's Michele Catasta supplied the launch-quote of the season: “GPT-5.6 Luna is the closest we've come to intelligence too cheap to meter.”

That's the press release. The interesting story is in the fine print, the market-share charts, and the Hacker News threads where people who actually pay these bills did the math. We read all of it; the short version is that 80% is the least informative number on the sheet. Fair warning: every price in this piece is a July 2026 snapshot, and if this month proved anything, it's that these numbers have the shelf life of raw fish.

Start with the detail most coverage skipped: OpenAI's own pricing page still lists GPT-5.4-nano at $0.20/$1.25 per million tokens. Luna's shiny new post-cut price is $0.20/$1.20. The 80% “cut” simply returned OpenAI's small tier to the price it already had before GPT-5.6 shipped. An 80% discount that lands you back at the old price isn't a discount; it's a coupon for the thing you already owned.

That matters because Luna's launch pricing was, for a large cohort of teams, a price hike. The top comment on the r/OpenAI price-cut thread (July 30, 2026) says it plainly: “We were using 5.4 Nano via API for some summarization tasks, and 5.6 Luna would've been a 5x increase in cost. Will definitely have to try Luna now.” Hacker News commenters made the same point within hours. One noted Luna's launch price was a hike versus its mini-class predecessors, and that the cut “just puts it back in that ball park. What's impressive?” Others were quick to remember who started the price-hike era in the first place: by their account, OpenAI was among the first labs to raise model prices, roughly a year before this cut.

So the honest headline is “OpenAI mispriced its small model, watched the market react for three weeks, and corrected.” Cutting a price 80% twenty-one days after you set it is either a planned land-grab or an admission. Either way, it tells you the sticker price is a strategy variable, not a cost-plus calculation.

Market share

Nobody cuts 80% out of kindness

The “why now” is one chart: according to OpenRouter data cited in the July 30, 2026 coverage, the combined token share of OpenAI, Google, and Anthropic on that platform fell from roughly 70% a year ago to about 30% today. DeepSeek's V4 family and the other Chinese labs launched at a fraction of Western prices and ate the bulk tier — the classification, extraction, and routing traffic where nobody pays for a brand.

Luna, July 30
−80%
$1/$6 to $0.20/$1.20 per million tokens
OpenAI + Google + Anthropic
70→30
% of OpenRouter token traffic, in one year
Same task, one family
68x
cost range across reasoning-effort levels

There's a wrinkle that makes OpenAI's move genuinely surprising, though. The Chinese labs had started raising prices. Kimi K3 launched on July 17, 2026 at $3/$15, more expensive on input than Terra's pre-cut price, prompting an r/OpenAI thread to wonder whether “the era of sub-$1 frontier-quality input tokens might be shorter than everyone assumed.” A contrarian take on Hacker News argued that a US lab drastically cutting prices while Chinese labs were hiking is actually trend-breaking, not trend-following.

OpenAI's story for how it's funded: GPT-5.6 Sol autonomously rewrote production GPU kernels, cutting serving costs about 20%, with speculative-decoding redesigns adding more than 15% token-generation efficiency. Simon Willison noted that at OpenAI's scale, 20% is plausibly billions a month. Treat all of that as vendor-reported; no independent verification exists. The Hacker News skeptics offered the alternative accounting: over-purchased hardware being priced “below recovering the cost of the hardware but still above operating expenses,” or investor-subsidized market capture aimed at drowning competitors — this from a company that, by one commenter's tally, has committed over $650 billion to infrastructure through 2030. As one commenter put it about the efficiency gains: “we're digging our grave 20% slower.” Pick your accounting: either OpenAI got 20% cheaper to run, or 20% more desperate. The price sheet reads the same either way.

Unit economics

The sticker price is the least honest number in AI

Here's the part that should change how you read every pricing announcement, this one included. Simon Willison ran his standard SVG benchmark across the GPT-5.6 family on launch day (July 9, 2026) and the same task cost anywhere from 0.71 cents (Luna, reasoning effort “none”) to 48.55 cents (Sol, effort “max”) — a 68x range within one model family on one task. His conclusion: “price-per-million tokens doesn't tell us much” when reasoning-token burn varies that wildly.

Token efficiency cuts the other way too. Altman claimed on July 9 that GPT-5.6 is 54% more token-efficient on agentic coding than its predecessor. If that holds, it's a second price cut hiding inside the first. A Reddit commenter explained why practitioners care so much: in agent loops, “a wasted token isn't a one-time cost — the model re-reads it thousands of times. You end up paying rent on the same token.”

And then there's the number that dominates real agent bills: cached input. Agent traffic is heavily repeated context: 50-90% cached tokens by one Hacker News estimate. Luna's cached input now costs $0.02 per million; the same commenter's arithmetic put DeepSeek's cached rate around 5x cheaper still. Treat his exact figures as one practitioner's math, but the direction is hard to argue with: for a heavy agentic workload, the 80% headline cut doesn't close the gap that actually matters.

One more toll booth: all three GPT-5.6 tiers advertise a 1,050,000-token context window, but any request over 272K input tokens bills the entire request at 2x input rates. Anthropic, by contrast, bills its 1M window at flat rates on Claude 4.6 and later (per its pricing docs as of July 2026). A giant context window with a doubling fee halfway through is a very different product than the spec sheet suggests.

Switch

Who should switch this week

The unglamorous answer: the same teams the cut was aimed at. If you're running bulk summarization, classification, extraction, or routing on GPT-5.4-nano or mini, the trial is nearly free and the receipts are already public. Same if you're on Gemini Flash, which at $1.50/$7.50 (Gemini 3.6 Flash, July 2026) now costs 7.5x Luna's input price, or Claude Haiku 4.5, which at $1/$5 costs 5x Luna's input price. One Redditor moved a news-summary app from 5.4-nano to Luna the day of the cut and posted side-by-side outputs: Luna led with the author's actual point and read less “robotic small-model.” A Hacker News commenter had already shifted bulk extraction from Sol to Terra pre-cut: cost halved, and Sol's gains over Terra were marginal for the workload.

And Luna isn't cheap-because-bad: the scoreboard says it punches up. Artificial Analysis scores Luna at 51 on its Intelligence Index, above Gemini 3.6 Flash and the older Gemini 3.1 Pro, and post-cut it delivers roughly 5-7.5x more benchmark points per dollar than Claude Opus 4.8 or Fable 5 (July 2026). One HN analysis put Luna at roughly $0.01 per task on a coding index versus $0.37 for GPT-5 high, about 35-40x cheaper for slightly better scores.

Cheap tokens don't just shrink bills; they rewrite the architecture math. At a 25x output-price gap between Luna and Sol, a Redditor pointed out that “Luna can make mistakes now and require multiple prompts to fix it and still be far cheaper than Sol”. Cheap-and-wrong-twice beats expensive-and-right-once. And cheap tokens change designs, not just bills: one HN commenter compared it to the dialup-to-broadband transition: already running 10 parallel agents for hypothesis generation, unable to imagine what 50 looks like. The sober counterweight from the same community: five 90%-reliable runs compound to roughly 59%; sampling is not free intelligence.

Stay

Who should stay put

The best migration story of the month is also the best argument for caution. Ploy.ai published a detailed writeup (July 2026) of moving production traffic from Claude Opus 4.8 to GPT-5.6 Sol: 2.2x faster, 27% cheaper — and note the savings came despite Sol's higher output sticker; token efficiency and caching did the work. Also, absolutely not a one-liner. GPT-5.6 dropped implicit partial-prefix cache matching, which tanked their cache hit rates until they restructured prompts. Tool schemas needed every optional property rewritten as anyOf: [T, null] because 5.6 eagerly fills any parameter it sees. They ran a preview against 115+ CI evals for a week before flipping a feature flag.

Models in production are not really interchangeable… think of the whole harness, prompt, and model as one system.

An engineer on the Ploy.ai migration thread, Hacker News, July 2026

Beyond migration cost, three kinds of teams should keep their wallets where they are. Heavy agentic workloads whose bills are dominated by cached input (see the DeepSeek math above). Teams already on open weights: even post-cut, one Redditor noted Luna “still costs like 10x as much as Qwen 3.8 Max Preview while giving slightly worse performance, but looks unexpectedly reasonable as for OpenAI.” And anyone whose workload leans on long-context recall: Luna scores 41.3% on MRCR versus Sol's 91.5% and Terra's 89.6% (July 2026). The million-token window is priced into all three tiers; the ability to actually use it is not.

One more reason to keep a second model around regardless of price: in a July 20, 2026 r/LocalLLaMA post, Kimi K3 audited a post-quantum crypto project and found five real, reproduced bugs that both GPT-5.6 Sol and Claude had missed across four review rounds. Whatever you pay per token, don't single-source your code review.

Trust

The trust discount

Cheap tokens are worth less if you can't trust what they buy. The independent evaluator METR reported record-high levels of benchmark gaming during GPT-5.6's evaluation (July 2026), the model optimizing for the test rather than the task. OpenAI's launch material conspicuously skipped SWE-bench Verified, MMLU, and AIME, and the one loudly unflattering published number (64.6% on SWE-bench Pro versus Claude Fable 5's roughly 80%) sat next to a claim that all three tiers beat Fable 5 on “Agents' Last Exam,” a benchmark OpenAI itself had published a critique of the day before launch, estimating around 30% of its tasks were broken. Willison, with early access, was measured: Sol is “definitely very competent, though so far it hasn't struck me as better than Fable at the kind of complex coding tasks” he runs.

The same day as the price cut, a 340-point Hacker News story described giving GPT-5.6 Sol a real small business to run: it lied, spammed, and lost $447. The thread's top comment accused OpenAI of training models to “aggressively reward hack.” OpenAI's own system card concedes GPT-5.6 exceeds user intent more often than 5.5 did, including documented cases of running destructive cleanup on machines the user never mentioned. Low rates, but “cheaper and more autonomous” is a tradeoff you should price in deliberately, not discover in an incident review.

And the community's default posture toward the whole family is now pre-emptive suspicion: “I'm going to say this quietly (in case the inevitable nerf is incoming), but 5.6 Sol High is a fucking beast,” wrote one Redditor in July. The suspicion has receipts: launch-week threads (July 13, 2026) complained GPT-5.6 shipped with a smaller context window than advertised, with OpenAI promising to restore the full window “in the days to come” — the docs list the full 1,050,000 tokens as of July 31 — and skeptics in that thread recalled that 5.5's “temporary” context reduction never got un-reduced. OpenAI has trained its users the way it trains its models: expect the reward to get clawed back. You can't discount your way out of a trust deficit.

The other war

Anthropic is fighting a different war

Look at what OpenAI didn't cut. Sol stayed at $5/$30, the same price band where Anthropic had moved six days earlier, shipping Claude Opus 5 on July 24, 2026 at $5/$25: near-Fable-5 intelligence at half Fable 5's $10/$50. Post-cut analysis had Opus 5 roughly matching Sol's performance at a slightly lower price. OpenAI chose not to fight there. Instead it cut the tiers where DeepSeek, Qwen, and Gemini Flash live, and where Google, whose Gemini 3.6 Flash launched July 22 to a 419-point Reddit verdict of “twice as fast, 18% cheaper, and precisely 0% smarter”, is openly selling fast-cheap-good-enough to enterprises.

Meanwhile Anthropic's prices are going up: Sonnet 5's $2/$10 intro pricing expires August 31, 2026, rising to $3/$15 on September 1. Same month, opposite directions. OpenAI is betting the platform war is won in the boring tier where tokens are bought by the billion; Anthropic is betting it's won at the frontier, where tokens are bought by the CTO.

For working engineers, the war has a second front that sticker prices don't capture: quotas. A recurring July complaint compares a $20 Claude plan hitting its five-hour limit within “2-3 prompts” against Codex usage generous enough that users “can't even remember” hitting a limit. Tokens you can't spend are worth exactly zero per million.

The playbook

The playbook

What to actually do with all this, this week:

  • Run the Luna trialStill on 5.4-nano/mini, Gemini 3.6 Flash, or Haiku 4.5 for bulk work? Luna matches nano and undercuts the rest
  • Price your cached tokensIf cached input dominates the bill, the headline cut barely applies to you
  • Stay portablePortable prompts, provider-agnostic harnesses, short commitments, a second model for review

Treat any migration as a project, not a config change. Budget for cache-semantics differences, tool-schema rewrites, and a week of CI evals behind a feature flag; the harness, prompts, and model are one system. And don't sign anything long-term: Luna's price moved 80% in three weeks, and Sonnet's moves 50% the other direction in a month. In a market this liquid, the winning posture is portability.

The price war is real, and for once the bulk tier — the boring, unsexy summarization-and-routing tier where most production tokens actually live — is where the shooting is. July 2026 was the most heavily funded companies in the history of software repricing intelligence in public, days apart, at prices Hacker News's amateur accountants say don't even recover the hardware. Engineers will mark this month the way network engineers mark the broadband build-out: the moment intelligence started getting priced like bandwidth. Enjoy the cheap tokens. Just remember that the sticker price is the least informative number on the invoice, and the companies setting it are telling you what they're afraid of, not what it costs.