The cost floor
The short answer starts at $40,000 a month
The smallest published Kimi K3 shape fits on one eight-B300 server. Buying that machine and applying the planning assumptions below produces a monthly cost of $40,149.56. A two-node, 16-B200 cluster comes to $48,479.55.
- Owned floor
- $40.1k
- 8 B300s, monthly model
- Rental floor
- $69.7k
- 16 H200s, raw compute
- API parity
- 17.4B
- tokens per month
- Provider spread
- ~5×
- measured output speed
Both owned figures include hardware amortization, electricity, cooling, and $20,000 of monthly platform labor. Neither includes a spare node. Renting moves the floor closer to $70,000 a month for raw compute. A 64-accelerator production footprint ranges from $278,918.40 to $467,200 across the public rates used here, before labor, storage, support, or network charges.
The official Hugging Face repository metadata reports 1,561,018,243,668 bytes of storage. The model card describes 2.8 trillion total parameters, 104 billion active per token, 896 experts, MXFP4 weights, and a 1,048,576-token context window. The checkpoint is 1.56 TB in decimal units before runtime memory and request state.
Fit versus production
K3 fits on 8 B300s or 16 B200s
The current SGLang K3 cookbook lists several working shapes: 8 B300s, 8 GB300s, or 8 MI350X/MI355X accelerators; 16 B200s or H200s; and 32 H100s. High-memory 288 GB accelerators cut the minimum node count in half. Hardware count alone still says nothing about useful concurrency.
Moonshot's launch post recommends 64 or more accelerators for efficient production inference because expert routing benefits from a larger high-bandwidth communication domain. Eight B300s answer whether the model fits. A 64-accelerator design addresses sustained production load.
SGLang publishes 16-, 32-, and 64-GPU Blackwell presets. Its 64-GPU sweep reached roughly 3,000 tokens per second per GPU in a fully data-parallel extreme on 288 GB accelerators, about 192,000 tokens per second across the sweep. The same page says the shape is not a recommended preset and that no preset completed a full serving round on final weights. Do not use that result as a B200 capacity guarantee.
A one-million-token context is a ceiling. Long prompts increase prefill time and request state, then reduce the sessions each replica can carry. Price the shape that fits first. Buy after the exact model build, context distribution, cache policy, and concurrency target survive a load test.
Live provider evidence
Provider speed currently varies by about 5×
Moonshot's Kimi Vendor Verifier lists submitted K3 results from Moonshot, Fireworks, Baseten, Together, DigitalOcean, Inferact, Nebius, and Modal. It checks model behavior and API compatibility; it does not rank serving speed.
For speed, Artificial Analysis was tracking five public endpoints on July 28, 2026. Its default methodology uses roughly 10,000 input tokens, measures output after the first token, and reports the median over the previous 72 hours.
| Provider | Output speed | First answer token |
|---|---|---|
| Fireworks | 164 tok/s | 13.3 s |
| Nebius | 128 tok/s | 17.2 s |
| Together | 56 tok/s | 36.9 s |
| Makora | 51 tok/s | 40.9 s |
| Kimi direct | 33 tok/s | 182.6 s |
Treat the table as a dated snapshot. Providers do not promise these numbers as a service level. Vercel AI Gateway measured 13 to 67 output tokens per second across six K3 routes on live gateway traffic the same day. Different requests, regions, traffic windows, and routing policies produce different rankings.
Together's current PTU table allocates 16,667 input tokens, 166,667 cached tokens, and 3,333 output tokens per minute for each $0.05 PTU-minute. One PTU running continuously costs $2,190 a month and supplies about 55.6 output tokens per second when output is the limiting dimension.
Cloud rental
Renting costs $70,000 to $467,000 a month
The rental model uses 730 hours for an average month and public rates available on July 28, 2026. Every figure below is raw compute. It excludes platform labor, storage, egress, support, and taxes.
| Cluster | Provider | Public rate | Monthly compute |
|---|---|---|---|
| 16 H200 | AWS | $5.97/GPU | $69,729.60 |
| 8 B300 | Fireworks | $12/GPU | $70,080.00 |
| 8 B300 | AWS | $14.04/GPU | $81,993.60 |
| 16 B200 | Together | $8.19/GPU | $95,659.20 |
| 16 B200 | Lambda | $9.86/GPU | $115,164.80 |
| 64 H200 | AWS | $5.97/GPU | $278,918.40 |
| 64 B200 | Together | $8.19/GPU | $382,636.80 |
| 64 B200 | Lambda | $9.36/GPU | $437,299.20 |
| 64 B200 | Fireworks | $10/GPU | $467,200.00 |
The inputs come from AWS Capacity Blocks, Fireworks pricing, Together's GPU cluster pricing, and Lambda's 1-Click Cluster pricing.
Add $20,000 of monthly platform labor to the 8- and 16-accelerator cases and $30,000 to the 64-accelerator cases for the comparison used here. The resulting range is $89,729.60 for 16 AWS H200s through $497,200 for 64 Fireworks B200s.
Capital and operations
Buying starts at $575,245
Exxact's configured eight-B300 server is $575,245 with a maximum draw of 15.575 kW. Its eight-B200 server is $395,048.50 with a 14.227 kW maximum draw.
The ownership model adds 15% to the server price for fabric, storage, rack equipment, and installation; amortizes the result over 36 months; applies a 1.3 power-usage-effectiveness factor; prices electricity at $0.12/kWh; and budgets $20,000 of monthly platform labor for 8 or 16 accelerators and $30,000 for 64.
| Cluster | Hardware | Monthly model | API parity |
|---|---|---|---|
| 8 B300 | $575,245 | $40,149.56 | 17.4B tokens |
| 16 B200 | $790,097 | $48,479.55 | 21.0B tokens |
| 64 B200 | $3,160,388 | $143,918.20 | 62.3B tokens |
| 64 B300 | $4,601,960 | $191,196.50 | 82.8B tokens |
The totals exclude tax, financing, colocation rent, support contracts, replacement inventory, and redundancy. One eight-GPU server is a hardware floor with a single failure domain. A spare node adds another $395,048.50 for B200 or $575,245 for B300 before installation.
Bill parity
API break-even starts at 17.4 billion tokens
Moonshot's K3 API price table lists cache-hit input at $0.30 per million tokens, cache-miss input at $3.00, and output at $15.00. Artificial Analysis uses a 7:2:1 cached-input, uncached-input, and output mix for general agentic traffic.
Renting pushes the threshold higher. Sixteen AWS H200s plus platform labor reach parity around 38.8 billion tokens a month. Sixty-four AWS H200s plus labor reach 133.7 billion. Together's provisioned option also matters: 18 PTUs cost $39,420 a month and allocate almost 1,000 output tokens per second, close to the monthly spend of the smallest owned B300 model without the hardware purchase.
At 17.4 billion tokens with the 7:2:1 mix, average traffic is about 6,614 total tokens per second, including 661 output tokens per second. That is roughly four Fireworks streams or twenty Kimi-direct streams generating continuously after the first token. Real traffic has pauses, bursts, rate limits, and input-processing time, so use this as a concurrency sanity check.
Caching moves the line. With 90% cache hits and a 10:1 input-to-output ratio, the 8-B300 threshold rises to 21.3 billion tokens. With the same ratio and no cache hits, it falls to 9.8 billion. Use production token logs rather than a vendor's average mix.
The decision
Self-host when control or steady load pays for it
Self-hosting earns its place when it solves a requirement the providers cannot meet or when measured steady load clears the owned cost by a useful margin. The custom K3 license requires review, especially for some Model-as-a-Service businesses above $20 million in trailing 12-month revenue and large commercial products subject to its attribution terms.
Security has an operating cost. vLLM's scaling guide says inter-node traffic is unencrypted and belongs on a private network. Logs, backups, prompt caches, generated code, access control, patching, failed workers, and rollback still need owners.
- 01Serverless
Use while demand is uncertain or below the modeled break-even range.
- 02Provisioned
Reserve throughput when traffic is steady and provider terms fit the security boundary.
- 03Self-hosted
Buy after a production-shaped load test and a named team accepts the on-call work.
Run a 30-day shadow test before buying. Record cache hits, input and output tokens, time to first answer, output speed, concurrency, queue time, provider errors, and cost per successful task. Test the exact prompts, tools, reasoning effort, and fallback policy planned for production.
If the decision still works after using cluster rates, current provider performance, a spare-node budget, and a named on-call team, the math is ready for finance.
