All model and GPU pairings
Kimi K3 2.8TNVIDIA Blackwell

Running Kimi K3 on GB300 NVL72

Quick answer

Kimi K3 sustains 8,047 tokens/s per GPU on GB300 NVL72 at 50 tokens/s per user, which works out to $0.080 per million tokens at hyperscaler pricing, served by Dynamo vLLM. Fastest measured TTFT: 1.6 ms; fastest TPOT: 0.0 ms (each the best across all configs, not one run).

Benchmarked configs

7

Serving engines

Dynamo vLLM

Precisions

fp4

Run dates

2026-08-182026-08-18

Throughput at every interactivity target

Serving is a trade-off: push more concurrent users through a GPU and each user's tokens arrive slower. The ladder below reads the measured frontier at each per-user speed target on the AgentX agentic coding workload, using the best engine and precision at that point.

Per-user targetTokens/s per GPU$ / 1M tokensEnginePrecision
30 tok/s9,624$0.067Dynamo vLLMfp4
50 tok/s8,047$0.080Dynamo vLLMfp4
75 tok/s6,474$0.099Dynamo vLLMfp4
100 tok/s5,257$0.12Dynamo vLLMfp4
150 tok/s3,517$0.18Dynamo vLLMfp4
200 tok/s2,280$0.28Dynamo vLLMfp4

What serving actually costs

Converting the 50 tokens/s per user operating point to $ per million total tokens across rental pricing tiers from the SemiAnalysis AI Cloud TCO model.

Pricing tier$ / GPU / hr$ / 1M tokens
Hyperscaler$2.31$0.080
Neocloud$2.79$0.096
Retail$3.30$0.11

Frequently asked questions

How fast is Kimi K3 on GB300 NVL72?
At an interactivity target of 50 tokens/s per user on the AgentX agentic coding workload, GB300 NVL72 sustains 8,047 tokens/s per GPU serving Kimi K3 with Dynamo vLLM in FP4. Peak measured throughput across all configs is 11,797 tokens/s per GPU.
How much does it cost to serve Kimi K3 on GB300 NVL72?
$0.080 per million total tokens at hyperscaler $/GPU/hr pricing, at 50 tokens/s per user. Neocloud and retail rental tiers are tabulated above; slower interactivity targets lower the cost further.
Which serving engines run Kimi K3 on GB300 NVL72?
The runs behind this page used Dynamo vLLM in FP4, including disaggregated prefill and multi-node serving. Engines are rebuilt and re-benchmarked continuously, so the best config can change between visits.
How are these Kimi K3 numbers measured?
Every number is measured on real GB300 NVL72 hardware by the InferenceX fleet, sweeping concurrency on the AgentX agentic coding workload to trace the throughput-versus-interactivity frontier; the newest run landed on 2026-08-18. The same derivation powers the InferenceX overview leaderboard.

Explore the data