All model and GPU pairings
Kimi K3 2.8TNVIDIA Hopper

Running Kimi K3 on H200

Quick answer

Kimi K3 runs on H200: 35 benchmarked configs so far. See the interactivity ladder below for measured operating points.

Benchmarked configs

35

Serving engines

vLLM

Precisions

fp4

Run dates

2026-08-072026-08-07

Throughput at every interactivity target

Serving is a trade-off: push more concurrent users through a GPU and each user's tokens arrive slower. The ladder below reads the measured frontier at each per-user speed target on the AgentX agentic coding workload, using the best engine and precision at that point.

Per-user targetTokens/s per GPU$ / 1M tokensEnginePrecision
30 tok/s--vLLMfp4
50 tok/s--vLLMfp4
75 tok/s--vLLMfp4
100 tok/s--vLLMfp4
150 tok/s--vLLMfp4
200 tok/s--vLLMfp4

Frequently asked questions

How fast is Kimi K3 on H200?
The InferenceX fleet has 35 benchmarked configs for this pairing; see the interactivity ladder above for the operating points reached so far.
How much does it cost to serve Kimi K3 on H200?
Cost per million tokens is derived from measured throughput and $/GPU/hr rates from the SemiAnalysis AI Cloud TCO model; it appears once this pairing reaches the primary interactivity tier.
Which serving engines run Kimi K3 on H200?
The runs behind this page used vLLM in FP4 and multi-node serving. Engines are rebuilt and re-benchmarked continuously, so the best config can change between visits.
How are these Kimi K3 numbers measured?
Every number is measured on real H200 hardware by the InferenceX fleet, sweeping concurrency on the AgentX agentic coding workload to trace the throughput-versus-interactivity frontier; the newest run landed on 2026-08-07. The same derivation powers the InferenceX overview leaderboard.

Explore the data