Fastest GPU for Kimi K3
GB300 NVL72 leads at 8,047 tokens/s per GPU, 32% ahead of B300.
Measured on the AgentX agentic coding workload at a matched interactivity target of 50 tokens/s per user. Hardware without a measurement at this operating point is not ranked.
| Rank | GPU | Tokens/s per GPU | $ / 1M tokens | Precision | Engine |
|---|---|---|---|---|---|
| 1 | GB300 NVL72 | 8,047 | $0.080 | fp4 | Dynamo vLLM |
| 2 | B300 | 6,083 | $0.10 | fp4 | vLLM |
| 3 | GB200 NVL72 | 4,579 | $0.11 | fp4 | Dynamo vLLM |
| 4 | B200 | 3,976 | $0.12 | fp4 | Dynamo vLLM |
| 5 | MI355X | 1,622 | $0.26 | fp4 | vLLM |
Optimizing for cost instead of speed? See the cheapest GPU for Kimi K3
Methodology
Every number on this page is a measurement, not a spec-sheet estimate. The InferenceX fleet serves Kimi K3 on real hardware with community serving engines, sweeping concurrency to trace each platform's throughput-versus-interactivity frontier.
Platforms are then read at the same operating point (50 tokens/s per user) so the comparison is iso-interactivity: a GPU cannot win by quoting throughput at an unusably slow per-user speed. Cost converts measured throughput to $ per million total tokens using $/GPU/hr rates from the SemiAnalysis AI Cloud TCO model.
The derivation is shared with the InferenceX overview leaderboard, and results re-run continuously, so this ranking updates as new engine releases and configs land.
Frequently asked questions
- What is the fastest GPU for Kimi K3 inference?
- As of the latest benchmark runs, GB300 NVL72 leads at 8,047 tokens/s per GPU on the AgentX agentic coding workload, measured at a matched interactivity target of 50 tokens/s per user.
- How is "fastest" measured?
- Every platform is read at the same interactivity tier (50 tokens/s per user) on the AgentX agentic coding workload, then ranked by measured throughput per GPU. This is the same derivation the InferenceX overview leaderboard renders, so the ranking can never disagree with the dashboard.
- How often is this ranking updated?
- Benchmarks re-run continuously on the InferenceX cluster fleet; the newest result feeding this ranking landed on 2026-08-21. The page always renders the latest data.