GPU rankings
Kimi K3 2.8T

Cheapest GPU for Kimi K3

GB300 NVL72 is the cheapest at $0.080 per million tokens, 23% below B300.

Measured on the AgentX agentic coding workload at a matched interactivity target of 50 tokens/s per user. Hardware without a measurement at this operating point is not ranked.

Cheapest GPU for Kimi K3: live benchmark ranking
RankGPUTokens/s per GPU$ / 1M tokensPrecisionEngine
1GB300 NVL728,047$0.080fp4Dynamo vLLM
2B3006,083$0.10fp4vLLM
3GB200 NVL724,579$0.11fp4Dynamo vLLM
4B2003,976$0.12fp4Dynamo vLLM
5MI355X1,622$0.26fp4vLLM

Optimizing for speed instead of cost? See the fastest GPU for Kimi K3

Methodology

Every number on this page is a measurement, not a spec-sheet estimate. The InferenceX fleet serves Kimi K3 on real hardware with community serving engines, sweeping concurrency to trace each platform's throughput-versus-interactivity frontier.

Platforms are then read at the same operating point (50 tokens/s per user) so the comparison is iso-interactivity: a GPU cannot win by quoting throughput at an unusably slow per-user speed. Cost converts measured throughput to $ per million total tokens using $/GPU/hr rates from the SemiAnalysis AI Cloud TCO model.

The derivation is shared with the InferenceX overview leaderboard, and results re-run continuously, so this ranking updates as new engine releases and configs land.

Frequently asked questions

What is the cheapest GPU to run Kimi K3?
As of the latest benchmark runs, GB300 NVL72 is cheapest at $0.080 per million total tokens on the AgentX agentic coding workload, at hyperscaler $/GPU/hr pricing and a matched interactivity target of 50 tokens/s per user.
How is cost per million tokens calculated?
Measured throughput per GPU at the 50 tokens/s per user tier is converted to $ per million total (input plus output) tokens using hyperscaler $/GPU/hr rates from the SemiAnalysis AI Cloud TCO model. Slower interactivity targets or cheaper rental tiers change the absolute numbers but rarely the order.
How often is this ranking updated?
Benchmarks re-run continuously on the InferenceX cluster fleet; the newest result feeding this ranking landed on 2026-08-21. The page always renders the latest data.

Explore the data