Tokens per dollar
Also known as tokens per $1 TCO
In plain English
Tokens per dollar asks how many tokens one dollar of infrastructure spend can produce under the cost basis named on the chart.
Technical definition
Total tokens per dollar divides total tokens produced per chip-hour by the modeled all-in infrastructure cost per chip-hour.
Typical unit
tokens per $1 TCO (tok/$)
Engineering details
The Owning at Large Hyperscaler Volume and Rent - 3 Year Commit variants use their corresponding TCO hourly rates. Historical Trends interpolates the matching total, input, or output throughput and then applies the hourly-cost multiplier.
Why it matters
The metric measures hardware and software cost efficiency, so comparisons must use the same model, workload, interactivity target, token type, and infrastructure cost basis.
How to read it in InferenceX
InferenceX exposes separate total-token axes for Owning at Large Hyperscaler Volume and Rent - 3 Year Commit costs. The Owning at Large Hyperscaler Volume axis is the dashboard default y-axis. Token Revenue per GPU Hour is the separate metric that uses normalized or OpenRouter token sale prices.
Source material
See the concept in real benchmarks
AgentX - InferenceXv3: Does the CUDA Moat Hold Up in Agentic Inferencing?
$3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355X, B200
InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper - Formerly InferenceMAX
GB300 NVL72, MI355X, B200, H100, Disaggregated Serving, Wide Expert Parallelism, Large Mixture of Experts, SGLang, vLLM, TRTLLM
B200 NVFP4 vs H200 FP8 on GLM-5: Up to 3.65x Better Performance per Dollar with SGLang MTP
Both SKUs run SGLang EAGLE MTP; the Blackwell generation lifts perf/$ by ~1.2x at the peak and the NVIDIA GLM-5-NVFP4 checkpoint on FlashInfer TRT-LLM sparse MLA stacks another ~2.4–3.0x on 8K/1K
Qwen3.5 397B on AgentX: B300 FP4 Delivers 12x the Performance per Dollar of H100
What four years of hardware and a 4-bit format buy on a long-context agentic workload