GB200 NVL72 FP4: MTP vs Off Speculative Decoding
Speculative decoding comparison of MTP versus Off on GB200 NVL72 FP4 (NVIDIA Blackwell) running GLM 5/5.1. Throughput, cost, and interactivity differences across LLM workloads. Use the chart controls below to switch sequences and metrics — same interactions as the main inference chart.
MTP acceptance-rate implementations differ across inference engines. Points from different engines are not directly comparable on the same curve — throughput and cost at matched interactivity may reflect engine-level differences rather than pure speculative decoding gains. Interpret cross-engine comparisons with caution.
Throughput at 50 tok/s/user on GLM 5/5.1 (GB200 NVL72 FP4): MTP hits 9984 tok/s/chip, Off hits 2207. Per-million costs land at $0.05 and $0.23 respectively. MTP is 352% cheaper per token; MTP delivers 352% more tok/s/chip. Speculative decoding trades extra compute on draft tokens for fewer decoding steps — the payoff depends on sequence length and batch size.
Around the middle of the 38–86 tok/s/user interactivity band, at 62 tok/s/user on GLM 5/5.1 (GB200 NVL72 FP4): MTP runs 9414 tok/s/chip at $0.05/M tokens, Off runs 1209 at $0.43/M. MTP is 679% cheaper per token; MTP delivers 679% more tok/s/chip. Gains from speculative decoding vary by workload; short-output prompts tend to benefit less.
At 74 tok/s/user on GLM 5/5.1 (GB200 NVL72 FP4), MTP delivers 8514 tok/s/chip at $0.06 per million tokens; Off delivers 466 tok/s/chip at $1.11. MTP is 1728% cheaper per token; MTP delivers 1728% more tok/s/chip. Speculative decoding accepts draft tokens to reduce per-token latency — gains vary by workload and prompt distribution. (Numbers reflect this URL's pinned 8k/1k · fp4 workload — changing sequence or model updates both the table and chart; the table stays pinned to this page's precision, so precision toggles in the controls affect the chart only.)

| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | MTP:9983.9Off:2206.8 | MTP:9414.2Off:1208.6 | MTP:8513.9Off:465.8 |
| Cost ($/M tok) | MTP:$0.052Off:$0.234 | MTP:$0.055Off:$0.427 | MTP:$0.061Off:$1.109 |
| tok/s/MW | MTP:5338994Off:1180133 | MTP:5034330Off:646310 | MTP:4552911Off:249075 |
| Concurrency | MTP:~1340Off:~227 | MTP:~818Off:~105 | MTP:~618Off:~74 |
Inference Performance
Agentic inference metrics from the AgentX scenario and fixed-sequence inference metrics across models, hardware configurations, and serving parameters.