GB200 NVL72 FP8: MTP vs Off Speculative Decoding
Speculative decoding comparison of MTP versus Off on GB200 NVL72 FP8 (NVIDIA Blackwell) running Qwen 3.5 397B-A17B. Throughput, cost, and interactivity differences across LLM workloads. Use the chart controls below to switch sequences and metrics — same interactions as the main inference chart.
MTP acceptance-rate implementations differ across inference engines. Points from different engines are not directly comparable on the same curve — throughput and cost at matched interactivity may reflect engine-level differences rather than pure speculative decoding gains. Interpret cross-engine comparisons with caution.
At 80 tok/s/user on Qwen 3.5 397B-A17B (GB200 NVL72 FP8), MTP delivers 10567 tok/s/GPU at $0.06 per million tokens; Off delivers 2904 tok/s/GPU at $0.21. MTP is 267% cheaper per token; MTP delivers 264% more tok/s/GPU. Speculative decoding accepts draft tokens to reduce per-token latency — gains vary by workload and prompt distribution.
MTP posts 6941 tok/s/GPU for $0.09 per million tokens at 111 tok/s/user on Qwen 3.5 397B-A17B (GB200 NVL72 FP8); Off posts 1523 tok/s/GPU for $0.40. MTP is 348% cheaper per token; MTP delivers 356% more tok/s/GPU. Draft-token acceptance rates determine whether speculative decoding helps or hurts at a given concurrency level.
Throughput at 142 tok/s/user on Qwen 3.5 397B-A17B (GB200 NVL72 FP8): MTP hits 3797 tok/s/GPU, Off hits 703. Per-million costs land at $0.16 and $0.86 respectively. MTP is 447% cheaper per token; MTP delivers 440% more tok/s/GPU. Speculative decoding trades extra compute on draft tokens for fewer decoding steps — the payoff depends on sequence length and batch size. (Numbers reflect this URL's pinned 8k/1k · fp8 workload — changing sequence or model updates both the table and chart; the table stays pinned to this page's precision, so precision toggles in the controls affect the chart only.)

| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/gpu) | MTP:10566.7Off:2903.8 | MTP:6941.5Off:1522.9 | MTP:3797.3Off:703.3 |
| Cost ($/M tok) | MTP:$0.058Off:$0.213 | MTP:$0.090Off:$0.404 | MTP:$0.158Off:$0.865 |
| tok/s/MW | MTP:5650620Off:1552851 | MTP:3712015Off:814380 | MTP:2030642Off:376096 |
| Concurrency | MTP:~1115Off:~39 | MTP:~350Off:~15 | MTP:~66Off:~5 |
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.