B200 FP4: MTP vs Off Speculative Decoding
Speculative decoding comparison of MTP versus Off on B200 FP4 (NVIDIA Blackwell) running MiniMax M2.5/M2.7. Throughput, cost, and interactivity differences across LLM workloads. Use the chart controls below to switch sequences and metrics — same interactions as the main inference chart.
MTP acceptance-rate implementations differ across inference engines. Points from different engines are not directly comparable on the same curve — throughput and cost at matched interactivity may reflect engine-level differences rather than pure speculative decoding gains. Interpret cross-engine comparisons with caution.
At 66 tok/s/user on MiniMax M2.5/M2.7 (B200 FP4), Off delivers 8554 tok/s/GPU at $0.06 per million tokens; MTP hasn't been benchmarked at this target.
Off hits 2744 tok/s/GPU for $0.20 per million tokens at 110 tok/s/user on MiniMax M2.5/M2.7 (B200 FP4). No MTP data at this operating point.
Off: 772 tok/s/GPU, $0.71 per million tokens at 154 tok/s/user on MiniMax M2.5/M2.7 (B200 FP4). MTP is unmeasured here. (Numbers reflect this URL's pinned 1k/1k · fp4 workload — changing sequence or model updates both the table and chart; the table stays pinned to this page's precision, so precision toggles in the controls affect the chart only.)

| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/gpu) | MTP:—Off:8553.8 | MTP:—Off:2744.2 | MTP:—Off:772.0 |
| Cost ($/M tok) | MTP:—Off:$0.063 | MTP:—Off:$0.198 | MTP:—Off:$0.710 |
| tok/s/MW | MTP:—Off:5002251 | MTP:—Off:1604791 | MTP:—Off:451468 |
| Concurrency | MTP:—Off:~1000 | MTP:—Off:~128 | MTP:—Off:~12 |
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.