MI300X FP8: MTP vs Off Speculative Decoding
Speculative decoding comparison of MTP versus Off on MI300X FP8 (AMD CDNA 3) running Qwen 3.5 397B-A17B. Throughput, cost, and interactivity differences across LLM workloads. Use the chart controls below to switch sequences and metrics — same interactions as the main inference chart.
MTP acceptance-rate implementations differ across inference engines. Points from different engines are not directly comparable on the same curve — throughput and cost at matched interactivity may reflect engine-level differences rather than pure speculative decoding gains. Interpret cross-engine comparisons with caution.
Off hits 346 tok/s/chip for $0.76 per million tokens at 43 tok/s/user on Qwen 3.5 397B-A17B (MI300X FP8). No MTP data at this operating point.
Off: 190 tok/s/chip, $1.39 per million tokens at 53 tok/s/user on Qwen 3.5 397B-A17B (MI300X FP8). MTP is unmeasured here.
At 62 tok/s/user on Qwen 3.5 397B-A17B (MI300X FP8), Off delivers 123 tok/s/chip at $2.15 per million tokens; MTP hasn't been benchmarked at this target. (Numbers reflect this URL's pinned 1k/1k · fp8 workload — changing sequence or model updates both the table and chart; the table stays pinned to this page's precision, so precision toggles in the controls affect the chart only.)

| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | MTP:—Off:345.8 | MTP:—Off:189.5 | MTP:—Off:122.7 |
| Cost ($/M tok) | MTP:—Off:$0.763 | MTP:—Off:$1.392 | MTP:—Off:$2.150 |
| tok/s/MW | MTP:—Off:248748 | MTP:—Off:136357 | MTP:—Off:88306 |
| Concurrency | MTP:—Off:~34 | MTP:—Off:~15 | MTP:—Off:~8 |
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.