Kimi K3 2.8T · Speculative Decoding
B300 FP4: MTP vs Off Speculative Decoding
Speculative decoding comparison of MTP versus Off on B300 FP4 (NVIDIA Blackwell) running Kimi K3 2.8T. Throughput, cost, and interactivity differences across LLM workloads. Use the chart controls below to switch sequences and metrics — same interactions as the main inference chart.
MTP acceptance-rate comparability
MTP acceptance-rate implementations differ across inference engines. Points from different engines are not directly comparable on the same curve — throughput and cost at matched interactivity may reflect engine-level differences rather than pure speculative decoding gains. Interpret cross-engine comparisons with caution.

No interpolated data available for the default workload on this configuration. Use the chart controls below to select a sequence and precision with benchmark data for both configurations.
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.
Vendor:
Deployment:
Spec Decoding: