Kimi K3 2.8T · Speculative Decoding

GB300 NVL72 FP4: DSpark vs Off Speculative Decoding

Speculative decoding comparison of DSpark versus Off on GB300 NVL72 FP4 (NVIDIA Blackwell) running Kimi K3 2.8T. Throughput, cost, and interactivity differences across LLM workloads. Use the chart controls below to switch sequences and metrics — same interactions as the main inference chart.

MTP acceptance-rate comparability

MTP acceptance-rate implementations differ across inference engines. Points from different engines are not directly comparable on the same curve — throughput and cost at matched interactivity may reflect engine-level differences rather than pure speculative decoding gains. Interpret cross-engine comparisons with caution.

Kimi K3 2.8T: GB300 NVL72 FP4 DSpark versus Off speculative decoding comparison at matched interactivity levels
GB300 NVL72 FP4 DSpark versus Off speculative decoding comparison for this page's canonical default workload.
No interpolated data available for the default workload on this configuration. Use the chart controls below to select a sequence and precision with benchmark data for both configurations.

Inference Performance

Agentic inference metrics from the AgentX scenario and fixed-sequence inference metrics across models, hardware configurations, and serving parameters.

Vendor:
Deployment:
Spec Decoding: