GB300 NVL72 FP4: DSpark vs Off Speculative Decoding
Speculative decoding comparison of DSpark versus Off on GB300 NVL72 FP4 (NVIDIA Blackwell) running Kimi K3 2.8T. Throughput, cost, and interactivity differences across LLM workloads. Use the chart controls below to switch sequences and metrics — same interactions as the main inference chart.
MTP acceptance-rate implementations differ across inference engines. Points from different engines are not directly comparable on the same curve — throughput and cost at matched interactivity may reflect engine-level differences rather than pure speculative decoding gains. Interpret cross-engine comparisons with caution.

Inference Performance
Agentic inference metrics from the AgentX scenario and fixed-sequence inference metrics across models, hardware configurations, and serving parameters.