Head-to-head AI inference benchmark comparison of B200 (NVIDIA Blackwell) and GB200 NVL72 (NVIDIA Blackwell) on DeepSeek V4.1 Flash 552B. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.
AgentX replays real coding-agent sessions rather than fixed-length prompts, so context grows turn over turn and most of each request is served from cache instead of being recomputed. That turns the comparison into a systems question: KV transfer between nodes, prefix-aware routing, and cache capacity all move the curve alongside raw chip throughput. Fixed-sequence workloads stay the clean baseline for kernel and silicon performance, so the two scenarios answer different questions about the same hardware. Learn more about AgentX →
Throughput at 81 tok/s/user on DeepSeek V4.1 Flash 552B: B200 hits 71402 tok/s/chip, GB200 NVL72 hits 54859. Per-million costs land at $0.01 and $0.01 respectively. B200 is 40% cheaper per token; B200 delivers 30% more tok/s/chip.
B200 / GB200 NVL72 on DeepSeek V4.1 Flash 552B at 153 tok/s/user: 37158 / 26900 tok/s/chip, $0.01 / $0.02 per million tokens. B200 is 49% cheaper per token; B200 delivers 38% more tok/s/chip.
Toward the upper edge of the 9–298 tok/s/user interactivity band, at 226 tok/s/user on DeepSeek V4.1 Flash 552B: B200 runs 19048 tok/s/chip at $0.03/M tokens, GB200 NVL72 runs 15745 at $0.03/M. B200 is 30% cheaper per token; B200 delivers 21% more tok/s/chip. (Numbers reflect the default agentic-traces · fp4 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)
| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | B200:71402.0GB200 NVL72:54859.0 | B200:37158.4GB200 NVL72:26900.3 | B200:19048.3GB200 NVL72:15745.3 |
| Cost ($/M tok) | B200:$0.007GB200 NVL72:$0.009 | B200:$0.013GB200 NVL72:$0.019 | B200:$0.025GB200 NVL72:$0.033 |
| tok/s/MW | B200:41755549GB200 NVL72:29336384 | B200:21730075GB200 NVL72:14385186 | B200:11139378GB200 NVL72:8419943 |
| Concurrency | B200:~47GB200 NVL72:~32 | B200:~20GB200 NVL72:~15 | B200:~10GB200 NVL72:~8 |
Agentic inference metrics from the AgentX scenario and fixed-sequence inference metrics across models, hardware configurations, and serving parameters.
No measurements to plot for this selection. Review the benchmark controls above or adjust quick filters.