Head-to-head AI inference benchmark comparison of B200 (NVIDIA Blackwell) and GB300 NVL72 (NVIDIA Blackwell) on Qwen 3.8 Flash Next 176B-A6B. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics β same interactions as the main inference chart.
AgentX replays real coding-agent sessions rather than fixed-length prompts, so context grows turn over turn and most of each request is served from cache instead of being recomputed. That turns the comparison into a systems question: KV transfer between nodes, prefix-aware routing, and cache capacity all move the curve alongside raw chip throughput. Fixed-sequence workloads stay the clean baseline for kernel and silicon performance, so the two scenarios answer different questions about the same hardware. Learn more about AgentX β
B200 / GB300 NVL72 on Qwen 3.8 Flash Next 176B-A6B at 99 tok/s/user: 57932 / 66781 tok/s/chip, $0.01 / $0.01 per million tokens. B200 is 16% cheaper per token; GB300 NVL72 delivers 15% more tok/s/chip.
Around the middle of the 53β239 tok/s/user interactivity band, at 146 tok/s/user on Qwen 3.8 Flash Next 176B-A6B: B200 runs 40327 tok/s/chip at $0.01/M tokens, GB300 NVL72 runs 36820 at $0.02/M. B200 is 46% cheaper per token; B200 delivers 10% more tok/s/chip.
Setting 192 tok/s/user as the target on Qwen 3.8 Flash Next 176B-A6B, B200 produces 20899 tok/s/chip ($0.02 per million tokens) and GB300 NVL72 produces 18080 ($0.04). B200 is 54% cheaper per token; B200 delivers 16% more tok/s/chip. (Numbers reflect the default agentic-traces Β· fp4 selection for this URL β table and chart below update if you change sequence, precision, or model in the controls.)
| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | B200:57931.6GB300 NVL72:66781.5 | B200:40326.8GB300 NVL72:36820.0 | B200:20899.1GB300 NVL72:18079.8 |
| Cost ($/M tok) | B200:$0.008GB300 NVL72:$0.010 | B200:$0.012GB300 NVL72:$0.017 | B200:$0.023GB300 NVL72:$0.035 |
| tok/s/MW | B200:33878132GB300 NVL72:31500685 | B200:23582896GB300 NVL72:17367942 | B200:12221693GB300 NVL72:8528194 |
| Concurrency | B200:~13GB300 NVL72:~13 | B200:~8GB300 NVL72:~7 | B200:~3GB300 NVL72:~2 |
Agentic inference metrics from the AgentX scenario and fixed-sequence inference metrics across models, hardware configurations, and serving parameters.
No measurements to plot for this selection. Review the benchmark controls above or adjust quick filters.