Head-to-head AI inference benchmark comparison of B200 (NVIDIA Blackwell) and H200 (NVIDIA Hopper) on Qwen 3.8 Flash Next 176B-A6B. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.
AgentX replays real coding-agent sessions rather than fixed-length prompts, so context grows turn over turn and most of each request is served from cache instead of being recomputed. That turns the comparison into a systems question: KV transfer between nodes, prefix-aware routing, and cache capacity all move the curve alongside raw chip throughput. Fixed-sequence workloads stay the clean baseline for kernel and silicon performance, so the two scenarios answer different questions about the same hardware. Learn more about AgentX →
At 43 tok/s/user on Qwen 3.8 Flash Next 176B-A6B, H200 delivers 16532 tok/s/chip at $0.02 per million tokens; B200 hasn't been benchmarked at this target.
H200 hits 15512 tok/s/chip for $0.02 per million tokens at 73 tok/s/user on Qwen 3.8 Flash Next 176B-A6B. No B200 data at this operating point.
H200: 13412 tok/s/chip, $0.03 per million tokens at 103 tok/s/user on Qwen 3.8 Flash Next 176B-A6B. B200 is unmeasured here. (Numbers reflect the default agentic-traces · fp8 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)
| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | B200:—H200:16532.0 | B200:—H200:15511.8 | B200:—H200:13411.7 |
| Cost ($/M tok) | B200:—H200:$0.020 | B200:—H200:$0.022 | B200:—H200:$0.025 |
| tok/s/MW | B200:—H200:12067129 | B200:—H200:11322490 | B200:—H200:9789569 |
| Concurrency | B200:—H200:~15 | B200:—H200:~13 | B200:—H200:~11 |
Agentic inference metrics from the AgentX scenario and fixed-sequence inference metrics across models, hardware configurations, and serving parameters.
No measurements to plot for this selection. Review the benchmark controls above or adjust quick filters.