Head-to-head AI inference benchmark comparison of B200 (NVIDIA Blackwell) and H100 (NVIDIA Hopper) on Qwen 3.8 Flash Next 176B-A6B. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.
AgentX replays real coding-agent sessions rather than fixed-length prompts, so context grows turn over turn and most of each request is served from cache instead of being recomputed. That turns the comparison into a systems question: KV transfer between nodes, prefix-aware routing, and cache capacity all move the curve alongside raw chip throughput. Fixed-sequence workloads stay the clean baseline for kernel and silicon performance, so the two scenarios answer different questions about the same hardware. Learn more about AgentX →
H100 hits 4980 tok/s/chip for $0.07 per million tokens at 46 tok/s/user on Qwen 3.8 Flash Next 176B-A6B. No B200 data at this operating point.
H100: 4944 tok/s/chip, $0.07 per million tokens at 77 tok/s/user on Qwen 3.8 Flash Next 176B-A6B. B200 is unmeasured here.
At 108 tok/s/user on Qwen 3.8 Flash Next 176B-A6B, H100 delivers 4908 tok/s/chip at $0.07 per million tokens; B200 hasn't been benchmarked at this target. (Numbers reflect the default agentic-traces · fp8 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)
| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | B200:—H100:4979.7 | B200:—H100:4944.2 | B200:—H100:4907.6 |
| Cost ($/M tok) | B200:—H100:$0.065 | B200:—H100:$0.066 | B200:—H100:$0.066 |
| tok/s/MW | B200:—H100:3634841 | B200:—H100:3608888 | B200:—H100:3582201 |
| Concurrency | B200:—H100:~11 | B200:—H100:~10 | B200:—H100:~9 |
Agentic inference metrics from the AgentX scenario and fixed-sequence inference metrics across models, hardware configurations, and serving parameters.
No measurements to plot for this selection. Review the benchmark controls above or adjust quick filters.