Head-to-head AI inference benchmark comparison of GB300 NVL72 (NVIDIA Blackwell) and H100 (NVIDIA Hopper) on Qwen 3.8 Flash Next 176B-A6B. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics β same interactions as the main inference chart.
AgentX replays real coding-agent sessions rather than fixed-length prompts, so context grows turn over turn and most of each request is served from cache instead of being recomputed. That turns the comparison into a systems question: KV transfer between nodes, prefix-aware routing, and cache capacity all move the curve alongside raw chip throughput. Fixed-sequence workloads stay the clean baseline for kernel and silicon performance, so the two scenarios answer different questions about the same hardware. Learn more about AgentX β
GB300 NVL72: 86018 tok/s/chip, $0.01 per million tokens at 76 tok/s/user on Qwen 3.8 Flash Next 176B-A6B. H100 is unmeasured here.
At 131 tok/s/user on Qwen 3.8 Flash Next 176B-A6B, GB300 NVL72 delivers 48666 tok/s/chip at $0.01 per million tokens; H100 hasn't been benchmarked at this target.
GB300 NVL72 hits 18817 tok/s/chip for $0.03 per million tokens at 185 tok/s/user on Qwen 3.8 Flash Next 176B-A6B. No H100 data at this operating point. (Numbers reflect the default agentic-traces Β· fp4 selection for this URL β table and chart below update if you change sequence, precision, or model in the controls.)
| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | GB300 NVL72:86018.0H100:β | GB300 NVL72:48666.0H100:β | GB300 NVL72:18816.8H100:β |
| Cost ($/M tok) | GB300 NVL72:$0.007H100:β | GB300 NVL72:$0.013H100:β | GB300 NVL72:$0.034H100:β |
| tok/s/MW | GB300 NVL72:40574517H100:β | GB300 NVL72:22955682H100:β | GB300 NVL72:8875868H100:β |
| Concurrency | GB300 NVL72:~17H100:β | GB300 NVL72:~10H100:β | GB300 NVL72:~3H100:β |
Agentic inference metrics from the AgentX scenario and fixed-sequence inference metrics across models, hardware configurations, and serving parameters.
No measurements to plot for this selection. Review the benchmark controls above or adjust quick filters.