GB200 NVL72 vs B200
Spec-sheet and pricing comparison of NVIDIA GB200 NVL72 and NVIDIA B200 with links to continuously measured LLM inference benchmarks on identical workloads.
Spec-sheet comparison
| GB200 NVL72 | B200 | GB200 NVL72 / B200 | |
|---|---|---|---|
| Memory per chip | 186 GB HBM3e | 180 GB HBM3e | 1.03x |
| Memory bandwidth | 8 TB/s | 8 TB/s | 1.0x |
| Dense FP8 compute | 5,000 TFLOP/s | 4,500 TFLOP/s | 1.11x |
| Dense FP4 compute | 10,000 TFLOP/s | 9,000 TFLOP/s | 1.11x |
| TDP | 1,200 W | 1,000 W | 1.2x |
| Hourly rate (neocloud tier) | $2.26/hr | $2.07/hr | 1.09x |
| Scale-up world size | 72 chips | 8 chips | 9.0x |
Ratios are spec-sheet values; see the live compare pages for measured deltas.
Frequently asked questions
- Which has more memory, GB200 NVL72 or B200?
- GB200 NVL72 offers 186 GB HBM3e per chip versus 180 GB HBM3e on B200 (1.03x the capacity).
- How do GB200 NVL72 and B200 prices compare?
- At the neocloud tier the SemiAnalysis TCO model rates GB200 NVL72 at $2.26/hr versus $2.07/hr for B200. Hourly price alone is misleading; the per-dollar compare pages divide measured throughput by these rates.
- Is GB200 NVL72 faster than B200 for LLM inference?
- On paper GB200 NVL72 has 1.11x the dense FP8 compute of B200, but delivered tokens per second depend on the model, framework, precision and interactivity target. InferenceX measures both chips daily on identical workloads; see the live compare pages for current results.
See live benchmark results
Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:
Go deeper with the SemiAnalysis models
InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.
SemiAnalysis Accelerator & HBM Model
SKU-level AI accelerator shipments, pricing and specifications, from foundry wafer starts and HBM supply through customer-level installed base, quarterly with multi-year forecasts.
SemiAnalysis AI Cloud TCO Model
The source of the hourly rates on this page: all-in GPU cost of ownership built up from server capex, power, colocation and cost of capital, with rental price scenarios and a full cluster finance suite.