All AI inference chips

GB300 NVL72 vs GB200 NVL72

Spec-sheet and pricing comparison of NVIDIA GB300 NVL72 and NVIDIA GB200 NVL72 with links to continuously measured LLM inference benchmarks on identical workloads.

Spec-sheet comparison

GB300 NVL72GB200 NVL72GB300 NVL72 / GB200 NVL72
Memory per chip278 GB HBM3e186 GB HBM3e1.49x
Memory bandwidth8 TB/s8 TB/s1.0x
Dense FP8 compute5,000 TFLOP/s5,000 TFLOP/s1.0x
Dense FP4 compute15,000 TFLOP/s10,000 TFLOP/s1.5x
TDP1,400 W1,200 W1.17x
Hourly rate (neocloud tier)$2.79/hr$2.26/hr1.23x
Scale-up world size72 chips72 chips1.0x

Ratios are spec-sheet values; see the live compare pages for measured deltas.

Frequently asked questions

Which has more memory, GB300 NVL72 or GB200 NVL72?
GB300 NVL72 offers 278 GB HBM3e per chip versus 186 GB HBM3e on GB200 NVL72 (1.49x the capacity).
How do GB300 NVL72 and GB200 NVL72 prices compare?
At the neocloud tier the SemiAnalysis TCO model rates GB300 NVL72 at $2.79/hr versus $2.26/hr for GB200 NVL72. Hourly price alone is misleading; the per-dollar compare pages divide measured throughput by these rates.
Is GB300 NVL72 faster than GB200 NVL72 for LLM inference?
On paper GB300 NVL72 has 1.0x the dense FP8 compute of GB200 NVL72, but delivered tokens per second depend on the model, framework, precision and interactivity target. InferenceX measures both chips daily on identical workloads; see the live compare pages for current results.

See live benchmark results

Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:

Go deeper with the SemiAnalysis models

InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.