All AI inference chips

B300 vs B200

Spec-sheet and pricing comparison of NVIDIA B300 and NVIDIA B200 with links to continuously measured LLM inference benchmarks on identical workloads.

Spec-sheet comparison

B300B200B300 / B200
Memory per chip268 GB HBM3e180 GB HBM3e1.49x
Memory bandwidth8 TB/s8 TB/s1.0x
Dense FP8 compute4,500 TFLOP/s4,500 TFLOP/s1.0x
Dense FP4 compute13,500 TFLOP/s9,000 TFLOP/s1.5x
TDP1,200 W1,000 W1.2x
Hourly rate (neocloud tier)$2.52/hr$2.07/hr1.22x
Scale-up world size8 chips8 chips1.0x

Ratios are spec-sheet values; see the live compare pages for measured deltas.

Frequently asked questions

Which has more memory, B300 or B200?
B300 offers 268 GB HBM3e per chip versus 180 GB HBM3e on B200 (1.49x the capacity).
How do B300 and B200 prices compare?
At the neocloud tier the SemiAnalysis TCO model rates B300 at $2.52/hr versus $2.07/hr for B200. Hourly price alone is misleading; the per-dollar compare pages divide measured throughput by these rates.
Is B300 faster than B200 for LLM inference?
On paper B300 has 1.0x the dense FP8 compute of B200, but delivered tokens per second depend on the model, framework, precision and interactivity target. InferenceX measures both chips daily on identical workloads; see the live compare pages for current results.

See live benchmark results

Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:

Go deeper with the SemiAnalysis models

InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.