All AI inference chips

B200 vs H200

Spec-sheet and pricing comparison of NVIDIA B200 and NVIDIA H200 SXM with links to continuously measured LLM inference benchmarks on identical workloads.

Spec-sheet comparison

B200H200B200 / H200
Memory per chip180 GB HBM3e141 GB HBM3e1.28x
Memory bandwidth8 TB/s4.8 TB/s1.67x
Dense FP8 compute4,500 TFLOP/s1,979 TFLOP/s2.27x
Dense FP4 compute9,000 TFLOP/sNot supportedn/a
TDP1,000 W700 W1.43x
Hourly rate (neocloud tier)$2.07/hr$1.59/hr1.3x
Scale-up world size8 chips8 chips1.0x

Ratios are spec-sheet values; see the live compare pages for measured deltas.

Frequently asked questions

Which has more memory, B200 or H200?
B200 offers 180 GB HBM3e per chip versus 141 GB HBM3e on H200 (1.28x the capacity).
How do B200 and H200 prices compare?
At the neocloud tier the SemiAnalysis TCO model rates B200 at $2.07/hr versus $1.59/hr for H200. Hourly price alone is misleading; the per-dollar compare pages divide measured throughput by these rates.
Is B200 faster than H200 for LLM inference?
On paper B200 has 2.27x the dense FP8 compute of H200, but delivered tokens per second depend on the model, framework, precision and interactivity target. InferenceX measures both chips daily on identical workloads; see the live compare pages for current results.

See live benchmark results

Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:

Go deeper with the SemiAnalysis models

InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.