AI Chips for LLM Inference

Specs, cloud pricing and continuously measured inference benchmarks for every chip InferenceX covers. Each page joins the hardware data used by the live dashboard with the SemiAnalysis AI Cloud TCO rates.

Chip pages

NVIDIA H100 SXM

NVIDIA H100 specs, cloud pricing and live LLM inference benchmarks: 80 GB HBM3, 3.35 TB/s bandwidth, FP8 tensor cores, NVLink 4.0, measured daily on vLLM, SGLang and TensorRT-LLM.

NVIDIA H200 SXM

NVIDIA H200 specs, cloud pricing and live LLM inference benchmarks: 141 GB HBM3e, 4.8 TB/s bandwidth, FP8 tensor cores and NVLink 4.0, measured daily against Blackwell and AMD Instinct.

NVIDIA B200

NVIDIA B200 specs, cloud pricing and live LLM inference benchmarks: 180 GB HBM3e, 8 TB/s bandwidth, 9,000 dense FP4 TFLOP/s and NVLink 5.0, measured daily on vLLM, SGLang and TensorRT-LLM.

NVIDIA B300

NVIDIA B300 (Blackwell Ultra) specs, cloud pricing and live LLM inference benchmarks: 268 GB usable HBM3e, 13,500 dense FP4 TFLOP/s, 800 Gbit/s scale-out, measured daily against B200, GB300 NVL72 and MI355X.

NVIDIA GB200 NVL72

NVIDIA GB200 NVL72 specs, rack pricing and live LLM inference benchmarks: 72 Blackwell chips in one NVLink domain, 186 GB HBM3e per chip, 900 GB/s scale-up, measured daily on disaggregated vLLM, SGLang and Dynamo TRT-LLM.

NVIDIA GB300 NVL72

NVIDIA GB300 NVL72 specs, rack pricing and live LLM inference benchmarks: 72 Blackwell Ultra chips, 278 GB HBM3e each (20 TB per rack), 15,000 dense FP4 TFLOP/s per chip, measured daily on disaggregated serving stacks.

AMD Instinct MI300X

AMD Instinct MI300X specs, cloud pricing and live LLM inference benchmarks: 192 GB HBM3, 5.3 TB/s bandwidth, 2,615 dense FP8 TFLOP/s, full-mesh Infinity Fabric, measured daily on vLLM and SGLang with ROCm.

AMD Instinct MI325X

AMD Instinct MI325X specs, cloud pricing and live LLM inference benchmarks: 256 GB HBM3e, 6 TB/s bandwidth, 2,615 dense FP8 TFLOP/s, full-mesh Infinity Fabric, measured daily on ROCm vLLM and SGLang.

AMD Instinct MI355X

AMD Instinct MI355X specs, cloud pricing and live LLM inference benchmarks: 288 GB HBM3e, 8 TB/s bandwidth, 10,066 dense FP4 TFLOP/s, the first AMD chip with FP4, measured daily on ROCm vLLM, SGLang and ATOM.

Head-to-head comparisons