AI Chips for LLM Inference
Specs, cloud pricing and continuously measured inference benchmarks for every chip InferenceX covers. Each page joins the hardware data used by the live dashboard with the SemiAnalysis AI Cloud TCO rates.
Chip pages
NVIDIA H100 SXM
NVIDIA H100 specs, cloud pricing and live LLM inference benchmarks: 80 GB HBM3, 3.35 TB/s bandwidth, FP8 tensor cores, NVLink 4.0, measured daily on vLLM, SGLang and TensorRT-LLM.
NVIDIA H200 SXM
NVIDIA H200 specs, cloud pricing and live LLM inference benchmarks: 141 GB HBM3e, 4.8 TB/s bandwidth, FP8 tensor cores and NVLink 4.0, measured daily against Blackwell and AMD Instinct.
NVIDIA B200
NVIDIA B200 specs, cloud pricing and live LLM inference benchmarks: 180 GB HBM3e, 8 TB/s bandwidth, 9,000 dense FP4 TFLOP/s and NVLink 5.0, measured daily on vLLM, SGLang and TensorRT-LLM.
NVIDIA B300
NVIDIA B300 (Blackwell Ultra) specs, cloud pricing and live LLM inference benchmarks: 268 GB usable HBM3e, 13,500 dense FP4 TFLOP/s, 800 Gbit/s scale-out, measured daily against B200, GB300 NVL72 and MI355X.
NVIDIA GB200 NVL72
NVIDIA GB200 NVL72 specs, rack pricing and live LLM inference benchmarks: 72 Blackwell chips in one NVLink domain, 186 GB HBM3e per chip, 900 GB/s scale-up, measured daily on disaggregated vLLM, SGLang and Dynamo TRT-LLM.
NVIDIA GB300 NVL72
NVIDIA GB300 NVL72 specs, rack pricing and live LLM inference benchmarks: 72 Blackwell Ultra chips, 278 GB HBM3e each (20 TB per rack), 15,000 dense FP4 TFLOP/s per chip, measured daily on disaggregated serving stacks.
AMD Instinct MI300X
AMD Instinct MI300X specs, cloud pricing and live LLM inference benchmarks: 192 GB HBM3, 5.3 TB/s bandwidth, 2,615 dense FP8 TFLOP/s, full-mesh Infinity Fabric, measured daily on vLLM and SGLang with ROCm.
AMD Instinct MI325X
AMD Instinct MI325X specs, cloud pricing and live LLM inference benchmarks: 256 GB HBM3e, 6 TB/s bandwidth, 2,615 dense FP8 TFLOP/s, full-mesh Infinity Fabric, measured daily on ROCm vLLM and SGLang.
AMD Instinct MI355X
AMD Instinct MI355X specs, cloud pricing and live LLM inference benchmarks: 288 GB HBM3e, 8 TB/s bandwidth, 10,066 dense FP4 TFLOP/s, the first AMD chip with FP4, measured daily on ROCm vLLM, SGLang and ATOM.