All AI inference chips
NVIDIA · Hopper

NVIDIA H200 SXM

Overview

The NVIDIA H200 SXM is the memory-upgraded Hopper part: the same 1,979 dense FP8 TFLOP/s as H100, but with 141 GB of HBM3e at 4.8 TB/s. The 76% larger memory pool and 43% higher bandwidth matter most for long-context serving and large KV caches, where H100 runs out of headroom first.

For memory-bound decode workloads the H200 delivers materially better tokens per second per chip than H100 at a similar hourly rate, which is why it stayed the volume Hopper SKU once Blackwell arrived. Like H100 it has no FP4 support, so FP8 and INT4 weight-only quantization are its serving precisions.

Specifications

VendorNVIDIA
ArchitectureHopper
Memory (usable)141 GB HBM3e
Memory bandwidth4.8 TB/s
FP4 dense TFLOP/sNot supported
FP8 dense TFLOP/s1,979
BF16 dense TFLOP/s989
Scale-up interconnectNVLink 4.0
Scale-up bandwidth per chip450 GB/s
Scale-up world size8
Scale-up topologySwitched 4-rail Optimized
Scale-out networkInfiniBand NDR
NICConnectX-7 400G
TDP per chip700 W
All-in power per chip1.37 kW
Hyperscaler $/chip/hr$1.22
Neocloud $/chip/hr$1.59
Retail $/chip/hr$2.05

Source: $/chip/hr rate tiers from the SemiAnalysis AI Cloud TCO Model

How InferenceX benchmarks it

InferenceX benchmarks H200 daily on vLLM, SGLang and TensorRT-LLM, and its compare pages line H200 FP8 results up against B200 NVFP4 and AMD MI325X so the Hopper-to-Blackwell and NVIDIA-to-AMD tradeoffs are measured rather than asserted.

Frequently asked questions

How much does H200 cost per hour in the cloud?
The SemiAnalysis AI Cloud TCO model rates H200 at about $1.22/hr at hyperscalers, $1.59/hr at neoclouds and $2.05/hr at the retail tier. InferenceX performance-per-dollar pages use these rates to turn measured throughput into $/M tokens.
How much memory does H200 have?
H200 has 141 GB of usable HBM3e per chip with 4.8 TB/s of memory bandwidth. A 8-chip NVLink 4.0 domain pools 1,128 GB.
What is the power consumption of H200?
H200 has a 700 W TDP per chip, and about 1.37 kW all-in per chip once the host CPU, NICs and cooling share are included. InferenceX uses the all-in figure for energy-per-token math.
Does H200 support FP4?
No. H200 tops out at FP8 with 1,979 dense TFLOP/s; FP4 serving requires a newer chip generation.
How fast is H200 for LLM inference?
It depends on the model, framework, precision and interactivity target, so InferenceX publishes continuously refreshed throughput-versus-interactivity Pareto frontiers for H200 instead of a single number. The live dashboard and compare pages show current results on every covered model.

See live benchmark results

Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:

Go deeper with the SemiAnalysis models

InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.