NVIDIA H200 SXM
Overview
The NVIDIA H200 SXM is the memory-upgraded Hopper part: the same 1,979 dense FP8 TFLOP/s as H100, but with 141 GB of HBM3e at 4.8 TB/s. The 76% larger memory pool and 43% higher bandwidth matter most for long-context serving and large KV caches, where H100 runs out of headroom first.
For memory-bound decode workloads the H200 delivers materially better tokens per second per chip than H100 at a similar hourly rate, which is why it stayed the volume Hopper SKU once Blackwell arrived. Like H100 it has no FP4 support, so FP8 and INT4 weight-only quantization are its serving precisions.
Specifications
| Vendor | NVIDIA |
|---|---|
| Architecture | Hopper |
| Memory (usable) | 141 GB HBM3e |
| Memory bandwidth | 4.8 TB/s |
| FP4 dense TFLOP/s | Not supported |
| FP8 dense TFLOP/s | 1,979 |
| BF16 dense TFLOP/s | 989 |
| Scale-up interconnect | NVLink 4.0 |
| Scale-up bandwidth per chip | 450 GB/s |
| Scale-up world size | 8 |
| Scale-up topology | Switched 4-rail Optimized |
| Scale-out network | InfiniBand NDR |
| NIC | ConnectX-7 400G |
| TDP per chip | 700 W |
| All-in power per chip | 1.37 kW |
| Hyperscaler $/chip/hr | $1.22 |
| Neocloud $/chip/hr | $1.59 |
| Retail $/chip/hr | $2.05 |
Source: $/chip/hr rate tiers from the SemiAnalysis AI Cloud TCO Model
How InferenceX benchmarks it
InferenceX benchmarks H200 daily on vLLM, SGLang and TensorRT-LLM, and its compare pages line H200 FP8 results up against B200 NVFP4 and AMD MI325X so the Hopper-to-Blackwell and NVIDIA-to-AMD tradeoffs are measured rather than asserted.
Frequently asked questions
- How much does H200 cost per hour in the cloud?
- The SemiAnalysis AI Cloud TCO model rates H200 at about $1.22/hr at hyperscalers, $1.59/hr at neoclouds and $2.05/hr at the retail tier. InferenceX performance-per-dollar pages use these rates to turn measured throughput into $/M tokens.
- How much memory does H200 have?
- H200 has 141 GB of usable HBM3e per chip with 4.8 TB/s of memory bandwidth. A 8-chip NVLink 4.0 domain pools 1,128 GB.
- What is the power consumption of H200?
- H200 has a 700 W TDP per chip, and about 1.37 kW all-in per chip once the host CPU, NICs and cooling share are included. InferenceX uses the all-in figure for energy-per-token math.
- Does H200 support FP4?
- No. H200 tops out at FP8 with 1,979 dense TFLOP/s; FP4 serving requires a newer chip generation.
- How fast is H200 for LLM inference?
- It depends on the model, framework, precision and interactivity target, so InferenceX publishes continuously refreshed throughput-versus-interactivity Pareto frontiers for H200 instead of a single number. The live dashboard and compare pages show current results on every covered model.
See live benchmark results
Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:
Go deeper with the SemiAnalysis models
InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.
SemiAnalysis Accelerator & HBM Model
SKU-level AI accelerator shipments, pricing and specifications, from foundry wafer starts and HBM supply through customer-level installed base, quarterly with multi-year forecasts.
SemiAnalysis AI Cloud TCO Model
The source of the hourly rates on this page: all-in GPU cost of ownership built up from server capex, power, colocation and cost of capital, with rental price scenarios and a full cluster finance suite.