NVIDIA H100 SXM
Overview
The NVIDIA H100 SXM is the Hopper-generation datacenter accelerator that powered the first wave of large-scale LLM deployment. Each chip pairs 80 GB of HBM3 at 3.35 TB/s with fourth-generation tensor cores that reach 1,979 dense FP8 TFLOP/s, connected in 8-chip nodes over NVLink 4.0 at 450 GB/s per chip.
H100 remains the baseline other accelerators are judged against: it has the deepest software support of any AI chip, wide cloud availability, and the lowest hourly rates of the NVIDIA lineup. It lacks FP4 tensor cores, so newer Blackwell parts pull ahead on quantized serving workloads where NVFP4 applies.
Specifications
| Vendor | NVIDIA |
|---|---|
| Architecture | Hopper |
| Memory (usable) | 80 GB HBM3 |
| Memory bandwidth | 3.35 TB/s |
| FP4 dense TFLOP/s | Not supported |
| FP8 dense TFLOP/s | 1,979 |
| BF16 dense TFLOP/s | 989 |
| Scale-up interconnect | NVLink 4.0 |
| Scale-up bandwidth per chip | 450 GB/s |
| Scale-up world size | 8 |
| Scale-up topology | Switched 4-rail Optimized |
| Scale-out network | RoCEv2 Ethernet |
| NIC | ConnectX-7 2x200GbE |
| TDP per chip | 700 W |
| All-in power per chip | 1.37 kW |
| Hyperscaler $/chip/hr | $1.17 |
| Neocloud $/chip/hr | $1.55 |
| Retail $/chip/hr | $1.78 |
Source: $/chip/hr rate tiers from the SemiAnalysis AI Cloud TCO Model
How InferenceX benchmarks it
InferenceX runs H100 continuously on vLLM, SGLang and TensorRT-LLM across fixed-sequence serving and long-context agentic traces, publishing throughput-versus-interactivity Pareto frontiers, cost per million tokens and energy per token alongside every newer chip so the upgrade math stays visible.
Frequently asked questions
- How much does H100 cost per hour in the cloud?
- The SemiAnalysis AI Cloud TCO model rates H100 at about $1.17/hr at hyperscalers, $1.55/hr at neoclouds and $1.78/hr at the retail tier. InferenceX performance-per-dollar pages use these rates to turn measured throughput into $/M tokens.
- How much memory does H100 have?
- H100 has 80 GB of usable HBM3 per chip with 3.35 TB/s of memory bandwidth. A 8-chip NVLink 4.0 domain pools 640 GB.
- What is the power consumption of H100?
- H100 has a 700 W TDP per chip, and about 1.37 kW all-in per chip once the host CPU, NICs and cooling share are included. InferenceX uses the all-in figure for energy-per-token math.
- Does H100 support FP4?
- No. H100 tops out at FP8 with 1,979 dense TFLOP/s; FP4 serving requires a newer chip generation.
- How fast is H100 for LLM inference?
- It depends on the model, framework, precision and interactivity target, so InferenceX publishes continuously refreshed throughput-versus-interactivity Pareto frontiers for H100 instead of a single number. The live dashboard and compare pages show current results on every covered model.
See live benchmark results
Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:
Go deeper with the SemiAnalysis models
InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.
SemiAnalysis Accelerator & HBM Model
SKU-level AI accelerator shipments, pricing and specifications, from foundry wafer starts and HBM supply through customer-level installed base, quarterly with multi-year forecasts.
SemiAnalysis AI Cloud TCO Model
The source of the hourly rates on this page: all-in GPU cost of ownership built up from server capex, power, colocation and cost of capital, with rental price scenarios and a full cluster finance suite.