All AI inference chips
NVIDIA · Blackwell

NVIDIA B200

Overview

The NVIDIA B200 is the volume Blackwell datacenter chip: 180 GB of usable HBM3e at 8 TB/s, NVLink 5.0 at 900 GB/s per chip, and tensor cores that double Hopper throughput per clock while adding FP4. At 9,000 dense FP4 TFLOP/s and 4,500 FP8, a single B200 more than doubles H100 compute in half the memory-bandwidth-bound regimes that dominate LLM decode.

B200 is where NVFP4 serving became mainstream: frontier open models ship FP4 checkpoints that keep accuracy within noise of FP8 while nearly doubling throughput per chip. Its 1,000 W TDP and roughly 1.7 kW all-in power draw price it well above Hopper per hour, so performance per dollar, not peak TFLOP/s, decides the upgrade.

Specifications

VendorNVIDIA
ArchitectureBlackwell
Memory (usable)180 GB HBM3e
Memory bandwidth8 TB/s
FP4 dense TFLOP/s9,000
FP8 dense TFLOP/s4,500
BF16 dense TFLOP/s2,250
Scale-up interconnectNVLink 5.0
Scale-up bandwidth per chip900 GB/s
Scale-up world size8
Scale-up topologySwitched 2-rail Optimized
Scale-out networkgIB RoCEv2 Ethernet
NICConnectX-7 400GbE
TDP per chip1,000 W
All-in power per chip1.71 kW
Hyperscaler $/chip/hr$1.73
Neocloud $/chip/hr$2.07
Retail $/chip/hr$2.60

Source: $/chip/hr rate tiers from the SemiAnalysis AI Cloud TCO Model

How InferenceX benchmarks it

B200 is one of the most heavily benchmarked chips on InferenceX: daily fixed-sequence sweeps and AgentX agentic-coding traces on vLLM, SGLang, TensorRT-LLM and Dynamo disaggregated serving, with per-dollar and precision compare pages tracking NVFP4 versus FP8 on every covered model.

Frequently asked questions

How much does B200 cost per hour in the cloud?
The SemiAnalysis AI Cloud TCO model rates B200 at about $1.73/hr at hyperscalers, $2.07/hr at neoclouds and $2.60/hr at the retail tier. InferenceX performance-per-dollar pages use these rates to turn measured throughput into $/M tokens.
How much memory does B200 have?
B200 has 180 GB of usable HBM3e per chip with 8 TB/s of memory bandwidth. A 8-chip NVLink 5.0 domain pools 1,440 GB.
What is the power consumption of B200?
B200 has a 1,000 W TDP per chip, and about 1.71 kW all-in per chip once the host CPU, NICs and cooling share are included. InferenceX uses the all-in figure for energy-per-token math.
Does B200 support FP4?
Yes. B200 reaches 9,000 dense FP4 TFLOP/s (4,500 at FP8), and InferenceX tracks FP4-versus-FP8 serving accuracy and throughput on its precision compare pages.
How fast is B200 for LLM inference?
It depends on the model, framework, precision and interactivity target, so InferenceX publishes continuously refreshed throughput-versus-interactivity Pareto frontiers for B200 instead of a single number. The live dashboard and compare pages show current results on every covered model.

See live benchmark results

Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:

Go deeper with the SemiAnalysis models

InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.