NVIDIA GB300 NVL72
Overview
The NVIDIA GB300 NVL72 is the Blackwell Ultra rack: 72 chips with 278 GB of usable HBM3e each, roughly 20 TB of pooled memory per rack, and 15,000 dense FP4 TFLOP/s per chip on the same 72-way NVLink 5.0 domain as GB200 NVL72. FP8 and BF16 throughput carry over from GB200; the Ultra uplift is FP4 and memory.
The extra memory per chip compounds at rack scale: bigger frontier MoE models, longer contexts and larger KV working sets fit without spilling, and each chip runs a higher 1,400 W TDP to feed the larger stacks. GB300 NVL72 is the current top of the NVIDIA inference lineup ahead of Vera Rubin.
Specifications
| Vendor | NVIDIA |
|---|---|
| Architecture | Blackwell |
| Memory (usable) | 278 GB HBM3e |
| Memory bandwidth | 8 TB/s |
| FP4 dense TFLOP/s | 15,000 |
| FP8 dense TFLOP/s | 5,000 |
| BF16 dense TFLOP/s | 2,500 |
| Scale-up interconnect | NVLink 5.0 |
| Scale-up bandwidth per chip | 900 GB/s |
| Scale-up world size | 72 |
| Scale-up topology | Switched 18-rail Optimized |
| Scale-out network | N/A (NVLink domain only) |
| NIC | N/A (NVLink domain only) |
| TDP per chip | 1,400 W |
| All-in power per chip | 2.12 kW |
| Hyperscaler $/chip/hr | $2.31 |
| Neocloud $/chip/hr | $2.79 |
| Retail $/chip/hr | $3.30 |
Source: $/chip/hr rate tiers from the SemiAnalysis AI Cloud TCO Model
How InferenceX benchmarks it
InferenceX runs GB300 NVL72 on frontier MoE models with disaggregated and wide-EP configurations, and its AgentX agentic-coding lane uses the rack as the reference ceiling that AMD MI355X ATOM results and smaller NVIDIA nodes are measured against.
Frequently asked questions
- How much does GB300 NVL72 cost per hour in the cloud?
- The SemiAnalysis AI Cloud TCO model rates GB300 NVL72 at about $2.31/hr at hyperscalers, $2.79/hr at neoclouds and $3.30/hr at the retail tier. InferenceX performance-per-dollar pages use these rates to turn measured throughput into $/M tokens.
- How much memory does GB300 NVL72 have?
- GB300 NVL72 has 278 GB of usable HBM3e per chip with 8 TB/s of memory bandwidth. A 72-chip NVLink 5.0 domain pools 20,016 GB.
- What is the power consumption of GB300 NVL72?
- GB300 NVL72 has a 1,400 W TDP per chip, and about 2.12 kW all-in per chip once the host CPU, NICs and cooling share are included. InferenceX uses the all-in figure for energy-per-token math.
- Does GB300 NVL72 support FP4?
- Yes. GB300 NVL72 reaches 15,000 dense FP4 TFLOP/s (5,000 at FP8), and InferenceX tracks FP4-versus-FP8 serving accuracy and throughput on its precision compare pages.
- How fast is GB300 NVL72 for LLM inference?
- It depends on the model, framework, precision and interactivity target, so InferenceX publishes continuously refreshed throughput-versus-interactivity Pareto frontiers for GB300 NVL72 instead of a single number. The live dashboard and compare pages show current results on every covered model.
See live benchmark results
Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:
Go deeper with the SemiAnalysis models
InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.
SemiAnalysis Accelerator & HBM Model
SKU-level AI accelerator shipments, pricing and specifications, from foundry wafer starts and HBM supply through customer-level installed base, quarterly with multi-year forecasts.
SemiAnalysis AI Cloud TCO Model
The source of the hourly rates on this page: all-in GPU cost of ownership built up from server capex, power, colocation and cost of capital, with rental price scenarios and a full cluster finance suite.