NVIDIA B300
Overview
The NVIDIA B300 is the Blackwell Ultra SXM part: 268 GB of usable HBM3e at 8 TB/s and 13,500 dense FP4 TFLOP/s, a 1.5x FP4 uplift over B200 while FP8 and BF16 stay at B200 levels. The 49% larger memory pool is the bigger serving story, holding larger KV working sets and bigger models per 8-chip node.
B300 also doubles scale-out networking to 800 Gbit/s per chip via ConnectX-8, which matters for disaggregated prefill/decode and wide expert parallelism. It slots between B200 and the rack-scale GB300 NVL72: same NVLink 5.0 world size of 8, but Ultra-class memory and FP4 compute.
Specifications
| Vendor | NVIDIA |
|---|---|
| Architecture | Blackwell |
| Memory (usable) | 268 GB HBM3e |
| Memory bandwidth | 8 TB/s |
| FP4 dense TFLOP/s | 13,500 |
| FP8 dense TFLOP/s | 4,500 |
| BF16 dense TFLOP/s | 2,250 |
| Scale-up interconnect | NVLink 5.0 |
| Scale-up bandwidth per chip | 900 GB/s |
| Scale-up world size | 8 |
| Scale-up topology | Switched 2-rail Optimized |
| Scale-out network | RoCEv2 Ethernet |
| NIC | ConnectX-8 2x400GbE |
| TDP per chip | 1,200 W |
| All-in power per chip | 1.9 kW |
| Hyperscaler $/chip/hr | $2.26 |
| Neocloud $/chip/hr | $2.52 |
| Retail $/chip/hr | $3.00 |
Source: $/chip/hr rate tiers from the SemiAnalysis AI Cloud TCO Model
How InferenceX benchmarks it
InferenceX runs B300 on the same daily cadence as B200 across vLLM, SGLang and TensorRT-LLM, including AgentX long-context agentic traces where its 268 GB of HBM3e keeps KV working sets resident that force smaller chips to preempt or offload.
Frequently asked questions
- How much does B300 cost per hour in the cloud?
- The SemiAnalysis AI Cloud TCO model rates B300 at about $2.26/hr at hyperscalers, $2.52/hr at neoclouds and $3.00/hr at the retail tier. InferenceX performance-per-dollar pages use these rates to turn measured throughput into $/M tokens.
- How much memory does B300 have?
- B300 has 268 GB of usable HBM3e per chip with 8 TB/s of memory bandwidth. A 8-chip NVLink 5.0 domain pools 2,144 GB.
- What is the power consumption of B300?
- B300 has a 1,200 W TDP per chip, and about 1.9 kW all-in per chip once the host CPU, NICs and cooling share are included. InferenceX uses the all-in figure for energy-per-token math.
- Does B300 support FP4?
- Yes. B300 reaches 13,500 dense FP4 TFLOP/s (4,500 at FP8), and InferenceX tracks FP4-versus-FP8 serving accuracy and throughput on its precision compare pages.
- How fast is B300 for LLM inference?
- It depends on the model, framework, precision and interactivity target, so InferenceX publishes continuously refreshed throughput-versus-interactivity Pareto frontiers for B300 instead of a single number. The live dashboard and compare pages show current results on every covered model.
See live benchmark results
Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:
Go deeper with the SemiAnalysis models
InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.
SemiAnalysis Accelerator & HBM Model
SKU-level AI accelerator shipments, pricing and specifications, from foundry wafer starts and HBM supply through customer-level installed base, quarterly with multi-year forecasts.
SemiAnalysis AI Cloud TCO Model
The source of the hourly rates on this page: all-in GPU cost of ownership built up from server capex, power, colocation and cost of capital, with rental price scenarios and a full cluster finance suite.