AMD Instinct MI300X
Overview
The AMD Instinct MI300X is the CDNA 3 accelerator that made AMD a serious LLM inference vendor: 192 GB of HBM3 at 5.3 TB/s, more than double H100 memory, with 2,615 dense FP8 TFLOP/s. Eight chips connect in a full-mesh Infinity Fabric node with no switches, 448 GB/s of scale-up bandwidth per chip.
The memory advantage lets MI300X serve models in fewer chips or hold larger KV caches per replica, and its hourly pricing sits well below comparable NVIDIA parts. Software is the historical caveat: ROCm builds of vLLM and SGLang have closed much of the gap, and the public benchmark record tracks exactly how far.
Specifications
| Vendor | AMD |
|---|---|
| Architecture | CDNA 3 |
| Memory (usable) | 192 GB HBM3 |
| Memory bandwidth | 5.3 TB/s |
| FP4 dense TFLOP/s | Not supported |
| FP8 dense TFLOP/s | 2,615 |
| BF16 dense TFLOP/s | 1,307 |
| Scale-up interconnect | Infinity Fabric |
| Scale-up bandwidth per chip | 448 GB/s |
| Scale-up world size | 8 |
| Scale-up topology | Full Mesh |
| Scale-out network | RoCEv2 Ethernet |
| NIC | Pollara 400GbE |
| TDP per chip | 750 W |
| All-in power per chip | 1.39 kW |
| Hyperscaler $/chip/hr | $0.95 |
| Neocloud $/chip/hr | $1.16 |
| Retail $/chip/hr | $1.30 |
Source: $/chip/hr rate tiers from the SemiAnalysis AI Cloud TCO Model
How InferenceX benchmarks it
MI300X was one of the original InferenceMAX chips and still runs in the continuous sweep on ROCm vLLM and SGLang, giving one of the longest public software-progress curves of any accelerator: the same chip, measured for months, as kernels and schedulers improved.
Frequently asked questions
- How much does MI300X cost per hour in the cloud?
- The SemiAnalysis AI Cloud TCO model rates MI300X at about $0.95/hr at hyperscalers, $1.16/hr at neoclouds and $1.30/hr at the retail tier. InferenceX performance-per-dollar pages use these rates to turn measured throughput into $/M tokens.
- How much memory does MI300X have?
- MI300X has 192 GB of usable HBM3 per chip with 5.3 TB/s of memory bandwidth. A 8-chip Infinity Fabric domain pools 1,536 GB.
- What is the power consumption of MI300X?
- MI300X has a 750 W TDP per chip, and about 1.39 kW all-in per chip once the host CPU, NICs and cooling share are included. InferenceX uses the all-in figure for energy-per-token math.
- Does MI300X support FP4?
- No. MI300X tops out at FP8 with 2,615 dense TFLOP/s; FP4 serving requires a newer chip generation.
- How fast is MI300X for LLM inference?
- It depends on the model, framework, precision and interactivity target, so InferenceX publishes continuously refreshed throughput-versus-interactivity Pareto frontiers for MI300X instead of a single number. The live dashboard and compare pages show current results on every covered model.
See live benchmark results
Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:
Go deeper with the SemiAnalysis models
InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.
SemiAnalysis Accelerator & HBM Model
SKU-level AI accelerator shipments, pricing and specifications, from foundry wafer starts and HBM supply through customer-level installed base, quarterly with multi-year forecasts.
SemiAnalysis AI Cloud TCO Model
The source of the hourly rates on this page: all-in GPU cost of ownership built up from server capex, power, colocation and cost of capital, with rental price scenarios and a full cluster finance suite.