AMD Instinct MI355X
Overview
The AMD Instinct MI355X is the CDNA 4 flagship and the first AMD accelerator with FP4 tensor throughput: 10,066 dense FP4 and 5,033 FP8 TFLOP/s with 288 GB of HBM3e at 8 TB/s, the largest single-chip memory in the field. Eight chips form a full-mesh node over fifth-generation Infinity Fabric at 538 GB/s per chip.
MI355X matches B200-class bandwidth and exceeds its memory while pricing meaningfully lower per hour, making it the sharpest performance-per-dollar challenge AMD has mounted. Its 1,400 W TDP and the maturity of ROCm serving stacks are the tradeoffs the live data quantifies.
Specifications
| Vendor | AMD |
|---|---|
| Architecture | CDNA 4 |
| Memory (usable) | 288 GB HBM3e |
| Memory bandwidth | 8 TB/s |
| FP4 dense TFLOP/s | 10,066 |
| FP8 dense TFLOP/s | 5,033 |
| BF16 dense TFLOP/s | 2,516 |
| Scale-up interconnect | 5th Gen Infinity Fabric |
| Scale-up bandwidth per chip | 538 GB/s |
| Scale-up world size | 8 |
| Scale-up topology | Full Mesh |
| Scale-out network | RoCEv2 Ethernet |
| NIC | Pollara 400GbE |
| TDP per chip | 1,400 W |
| All-in power per chip | 2.09 kW |
| Hyperscaler $/chip/hr | $1.50 |
| Neocloud $/chip/hr | $2.09 |
| Retail $/chip/hr | $2.10 |
Source: $/chip/hr rate tiers from the SemiAnalysis AI Cloud TCO Model
How InferenceX benchmarks it
MI355X is the most intensively tracked AMD chip on InferenceX: daily vLLM, SGLang and ATOM runs, AgentX agentic-coding traces against GB300 NVL72 and B300, and some of the fastest published software-progress curves, including order-of-magnitude gains within weeks of a model release.
Frequently asked questions
- How much does MI355X cost per hour in the cloud?
- The SemiAnalysis AI Cloud TCO model rates MI355X at about $1.50/hr at hyperscalers, $2.09/hr at neoclouds and $2.10/hr at the retail tier. InferenceX performance-per-dollar pages use these rates to turn measured throughput into $/M tokens.
- How much memory does MI355X have?
- MI355X has 288 GB of usable HBM3e per chip with 8 TB/s of memory bandwidth. A 8-chip 5th Gen Infinity Fabric domain pools 2,304 GB.
- What is the power consumption of MI355X?
- MI355X has a 1,400 W TDP per chip, and about 2.09 kW all-in per chip once the host CPU, NICs and cooling share are included. InferenceX uses the all-in figure for energy-per-token math.
- Does MI355X support FP4?
- Yes. MI355X reaches 10,066 dense FP4 TFLOP/s (5,033 at FP8), and InferenceX tracks FP4-versus-FP8 serving accuracy and throughput on its precision compare pages.
- How fast is MI355X for LLM inference?
- It depends on the model, framework, precision and interactivity target, so InferenceX publishes continuously refreshed throughput-versus-interactivity Pareto frontiers for MI355X instead of a single number. The live dashboard and compare pages show current results on every covered model.
See live benchmark results
Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:
Go deeper with the SemiAnalysis models
InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.
SemiAnalysis Accelerator & HBM Model
SKU-level AI accelerator shipments, pricing and specifications, from foundry wafer starts and HBM supply through customer-level installed base, quarterly with multi-year forecasts.
SemiAnalysis AI Cloud TCO Model
The source of the hourly rates on this page: all-in GPU cost of ownership built up from server capex, power, colocation and cost of capital, with rental price scenarios and a full cluster finance suite.