All AI inference chips
AMD · CDNA 3

AMD Instinct MI325X

Overview

The AMD Instinct MI325X is the memory-bumped CDNA 3 part: the same 2,615 dense FP8 TFLOP/s as MI300X but with 256 GB of HBM3e at 6 TB/s, the largest memory pool of its generation. An 8-chip full-mesh node carries over 2 TB of HBM, enough to serve very large models without leaving the node.

MI325X competes on capacity economics against H200: more memory per chip and lower hourly rates, traded against the NVIDIA software ecosystem. For memory-bound decode and long-context serving its bandwidth-per-dollar is among the best of the pre-CDNA 4 field.

Specifications

VendorAMD
ArchitectureCDNA 3
Memory (usable)256 GB HBM3e
Memory bandwidth6 TB/s
FP4 dense TFLOP/sNot supported
FP8 dense TFLOP/s2,615
BF16 dense TFLOP/s1,307
Scale-up interconnectInfinity Fabric
Scale-up bandwidth per chip448 GB/s
Scale-up world size8
Scale-up topologyFull Mesh
Scale-out networkRoCEv2 Ethernet
NICPollara 400GbE
TDP per chip1,000 W
All-in power per chip1.69 kW
Hyperscaler $/chip/hr$1.10
Neocloud $/chip/hr$1.32
Retail $/chip/hr$1.60

Source: $/chip/hr rate tiers from the SemiAnalysis AI Cloud TCO Model

How InferenceX benchmarks it

InferenceX benchmarks MI325X on ROCm vLLM and SGLang across the shared model set, with compare pages pairing it against H200 and MI355X so both the NVIDIA-versus-AMD and generation-over-generation deltas stay continuously measured.

Frequently asked questions

How much does MI325X cost per hour in the cloud?
The SemiAnalysis AI Cloud TCO model rates MI325X at about $1.10/hr at hyperscalers, $1.32/hr at neoclouds and $1.60/hr at the retail tier. InferenceX performance-per-dollar pages use these rates to turn measured throughput into $/M tokens.
How much memory does MI325X have?
MI325X has 256 GB of usable HBM3e per chip with 6 TB/s of memory bandwidth. A 8-chip Infinity Fabric domain pools 2,048 GB.
What is the power consumption of MI325X?
MI325X has a 1,000 W TDP per chip, and about 1.69 kW all-in per chip once the host CPU, NICs and cooling share are included. InferenceX uses the all-in figure for energy-per-token math.
Does MI325X support FP4?
No. MI325X tops out at FP8 with 2,615 dense TFLOP/s; FP4 serving requires a newer chip generation.
How fast is MI325X for LLM inference?
It depends on the model, framework, precision and interactivity target, so InferenceX publishes continuously refreshed throughput-versus-interactivity Pareto frontiers for MI325X instead of a single number. The live dashboard and compare pages show current results on every covered model.

See live benchmark results

Every number above is static hardware data. Delivered tokens per second, cost per million tokens and energy per token are measured continuously on the dashboard:

Go deeper with the SemiAnalysis models

InferenceX measures delivered inference performance. The SemiAnalysis institutional models cover the market behind these chips: who ships them, who buys them, and what they cost to own.