All model and GPU pairings
Llama 3.3 70BAMD CDNA 3

Running Llama 3.3 70B on MI300X

Quick answer

Llama 3.3 70B sustains 1,482 tokens/s per GPU on MI300X at 50 tokens/s per user, which works out to $0.18 per million tokens at hyperscaler pricing, served by vLLM. Fastest measured TTFT: 0.1 ms; fastest TPOT: 0.0 ms (each the best across all configs, not one run).

Benchmarked configs

60

Serving engines

vLLM

Precisions

fp8

Run dates

2025-10-292025-10-29

Throughput at every interactivity target

Serving is a trade-off: push more concurrent users through a GPU and each user's tokens arrive slower. The ladder below reads the measured frontier at each per-user speed target on a single-turn chat workload (8k input / 1k output), using the best engine and precision at that point.

Per-user targetTokens/s per GPU$ / 1M tokensEnginePrecision
30 tok/s1,985$0.13vLLMfp8
50 tok/s1,482$0.18vLLMfp8
75 tok/s718$0.37vLLMfp8
100 tok/s457$0.58vLLMfp8
150 tok/s--vLLMfp8
200 tok/s--vLLMfp8

What serving actually costs

Converting the 50 tokens/s per user operating point to $ per million total tokens across rental pricing tiers from the SemiAnalysis AI Cloud TCO model.

Pricing tier$ / GPU / hr$ / 1M tokens
Hyperscaler$0.95$0.18
Neocloud$1.16$0.22
Retail$1.30$0.24

Frequently asked questions

How fast is Llama 3.3 70B on MI300X?
At an interactivity target of 50 tokens/s per user on a single-turn chat workload (8k input / 1k output), MI300X sustains 1,482 tokens/s per GPU serving Llama 3.3 70B with vLLM in FP8. Peak measured throughput across all configs is 3,520 tokens/s per GPU.
How much does it cost to serve Llama 3.3 70B on MI300X?
$0.18 per million total tokens at hyperscaler $/GPU/hr pricing, at 50 tokens/s per user. Neocloud and retail rental tiers are tabulated above; slower interactivity targets lower the cost further.
Which serving engines run Llama 3.3 70B on MI300X?
The runs behind this page used vLLM in FP8. Engines are rebuilt and re-benchmarked continuously, so the best config can change between visits.
How are these Llama 3.3 70B numbers measured?
Every number is measured on real MI300X hardware by the InferenceX fleet, sweeping concurrency on a single-turn chat workload (8k input / 1k output) to trace the throughput-versus-interactivity frontier; the newest run landed on 2025-10-29. The same derivation powers the InferenceX overview leaderboard.

Explore the data