All model and GPU pairings
DeepSeekv4 Pro 0813 1.6TAMD CDNA 3

Running DeepSeek V4 Pro on MI300X

Quick answer

DeepSeek V4 Pro runs on MI300X: 29 benchmarked configs so far. See the interactivity ladder below for measured operating points.

Benchmarked configs

29

Serving engines

vLLM

Precisions

fp8

Run dates

2026-07-162026-07-16

Throughput at every interactivity target

Serving is a trade-off: push more concurrent users through a GPU and each user's tokens arrive slower. The ladder below reads the measured frontier at each per-user speed target on a single-turn chat workload (8k input / 1k output), using the best engine and precision at that point.

Per-user targetTokens/s per GPU$ / 1M tokensEnginePrecision
30 tok/s--vLLMfp8
50 tok/s--vLLMfp8
75 tok/s--vLLMfp8
100 tok/s--vLLMfp8
150 tok/s--vLLMfp8
200 tok/s--vLLMfp8

Frequently asked questions

How fast is DeepSeek V4 Pro on MI300X?
The InferenceX fleet has 29 benchmarked configs for this pairing; see the interactivity ladder above for the operating points reached so far.
How much does it cost to serve DeepSeek V4 Pro on MI300X?
Cost per million tokens is derived from measured throughput and $/GPU/hr rates from the SemiAnalysis AI Cloud TCO model; it appears once this pairing reaches the primary interactivity tier.
Which serving engines run DeepSeek V4 Pro on MI300X?
The runs behind this page used vLLM in FP8. Engines are rebuilt and re-benchmarked continuously, so the best config can change between visits.
How are these DeepSeek V4 Pro numbers measured?
Every number is measured on real MI300X hardware by the InferenceX fleet, sweeping concurrency on a single-turn chat workload (8k input / 1k output) to trace the throughput-versus-interactivity frontier; the newest run landed on 2026-07-16. The same derivation powers the InferenceX overview leaderboard.

Explore the data