All model and GPU pairings
MiniMax M2.5/M2.7AMD CDNA 4

Running MiniMax M2.7 on MI355X

Quick answer

MiniMax M2.7 sustains 7,692 tokens/s per GPU on MI355X at 50 tokens/s per user, which works out to $0.054 per million tokens at hyperscaler pricing, served by vLLM. Fastest measured TTFT: 0.1 ms; fastest TPOT: 0.0 ms (each the best across all configs, not one run).

Benchmarked configs

172

Serving engines

ATOM¹, VLLM-DISAGG, vLLM

Precisions

fp4, fp8

Run dates

2026-03-282026-06-08

Throughput at every interactivity target

Serving is a trade-off: push more concurrent users through a GPU and each user's tokens arrive slower. The ladder below reads the measured frontier at each per-user speed target on a single-turn chat workload (8k input / 1k output), using the best engine and precision at that point.

Per-user targetTokens/s per GPU$ / 1M tokensEnginePrecision
30 tok/s5,699$0.073vLLMfp8
50 tok/s7,692$0.054vLLMfp4
75 tok/s4,861$0.086vLLMfp4
100 tok/s2,575$0.16vLLMfp4
150 tok/s--vLLMfp4
200 tok/s--vLLMfp4

What serving actually costs

Converting the 50 tokens/s per user operating point to $ per million total tokens across rental pricing tiers from the SemiAnalysis AI Cloud TCO model.

Pricing tier$ / GPU / hr$ / 1M tokens
Hyperscaler$1.50$0.054
Neocloud$2.09$0.075
Retail$2.10$0.076

Frequently asked questions

How fast is MiniMax M2.7 on MI355X?
At an interactivity target of 50 tokens/s per user on a single-turn chat workload (8k input / 1k output), MI355X sustains 7,692 tokens/s per GPU serving MiniMax M2.7 with vLLM in FP4. Peak measured throughput across all configs is 17,666 tokens/s per GPU.
How much does it cost to serve MiniMax M2.7 on MI355X?
$0.054 per million total tokens at hyperscaler $/GPU/hr pricing, at 50 tokens/s per user. Neocloud and retail rental tiers are tabulated above; slower interactivity targets lower the cost further.
Which serving engines run MiniMax M2.7 on MI355X?
The runs behind this page used ATOM¹, VLLM-DISAGG, vLLM in FP4, FP8, including disaggregated prefill and multi-node serving. Engines are rebuilt and re-benchmarked continuously, so the best config can change between visits.
How are these MiniMax M2.7 numbers measured?
Every number is measured on real MI355X hardware by the InferenceX fleet, sweeping concurrency on a single-turn chat workload (8k input / 1k output) to trace the throughput-versus-interactivity frontier; the newest run landed on 2026-06-08. The same derivation powers the InferenceX overview leaderboard.

Explore the data