All model and GPU pairings
Qwen 3.5 397B-A17BNVIDIA Blackwell

Running Qwen3.5 on GB200 NVL72

Quick answer

Qwen3.5 runs on GB200 NVL72: 39 benchmarked configs so far. See the interactivity ladder below for measured operating points.

Benchmarked configs

39

Serving engines

Dynamo SGLang

Precisions

fp4, fp8

Run dates

2026-06-222026-08-25

Throughput at every interactivity target

Serving is a trade-off: push more concurrent users through a GPU and each user's tokens arrive slower. The ladder below reads the measured frontier at each per-user speed target on a single-turn chat workload (8k input / 1k output), using the best engine and precision at that point.

Per-user targetTokens/s per GPU$ / 1M tokensEnginePrecision
30 tok/s--Dynamo SGLangfp8
50 tok/s--Dynamo SGLangfp8
75 tok/s--Dynamo SGLangfp8
100 tok/s4,603$0.11Dynamo SGLangfp8
150 tok/s1,780$0.29Dynamo SGLangfp8
200 tok/s964$0.54Dynamo SGLangfp8

Frequently asked questions

How fast is Qwen3.5 on GB200 NVL72?
The InferenceX fleet has 39 benchmarked configs for this pairing; see the interactivity ladder above for the operating points reached so far.
How much does it cost to serve Qwen3.5 on GB200 NVL72?
Cost per million tokens is derived from measured throughput and $/GPU/hr rates from the SemiAnalysis AI Cloud TCO model; it appears once this pairing reaches the primary interactivity tier.
Which serving engines run Qwen3.5 on GB200 NVL72?
The runs behind this page used Dynamo SGLang in FP4, FP8, including disaggregated prefill and multi-node serving. Engines are rebuilt and re-benchmarked continuously, so the best config can change between visits.
How are these Qwen3.5 numbers measured?
Every number is measured on real GB200 NVL72 hardware by the InferenceX fleet, sweeping concurrency on a single-turn chat workload (8k input / 1k output) to trace the throughput-versus-interactivity frontier; the newest run landed on 2026-08-25. The same derivation powers the InferenceX overview leaderboard.

Explore the data