All model and GPU pairings
DeepSeek V4.1 Flash 552BNVIDIA Hopper

Running DeepSeek V4.1 Flash on H200

Quick answer

DeepSeek V4.1 Flash sustains 16,125 tokens/s per GPU on H200 at 50 tokens/s per user, which works out to $0.021 per million tokens at hyperscaler pricing, served by vLLM. Fastest measured TTFT: 0.3 ms; fastest TPOT: 0.0 ms (each the best across all configs, not one run).

Benchmarked configs

8

Serving engines

vLLM

Precisions

fp4

Run dates

2026-09-112026-09-11

Throughput at every interactivity target

Serving is a trade-off: push more concurrent users through a GPU and each user's tokens arrive slower. The ladder below reads the measured frontier at each per-user speed target on the AgentX agentic coding workload, using the best engine and precision at that point.

Per-user targetTokens/s per GPU$ / 1M tokensEnginePrecision
30 tok/s18,899$0.018vLLMfp4
50 tok/s16,125$0.021vLLMfp4
75 tok/s11,007$0.031vLLMfp4
100 tok/s8,444$0.040vLLMfp4
150 tok/s5,571$0.061vLLMfp4
200 tok/s4,184$0.081vLLMfp4

What serving actually costs

Converting the 50 tokens/s per user operating point to $ per million total tokens across rental pricing tiers from the SemiAnalysis AI Cloud TCO model.

Pricing tier$ / GPU / hr$ / 1M tokens
Owning at Large Hyperscaler Volume$1.22$0.021
Retail$2.90$0.050

Frequently asked questions

How fast is DeepSeek V4.1 Flash on H200?
At an interactivity target of 50 tokens/s per user on the AgentX agentic coding workload, H200 sustains 16,125 tokens/s per GPU serving DeepSeek V4.1 Flash with vLLM in FP4. Peak measured throughput across all configs is 19,341 tokens/s per GPU.
How much does it cost to serve DeepSeek V4.1 Flash on H200?
$0.021 per million total tokens at large-hyperscaler-volume ownership $/GPU/hr pricing, at 50 tokens/s per user. The retail rental tier is tabulated above; slower interactivity targets lower the cost further.
Which serving engines run DeepSeek V4.1 Flash on H200?
The runs behind this page used vLLM in FP4. Engines are rebuilt and re-benchmarked continuously, so the best config can change between visits.
How are these DeepSeek V4.1 Flash numbers measured?
Every number is measured on real H200 hardware by the InferenceX fleet, sweeping concurrency on the AgentX agentic coding workload to trace the throughput-versus-interactivity frontier; the newest run landed on 2026-09-11. The same derivation powers the InferenceX overview leaderboard.

Explore the data