All model and GPU pairings
DeepSeek R1NVIDIA Hopper

Running DeepSeek R1 on H100

Quick answer

DeepSeek R1 sustains 74.2 tokens/s per GPU on H100 at 50 tokens/s per user, which works out to $4.38 per million tokens at hyperscaler pricing, served by Dynamo SGLang. Fastest measured TTFT: 0.2 ms; fastest TPOT: 0.0 ms (each the best across all configs, not one run).

Benchmarked configs

72

Serving engines

Dynamo SGLang, Dynamo TRTLLM

Precisions

fp8

Run dates

2026-02-082026-02-13

Throughput at every interactivity target

Serving is a trade-off: push more concurrent users through a GPU and each user's tokens arrive slower. The ladder below reads the measured frontier at each per-user speed target on a single-turn chat workload (8k input / 1k output), using the best engine and precision at that point.

Per-user targetTokens/s per GPU$ / 1M tokensEnginePrecision
30 tok/s116$2.81Dynamo SGLangfp8
50 tok/s74.2$4.38Dynamo SGLangfp8
75 tok/s49.5$6.57Dynamo SGLangfp8
100 tok/s25.7$12.64Dynamo SGLangfp8
150 tok/s--Dynamo SGLangfp8
200 tok/s--Dynamo SGLangfp8

What serving actually costs

Converting the 50 tokens/s per user operating point to $ per million total tokens across rental pricing tiers from the SemiAnalysis AI Cloud TCO model.

Pricing tier$ / GPU / hr$ / 1M tokens
Hyperscaler$1.17$4.38
Neocloud$1.55$5.81
Retail$1.78$6.67

Frequently asked questions

How fast is DeepSeek R1 on H100?
At an interactivity target of 50 tokens/s per user on a single-turn chat workload (8k input / 1k output), H100 sustains 74.2 tokens/s per GPU serving DeepSeek R1 with Dynamo SGLang in FP8. Peak measured throughput across all configs is 1,376 tokens/s per GPU.
How much does it cost to serve DeepSeek R1 on H100?
$4.38 per million total tokens at hyperscaler $/GPU/hr pricing, at 50 tokens/s per user. Neocloud and retail rental tiers are tabulated above; slower interactivity targets lower the cost further.
Which serving engines run DeepSeek R1 on H100?
The runs behind this page used Dynamo SGLang, Dynamo TRTLLM in FP8, including disaggregated prefill and multi-node serving. Engines are rebuilt and re-benchmarked continuously, so the best config can change between visits.
How are these DeepSeek R1 numbers measured?
Every number is measured on real H100 hardware by the InferenceX fleet, sweeping concurrency on a single-turn chat workload (8k input / 1k output) to trace the throughput-versus-interactivity frontier; the newest run landed on 2026-02-13. The same derivation powers the InferenceX overview leaderboard.

Explore the data