DeepSeek V4 Pro 1.6T
8K/1K
Hyperscaler cost ↓ Lower is better Source: InferenceX & SemiAnalysis Market July 2026 AI Cloud TCO Model
| Model · Scenario | B200 · Reference | MI355X | B300 | GB200 NVL72 | GB300 NVL72 |
|---|---|---|---|---|---|
DeepSeek V4 Pro 1.6T8K/1K | vLLM · FP4 | — cannot reach @150 | $0.64250% cheaper than B200 Dynamo SGLang · FP4 | $1.13311% cheaper than B200 Dynamo SGLang · FP4 | $0.52559% cheaper than B200 Dynamo SGLang · FP4 |
DeepSeek V4 Pro 1.6TLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | — cannot reach @150 | — cannot reach @150 | — cannot reach @150 | — no data for this scenario | $0.107No B200 baseline to compare against Dynamo vLLM · FP4 |
Kimi K3 2.8TLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | — cannot reach @150 | — no data for this scenario | $0.229No B200 baseline to compare against vLLM · FP4 | — cannot reach @150 | — no data for this scenario |
MiniMax M3 428B8K/1K | vLLM · FP4 | $0.248145% more expensive than B200 vLLM · FP4 | $0.12321% more expensive than B200 vLLM · FP4 | $0.904791% more expensive than B200 Dynamo vLLM · FP8 · STP | $0.268164% more expensive than B200 Dynamo vLLM · FP8 |
MiniMax M3 428BLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | vLLM · FP4 | — no data for this scenario | $0.01911% cheaper than B200 vLLM · FP4 | — no data for this scenario | — no data for this scenario |
GLM5.2Long Context Multi-Turn Realistic Agentic Scenario (AgentX) | SGLang · FP4 | — no data for this scenario | $0.077About the same cost as B200 SGLang · FP4 | — no data for this scenario | — no data for this scenario |
Qwen3.5 397B8K/1K | SGLang · FP4 | $0.14779% more expensive than B200 SGLang · FP4 | $0.0887% more expensive than B200 SGLang · FP4 | $0.296260% more expensive than B200 Dynamo SGLang · FP8 | $0.335308% more expensive than B200 Dynamo SGLang · FP8 |
Qwen3.5 397BLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | SGLang · FP4 | — no data for this scenario | $0.01559% more expensive than B200 SGLang · FP4 | — no data for this scenario | $0.01116% more expensive than B200 Dynamo SGLang · FP4 |
DeepSeek R1 0528 671BMaintenance8K/1K | SGLang · FP4 | $0.26619% more expensive than B200 MoRI SGLang · FP4 | $0.728225% more expensive than B200 SGLang · FP8 | — cannot reach @150 | $3.1811,321% more expensive than B200 Dynamo SGLang · FP4 · STP |
Kimi K2.5/2.6/2.7-Code 1TDeprecated8K/1K | vLLM · FP4 · STP | — cannot reach @150 | $1.44446% more expensive than B200 vLLM · FP4 · STP | $1.77179% more expensive than B200 Dynamo vLLM · FP4 · STP | $2.218124% more expensive than B200 Dynamo vLLM · FP4 · STP |
GLM5/5.1 744BDeprecated8K/1K | — cannot reach @150 | — cannot reach @150 | — cannot reach @150 | $1.524No B200 baseline to compare against Dynamo SGLang · FP4 | $2.560No B200 baseline to compare against Dynamo SGLang · FP4 |
gpt-oss 120BDeprecated8K/1K | vLLM · FP4 · STP | $0.02928% more expensive than B200 vLLM · FP4 · STP | — no data for this scenario | — no data for this scenario | — no data for this scenario |
MiniMax M2.5/2.7 230BDeprecated8K/1K | vLLM · FP4 · STP | — cannot reach @150 | $0.27538% more expensive than B200 vLLM · FP4 · STP | $0.28040% more expensive than B200 Dynamo vLLM · FP4 · STP | $0.29749% more expensive than B200 Dynamo vLLM · FP4 · STP |
Llama 3.3 70B InstructDeprecated8K/1K | — cannot reach @150 | — cannot reach @150 | — no data for this scenario | — no data for this scenario | — no data for this scenario |
8K/1K
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
8K/1K
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
8K/1K
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
8K/1K
8K/1K
8K/1K
8K/1K
8K/1K
8K/1K
— = no result. ∞ = B200 baseline unavailable.
If a chip does not have FP4 spec decoding available, the next best available configuration is used.