DeepSeek V4 Pro 1.6T
8K/1K
Hyperscaler cost ↓ Lower is better Source: InferenceX & SemiAnalysis Market July 2026 AI Cloud TCO Model
| Model · Scenario | B200 · Reference | MI355X | B300 | GB200 NVL72 | GB300 NVL72 |
|---|---|---|---|---|---|
DeepSeek V4 Pro 1.6T8K/1K | TRTLLM · FP4 | — cannot reach @150 | $0.64214% cheaper than B200 Dynamo SGLang · FP4 | $1.13352% more expensive than B200 Dynamo SGLang · FP4 | $0.52530% cheaper than B200 Dynamo SGLang · FP4 |
DeepSeek V4 Pro 1.6TLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | — cannot reach @150 | — cannot reach @150 | — cannot reach @150 | — no data for this scenario | $0.107No B200 baseline to compare against Dynamo vLLM · FP4 |
Kimi K3 2.8TLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | — cannot reach @150 | — no data for this scenario | $0.229No B200 baseline to compare against vLLM · FP4 | — cannot reach @150 | — no data for this scenario |
MiniMax M3 428B8K/1K | vLLM · FP4 | $0.102About the same cost as B200 ATOM¹ · FP4 | $0.12321% more expensive than B200 vLLM · FP4 | $0.904791% more expensive than B200 Dynamo vLLM · FP8 · STP | $0.268164% more expensive than B200 Dynamo vLLM · FP8 |
MiniMax M3 428BLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | vLLM · FP4 | — no data for this scenario | $0.01911% cheaper than B200 vLLM · FP4 | — no data for this scenario | — no data for this scenario |
GLM5.2Long Context Multi-Turn Realistic Agentic Scenario (AgentX) | SGLang · FP4 | — no data for this scenario | $0.077About the same cost as B200 SGLang · FP4 | — no data for this scenario | — no data for this scenario |
Qwen3.5 397B8K/1K | TRTLLM · FP4 | $0.14781% more expensive than B200 SGLang · FP4 | $0.0888% more expensive than B200 SGLang · FP4 | $0.296263% more expensive than B200 Dynamo SGLang · FP8 | $0.06915% cheaper than B200 Dynamo TRTLLM · FP4 |
Qwen3.5 397BLong Context Multi-Turn Realistic Agentic Scenario (AgentX) | SGLang · FP4 | — no data for this scenario | $0.01559% more expensive than B200 SGLang · FP4 | — no data for this scenario | $0.01116% more expensive than B200 Dynamo SGLang · FP4 |
DeepSeek R1 0528 671BMaintenance8K/1K | SGLang · FP4 | $0.26619% more expensive than B200 MoRI SGLang · FP4 | $0.34554% more expensive than B200 Dynamo TRTLLM · FP4 | $0.14436% cheaper than B200 Dynamo TRTLLM · FP4 | $0.13838% cheaper than B200 Dynamo TRTLLM · FP4 |
Kimi K2.5/2.6/2.7-Code 1TDeprecated8K/1K | Dynamo TRTLLM · FP4 · STP | — cannot reach @150 | $1.44450% more expensive than B200 vLLM · FP4 · STP | $1.0418% more expensive than B200 Dynamo TRTLLM · FP4 · STP | $1.20925% more expensive than B200 Dynamo TRTLLM · FP4 · STP |
GLM5/5.1 744BDeprecated8K/1K | — no exact @150 result | — cannot reach @150 | — cannot reach @150 | $1.036No B200 baseline to compare against Dynamo TRTLLM · FP4 | $0.998No B200 baseline to compare against Dynamo TRTLLM · FP4 |
gpt-oss 120BDeprecated8K/1K | TRTLLM · FP4 · STP | $0.02731% more expensive than B200 ATOM¹ · FP4 · STP | — no data for this scenario | $0.0229% more expensive than B200 Dynamo TRTLLM · FP4 · STP | — no data for this scenario |
MiniMax M2.5/2.7 230BDeprecated8K/1K | vLLM · FP4 · STP | — cannot reach @150 | $0.27538% more expensive than B200 vLLM · FP4 · STP | $0.28040% more expensive than B200 Dynamo vLLM · FP4 · STP | $0.29749% more expensive than B200 Dynamo vLLM · FP4 · STP |
Llama 3.3 70B InstructDeprecated8K/1K | TRTLLM · FP8 · STP | — cannot reach @150 | — no data for this scenario | — no data for this scenario | — no data for this scenario |
8K/1K
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
8K/1K
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
8K/1K
Long Context Multi-Turn Realistic Agentic Scenario (AgentX)
8K/1K
8K/1K
8K/1K
8K/1K
8K/1K
8K/1K
— = no result. ∞ = B200 baseline unavailable.
If a chip does not have FP4 spec decoding available, the next best available configuration is used.