Inference Cost per Million Tokens

Hyperscaler cost ↓ Lower is better Source: InferenceX & SemiAnalysis Market July 2026 AI Cloud TCO Model

  • DeepSeek V4 Pro 1.6T

    8K/1K

    B200
    vLLM · FP4
    MI355X
    cannot reach @150
    B300
    $0.64233% cheaper than this platform’s Jul 12 result
    Dynamo SGLang · FP4
    Compare curves
    GB200 NVL72
    Dynamo SGLang · FP4
    GB300 NVL72
    Dynamo SGLang · FP4
  • DeepSeek V4 Pro 1.6T

    Long Context Multi-Turn Realistic Agentic Scenario (AgentX)

    B200
    cannot reach @150
    MI355X
    cannot reach @150
    B300
    cannot reach @150
    GB200 NVL72
    no data for this scenario
    GB300 NVL72
    Dynamo vLLM · FP4
  • Kimi K3 2.8T

    Long Context Multi-Turn Realistic Agentic Scenario (AgentX)

    B200
    cannot reach @150
    MI355X
    no data for this scenario
    B300
    vLLM · FP4
    GB200 NVL72
    cannot reach @150
    GB300 NVL72
    no data for this scenario
  • MiniMax M3 428B

    Long Context Multi-Turn Realistic Agentic Scenario (AgentX)

    B200
    vLLM · FP4
    MI355X
    no data for this scenario
    B300
    vLLM · FP4
    GB200 NVL72
    no data for this scenario
    GB300 NVL72
    no data for this scenario
  • GLM5.2

    Long Context Multi-Turn Realistic Agentic Scenario (AgentX)

    B200
    SGLang · FP4
    MI355X
    no data for this scenario
    B300
    SGLang · FP4
    GB200 NVL72
    no data for this scenario
    GB300 NVL72
    no data for this scenario
  • Qwen3.5 397B

    Long Context Multi-Turn Realistic Agentic Scenario (AgentX)

    B200
    SGLang · FP4
    MI355X
    no data for this scenario
    B300
    SGLang · FP4
    GB200 NVL72
    no data for this scenario
    GB300 NVL72
    Dynamo SGLang · FP4

Current cost and change versus the latest validated platform result 30–60 days earlier.

Platforms without a valid 30-day comparison show current cost only.

If a chip does not have FP4 spec decoding available, the next best available configuration is used.