AI Inference Overview

Every active model across MI355X, B200, B300, GB200 and GB300 at a glance.

8K→1K · Single-turn · Output tok/s/GPU @100 tok/s/user · Speculative decode only · Best validated stack per platform

Database snapshot through Jul 22

  • DeepSeek V4 Pro 1.6T

    B200
    117TRTLLM · FP4Jun 12
    MI355X
    57SGLang · FP4Jul 14

    −52% vs B200

    B300
    193SGLang · FP4Jun 11

    +64% vs B200

    GB200 NVL72
    no exact @100 result
    GB300 NVL72
    no exact @100 result
    View details
  • Kimi K2.5/2.6/2.7-Code 1T

    B200
    standard decode only
    MI355X
    standard decode only
    B300
    standard decode only
    GB200 NVL72
    standard decode only
    GB300 NVL72
    standard decode only
    View details
  • MiniMax M3 428B

    B200
    685vLLM · FP4Jul 6
    MI355X
    618ATOM¹ · FP4Jul 3

    −10% vs B200

    B300
    801vLLM · FP4Jul 6

    +17% vs B200

    GB200 NVL72
    standard decode only
    GB300 NVL72
    standard decode only
    View details
  • GLM5.2

    B200
    no 8K/1K data
    MI355X
    no 8K/1K data
    B300
    no 8K/1K data
    GB200 NVL72
    no 8K/1K data
    GB300 NVL72
    no 8K/1K data
    View details
  • Qwen3.5 397B

    B200
    855SGLang · FP4Jul 5
    MI355X
    452SGLang · FP4Jul 16

    −47% vs B200

    B300
    1,053SGLang · FP4Jul 4

    +23% vs B200

    GB200 NVL72
    standard decode only
    GB300 NVL72
    standard decode only
    View details

Each cell shows the platform's best validated speculative-decode serving configuration for that model, labeled with its precision. Deltas against B200 compare only same-precision, same-release results — FP4 is never measured against FP8. Results compare complete serving stacks rather than isolated silicon.

∞ = no comparable result

Tier values interpolate each configuration’s official Pareto frontier — no extrapolation.