AI Inference Overview

Every active model across MI355X, B200, B300, GB200 and GB300 at a glance.

8K→1K · Single-turn · Output tok/s/GPU @75 tok/s/user · Speculative decode only · Best validated stack per platform

Database snapshot through Jul 22

  • DeepSeek V4 Pro 1.6T

    B200
    255TRTLLM · FP4Jun 12
    MI355X
    91ATOM¹ · FP4Jun 2

    −64% vs B200

    B300
    247SGLang · FP4Jun 11

    −3% vs B200

    At 100, B300 leads

    GB200 NVL72
    no exact @75 result
    GB300 NVL72
    no exact @75 result
    View details
  • Kimi K2.5/2.6/2.7-Code 1T

    B200
    standard decode only
    MI355X
    standard decode only
    B300
    standard decode only
    GB200 NVL72
    standard decode only
    GB300 NVL72
    standard decode only
    View details
  • MiniMax M3 428B

    B200
    956vLLM · FP4Jul 6
    MI355X
    728ATOM¹ · FP4Jul 3

    −24% vs B200

    B300
    972vLLM · FP4Jul 6

    +2% vs B200

    GB200 NVL72
    standard decode only
    GB300 NVL72
    standard decode only
    View details
  • GLM5.2

    B200
    no 8K/1K data
    MI355X
    no 8K/1K data
    B300
    no 8K/1K data
    GB200 NVL72
    no 8K/1K data
    GB300 NVL72
    no 8K/1K data
    View details
  • Qwen3.5 397B

    B200
    1,244TRTLLM · FP4Jun 24
    MI355X
    519SGLang · FP4Jul 16

    −58% vs B200

    B300
    1,315SGLang · FP4Jul 4

    +6% vs B200

    GB200 NVL72
    standard decode only
    GB300 NVL72
    standard decode only
    View details

Each cell shows the platform's best validated speculative-decode serving configuration for that model, labeled with its precision. Deltas against B200 compare only same-precision, same-release results — FP4 is never measured against FP8. Results compare complete serving stacks rather than isolated silicon.

∞ = no comparable result

Tier values interpolate each configuration’s official Pareto frontier — no extrapolation.