DeepSeek V4 Pro 1.6T
−64% vs B200
Every active model across MI355X, B200, B300, GB200 and GB300 at a glance.
8K→1K · Single-turn · Output tok/s/GPU @75 tok/s/user · Speculative decode only · Best validated stack per platform
Database snapshot through Jul 22
| Model | B200 | MI355X | B300 | GB200 NVL72 | GB300 NVL72 | Details |
|---|---|---|---|---|---|---|
DeepSeek V4 Pro 1.6T | −64% vs B200 | no exact @75 result | no exact @75 result | View details | ||
Kimi K2.5/2.6/2.7-Code 1T | standard decode only | standard decode only | standard decode only | standard decode only | standard decode only | View details |
MiniMax M3 428B | −24% vs B200 | +2% vs B200 | standard decode only | standard decode only | View details | |
GLM5.2 | no 8K/1K data | no 8K/1K data | no 8K/1K data | no 8K/1K data | no 8K/1K data | View details |
Qwen3.5 397B | −58% vs B200 | +6% vs B200 | standard decode only | standard decode only | View details |
−64% vs B200
−24% vs B200
+2% vs B200
−58% vs B200
+6% vs B200
Each cell shows the platform's best validated speculative-decode serving configuration for that model, labeled with its precision. Deltas against B200 compare only same-precision, same-release results — FP4 is never measured against FP8. Results compare complete serving stacks rather than isolated silicon.
∞ = no comparable result
Tier values interpolate each configuration’s official Pareto frontier — no extrapolation.