Run Any Model on Any GPU: Measured Results
Pick a model and a GPU: each page below answers how fast it runs, what it costs per million tokens, and which serving engine produced the number, from continuously re-benchmarked runs on real hardware.
DeepSeek V4 Pro
8 GPUs- DeepSeek V4 Pro on H200Measured throughput, latency & cost
- DeepSeek V4 Pro on B200Measured throughput, latency & cost
- DeepSeek V4 Pro on B300Measured throughput, latency & cost
- DeepSeek V4 Pro on GB200 NVL72Measured throughput, latency & cost
- DeepSeek V4 Pro on GB300 NVL72Measured throughput, latency & cost
- DeepSeek V4 Pro on MI300XMeasured throughput, latency & cost
- DeepSeek V4 Pro on MI325XMeasured throughput, latency & cost
- DeepSeek V4 Pro on MI355XMeasured throughput, latency & cost
DeepSeek V4.1 Flash
7 GPUs- DeepSeek V4.1 Flash on H100Measured throughput, latency & cost
- DeepSeek V4.1 Flash on H200Measured throughput, latency & cost
- DeepSeek V4.1 Flash on B200Measured throughput, latency & cost
- DeepSeek V4.1 Flash on B300Measured throughput, latency & cost
- DeepSeek V4.1 Flash on GB200 NVL72Measured throughput, latency & cost
- DeepSeek V4.1 Flash on GB300 NVL72Measured throughput, latency & cost
- DeepSeek V4.1 Flash on MI355XMeasured throughput, latency & cost
DeepSeek R1
9 GPUs- DeepSeek R1 on H100Measured throughput, latency & cost
- DeepSeek R1 on H200Measured throughput, latency & cost
- DeepSeek R1 on B200Measured throughput, latency & cost
- DeepSeek R1 on B300Measured throughput, latency & cost
- DeepSeek R1 on GB200 NVL72Measured throughput, latency & cost
- DeepSeek R1 on GB300 NVL72Measured throughput, latency & cost
- DeepSeek R1 on MI300XMeasured throughput, latency & cost
- DeepSeek R1 on MI325XMeasured throughput, latency & cost
- DeepSeek R1 on MI355XMeasured throughput, latency & cost
Kimi K3
6 GPUs- Kimi K3 on H200Measured throughput, latency & cost
- Kimi K3 on B200Measured throughput, latency & cost
- Kimi K3 on B300Measured throughput, latency & cost
- Kimi K3 on GB200 NVL72Measured throughput, latency & cost
- Kimi K3 on GB300 NVL72Measured throughput, latency & cost
- Kimi K3 on MI355XMeasured throughput, latency & cost
Kimi K2.6
8 GPUs- Kimi K2.6 on H200Measured throughput, latency & cost
- Kimi K2.6 on B200Measured throughput, latency & cost
- Kimi K2.6 on B300Measured throughput, latency & cost
- Kimi K2.6 on GB200 NVL72Measured throughput, latency & cost
- Kimi K2.6 on GB300 NVL72Measured throughput, latency & cost
- Kimi K2.6 on MI300XMeasured throughput, latency & cost
- Kimi K2.6 on MI325XMeasured throughput, latency & cost
- Kimi K2.6 on MI355XMeasured throughput, latency & cost
GLM-5
7 GPUs- GLM-5 on H200Measured throughput, latency & cost
- GLM-5 on B200Measured throughput, latency & cost
- GLM-5 on B300Measured throughput, latency & cost
- GLM-5 on GB200 NVL72Measured throughput, latency & cost
- GLM-5 on GB300 NVL72Measured throughput, latency & cost
- GLM-5 on MI325XMeasured throughput, latency & cost
- GLM-5 on MI355XMeasured throughput, latency & cost
GLM-5.3
7 GPUs- GLM-5.3 on H200Measured throughput, latency & cost
- GLM-5.3 on B200Measured throughput, latency & cost
- GLM-5.3 on B300Measured throughput, latency & cost
- GLM-5.3 on GB200 NVL72Measured throughput, latency & cost
- GLM-5.3 on GB300 NVL72Measured throughput, latency & cost
- GLM-5.3 on MI325XMeasured throughput, latency & cost
- GLM-5.3 on MI355XMeasured throughput, latency & cost
MiniMax M3
9 GPUs- MiniMax M3 on H100Measured throughput, latency & cost
- MiniMax M3 on H200Measured throughput, latency & cost
- MiniMax M3 on B200Measured throughput, latency & cost
- MiniMax M3 on B300Measured throughput, latency & cost
- MiniMax M3 on GB200 NVL72Measured throughput, latency & cost
- MiniMax M3 on GB300 NVL72Measured throughput, latency & cost
- MiniMax M3 on MI300XMeasured throughput, latency & cost
- MiniMax M3 on MI325XMeasured throughput, latency & cost
- MiniMax M3 on MI355XMeasured throughput, latency & cost
MiniMax M2.7
9 GPUs- MiniMax M2.7 on H100Measured throughput, latency & cost
- MiniMax M2.7 on H200Measured throughput, latency & cost
- MiniMax M2.7 on B200Measured throughput, latency & cost
- MiniMax M2.7 on B300Measured throughput, latency & cost
- MiniMax M2.7 on GB200 NVL72Measured throughput, latency & cost
- MiniMax M2.7 on GB300 NVL72Measured throughput, latency & cost
- MiniMax M2.7 on MI300XMeasured throughput, latency & cost
- MiniMax M2.7 on MI325XMeasured throughput, latency & cost
- MiniMax M2.7 on MI355XMeasured throughput, latency & cost
Qwen3.8-Flash-Next
4 GPUsQwen3.5
9 GPUs- Qwen3.5 on H100Measured throughput, latency & cost
- Qwen3.5 on H200Measured throughput, latency & cost
- Qwen3.5 on B200Measured throughput, latency & cost
- Qwen3.5 on B300Measured throughput, latency & cost
- Qwen3.5 on GB200 NVL72Measured throughput, latency & cost
- Qwen3.5 on GB300 NVL72Measured throughput, latency & cost
- Qwen3.5 on MI300XMeasured throughput, latency & cost
- Qwen3.5 on MI325XMeasured throughput, latency & cost
- Qwen3.5 on MI355XMeasured throughput, latency & cost
gpt-oss-120b
7 GPUs- gpt-oss-120b on H100Measured throughput, latency & cost
- gpt-oss-120b on H200Measured throughput, latency & cost
- gpt-oss-120b on B200Measured throughput, latency & cost
- gpt-oss-120b on GB200 NVL72Measured throughput, latency & cost
- gpt-oss-120b on MI300XMeasured throughput, latency & cost
- gpt-oss-120b on MI325XMeasured throughput, latency & cost
- gpt-oss-120b on MI355XMeasured throughput, latency & cost
Llama 3.3 70B
6 GPUs- Llama 3.3 70B on H100Measured throughput, latency & cost
- Llama 3.3 70B on H200Measured throughput, latency & cost
- Llama 3.3 70B on B200Measured throughput, latency & cost
- Llama 3.3 70B on MI300XMeasured throughput, latency & cost
- Llama 3.3 70B on MI325XMeasured throughput, latency & cost
- Llama 3.3 70B on MI355XMeasured throughput, latency & cost