MiniMax M3 428B — B300 vs GB200 NVL72
Head-to-head AI inference benchmark comparison of B300 (NVIDIA Blackwell) and GB200 NVL72 (NVIDIA Blackwell) on MiniMax M3 428B. Latency, throughput, and cost across LLM workloads. Use the chart controls below to switch sequences, precisions, and metrics — same interactions as the main inference chart.
B300 posts 6892 tok/s/chip for $0.09 per million tokens at 66 tok/s/user on MiniMax M3 428B; GB200 NVL72 posts 3927 tok/s/chip for $0.13. B300 is 44% cheaper per token; B300 delivers 76% more tok/s/chip.
Throughput at 106 tok/s/user on MiniMax M3 428B: B300 hits 4403 tok/s/chip, GB200 NVL72 hits 1912. Per-million costs land at $0.14 and $0.27 respectively. B300 is 89% cheaper per token; B300 delivers 130% more tok/s/chip.
B300 / GB200 NVL72 on MiniMax M3 428B at 147 tok/s/user: 3218 / 811 tok/s/chip, $0.20 / $0.64 per million tokens. B300 is 226% cheaper per token; B300 delivers 297% more tok/s/chip. (Numbers reflect the default 8k/1k · fp8 selection for this URL — table and chart below update if you change sequence, precision, or model in the controls.)
| Metric | Interactivity (tok/s/user) | Interactivity (tok/s/user) | Interactivity (tok/s/user) |
|---|---|---|---|
| Throughput (tok/s/chip) | B300:6892.0GB200 NVL72:3926.9 | B300:4403.3GB200 NVL72:1912.4 | B300:3217.5GB200 NVL72:811.5 |
| Cost ($/M tok) | B300:$0.091GB200 NVL72:$0.132 | B300:$0.143GB200 NVL72:$0.270 | B300:$0.195GB200 NVL72:$0.637 |
| tok/s/MW | B300:3627393GB200 NVL72:2099968 | B300:2317538GB200 NVL72:1022690 | B300:1693422GB200 NVL72:433944 |
| Concurrency | B300:~50GB200 NVL72:~249 | B300:~21GB200 NVL72:~45 | B300:~11GB200 NVL72:~13 |
Inference Performance
Inference performance metrics across different models, hardware configurations, and serving parameters.