Chip Performance per Dollar

302 head-to-head cost-per-million-tokens comparisons across DeepSeekv4 Pro 0813 1.6T, DeepSeek R1, Kimi K3 2.8T, Kimi K2.5/K2.6/K2.7-Code 1T, GLM 5/5.1, GLM 5.3 744B, MiniMax M3 428B, MiniMax M2.5/M2.7, Qwen 3.5 397B-A17B, gpt-oss 120B, and Llama 3.3 70B. Performance normalized by owning-hyperscaler TCO — each page renders the cost-per-token chart and an interpolated dollars-per-million comparison table so you can pick the cheaper SKU at any target interactivity level.

Models with AgentX data open long-context, multi-turn trace replay results. Models not yet covered by AgentX open the controlled 8K→1K workload. Each card identifies its scenario.

DeepSeekv4 Pro 0813 1.6T

28 chip pairs with cost-per-token benchmark data on DeepSeekv4 Pro 0813 1.6T.

DeepSeek R1

36 chip pairs with cost-per-token benchmark data on DeepSeek R1.

Kimi K2.5/K2.6/K2.7-Code 1T

28 chip pairs with cost-per-token benchmark data on Kimi K2.5/K2.6/K2.7-Code 1T.

MiniMax M3 428B

36 chip pairs with cost-per-token benchmark data on MiniMax M3 428B.

MiniMax M2.5/M2.7

36 chip pairs with cost-per-token benchmark data on MiniMax M2.5/M2.7.

Qwen 3.5 397B-A17B

45 chip pairs with cost-per-token benchmark data on Qwen 3.5 397B-A17B.