Inference Dashboard / Model
Model Architectures
Architecture deep-dives for every model benchmarked on InferenceX: MoE and attention design, official vendor eval scores, and live inference performance data.
- DeepSeek V4 Pro→MoEHybrid1.6TDeepSeek's 1.6T-parameter / 49B-active MoE flagship with hybrid CSA+HCA sparse attention, a 1M-token context window, and MIT-licensed open weights.DeepSeek · Released April 24, 2026 (preview); August 13, 2026 (V4-Pro-0813 GA)
- DeepSeek R1 0528→MoEMLA671BDeepSeek's 671B-parameter / 37B-active MoE reasoning model with Multi-head Latent Attention, released under the MIT license.DeepSeek · Released May 28, 2025
- Kimi K3→MoEHybrid2.8TMoonshot AI's 2.8T-parameter / 104B-active MoE flagship built on Kimi Delta Attention and Gated MLA, with native vision and a 1M-token context window.Moonshot AI · Released July 16, 2026 (open weights by July 27, 2026)
- Kimi K2.5 / K2.6 / K2.7-Code→MoEMLA1.0TMoonshot AI's 1T-parameter / 32B-active MoE family sharing a DeepSeek-V3-style MLA backbone with native INT4 quantization and a 256K context window.Moonshot AI · Released January 27, 2026 (K2.5); April 20, 2026 (K2.6); June 12, 2026 (K2.7-Code)
- GLM-5 / GLM-5.1→Z.ai's 744B-parameter / 40B-active MoE family with DeepSeek Sparse Attention (DSA), trained on 28.5T tokens and released under the MIT license.Z.ai (Zhipu AI) · Released February 11, 2026 (GLM-5); April 7, 2026 (GLM-5.1)
- GLM-5.2 / GLM-5.3→Z.ai's 744B-class MoE line extending sparse attention with IndexShare for 1M-token contexts, with GLM-5.3 adding a major agentic post-training pass.Z.ai (Zhipu AI) · Released June 16, 2026 (GLM-5.2); August 14, 2026 (GLM-5.3)
- MiniMax M3→MoEGQA428BMiniMax's ~428B-parameter / ~23B-active multimodal MoE with GQA attention, a 1M-token context window, and 7 multi-token-prediction modules.MiniMax · Released June 1, 2026
- MiniMax M2.5 / M2.7→MoEGQA230BMiniMax's 229B-parameter / ~10B-active MoE family with GQA, FP8 block quantization, and a 192K context window.MiniMax · Released February 12, 2026 (M2.5); March 18, 2026 (M2.7)
- Qwen3.5-397B-A17B→Alibaba's 397B-parameter / 17B-active MoE with a 3:1 Gated DeltaNet to Gated Attention hybrid stack, 262K native context, and Apache 2.0 weights.Alibaba (Qwen) · Released February 2026
- gpt-oss-120b→MoESink/Full GQA120BOpenAI's 117B-parameter / 5.1B-active open-weight MoE with alternating dense and sliding-window attention and native MXFP4 quantization, under Apache 2.0.OpenAI · Released August 5, 2025
- Llama 3.3 70B Instruct→DenseGQA70BMeta's 70B dense instruction-tuned transformer with GQA and a 128K context window, released under the Llama 3.3 Community License.Meta · Released December 6, 2024