# InferenceX by SemiAnalysis > InferenceX is an open-source agentic inference benchmark dashboard. It compares the AgentX long-context, multi-turn coding scenario with fixed-sequence serving across NVIDIA, AMD, and other accelerators using public runs. ## Links - [Dashboard](https://inferencex.semianalysis.com) - [AgentX](https://inferencex.semianalysis.com/agentx) - [AgentX Methodology](https://inferencex.semianalysis.com/agentx/methodology) - [Articles](https://inferencex.semianalysis.com/blog) - [Whitepapers](https://inferencex.semianalysis.com/whitepaper) - [API Reference](https://inferencex.semianalysis.com/api) - [OpenAPI 3.1 Specification](https://inferencex.semianalysis.com/api/openapi.json) - [RSS Feed](https://inferencex.semianalysis.com/feed.xml) - [Full content for LLMs](https://inferencex.semianalysis.com/llms-full.txt) - [GitHub](https://github.com/SemiAnalysisAI/InferenceX) ## Model Benchmark Pages - [DeepSeek V4 Pro Inference Benchmarks](https://inferencex.semianalysis.com/inference/deepseek-v4): DeepSeekv4 Pro 0813 1.6T - [DeepSeek V4.1 Flash Inference Benchmarks](https://inferencex.semianalysis.com/inference/deepseek-v41-flash): DeepSeek V4.1 Flash 552B - [DeepSeek R1 Inference Benchmarks](https://inferencex.semianalysis.com/inference/deepseek-r1): DeepSeek R1 - [Kimi K3 Inference Benchmarks](https://inferencex.semianalysis.com/inference/kimi-k3): Kimi K3 2.8T - [Kimi K2.6 Inference Benchmarks](https://inferencex.semianalysis.com/inference/kimi-k26): Kimi K2.5/K2.6/K2.7-Code 1T - [GLM-5 Inference Benchmarks](https://inferencex.semianalysis.com/inference/glm-5-1): GLM 5/5.1 - [GLM-5.3 Inference Benchmarks](https://inferencex.semianalysis.com/inference/glm-5-3): GLM 5.3 744B - [MiniMax M3 Inference Benchmarks](https://inferencex.semianalysis.com/inference/minimax-m3): MiniMax M3 428B - [MiniMax M2.7 Inference Benchmarks](https://inferencex.semianalysis.com/inference/minimax-m27): MiniMax M2.5/M2.7 - [Qwen3.8-Flash-Next Inference Benchmarks](https://inferencex.semianalysis.com/inference/qwen-3-8-flash-next): Qwen 3.8 Flash Next 176B-A6B - [Qwen3.8-27B Inference Benchmarks](https://inferencex.semianalysis.com/inference/qwen-3-8-27b): Qwen 3.8 27B - [Qwen3.8-27B Eager Inference Benchmarks](https://inferencex.semianalysis.com/inference/qwen-3-8-27b-eager): Qwen 3.8 27B (eager) - [Qwen3.5 Inference Benchmarks](https://inferencex.semianalysis.com/inference/qwen-3-5): Qwen 3.5 397B-A17B - [gpt-oss-120b Inference Benchmarks](https://inferencex.semianalysis.com/inference/gptoss-120b): gpt-oss 120B - [Llama 3.3 70B Inference Benchmarks](https://inferencex.semianalysis.com/inference/llama-3-3-70b): Llama 3.3 70B ## Chip Pages - [NVIDIA H100 SXM](https://inferencex.semianalysis.com/chips/h100): specs, pricing and benchmarks - [NVIDIA H200 SXM](https://inferencex.semianalysis.com/chips/h200): specs, pricing and benchmarks - [NVIDIA B200](https://inferencex.semianalysis.com/chips/b200): specs, pricing and benchmarks - [NVIDIA B300](https://inferencex.semianalysis.com/chips/b300): specs, pricing and benchmarks - [NVIDIA GB200 NVL72](https://inferencex.semianalysis.com/chips/gb200-nvl72): specs, pricing and benchmarks - [NVIDIA GB300 NVL72](https://inferencex.semianalysis.com/chips/gb300-nvl72): specs, pricing and benchmarks - [AMD Instinct MI300X](https://inferencex.semianalysis.com/chips/mi300x): specs, pricing and benchmarks - [AMD Instinct MI325X](https://inferencex.semianalysis.com/chips/mi325x): specs, pricing and benchmarks - [AMD Instinct MI355X](https://inferencex.semianalysis.com/chips/mi355x): specs, pricing and benchmarks - [B200 vs H100](https://inferencex.semianalysis.com/chips/b200-vs-h100): head-to-head comparison - [B200 vs H200](https://inferencex.semianalysis.com/chips/b200-vs-h200): head-to-head comparison - [B300 vs B200](https://inferencex.semianalysis.com/chips/b300-vs-b200): head-to-head comparison - [GB200 NVL72 vs B200](https://inferencex.semianalysis.com/chips/gb200-nvl72-vs-b200): head-to-head comparison - [GB300 NVL72 vs GB200 NVL72](https://inferencex.semianalysis.com/chips/gb300-nvl72-vs-gb200-nvl72): head-to-head comparison - [GB300 NVL72 vs B300](https://inferencex.semianalysis.com/chips/gb300-nvl72-vs-b300): head-to-head comparison - [MI355X vs MI325X](https://inferencex.semianalysis.com/chips/mi355x-vs-mi325x): head-to-head comparison - [MI325X vs MI300X](https://inferencex.semianalysis.com/chips/mi325x-vs-mi300x): head-to-head comparison - [MI300X vs H100](https://inferencex.semianalysis.com/chips/mi300x-vs-h100): head-to-head comparison - [MI355X vs B200](https://inferencex.semianalysis.com/chips/mi355x-vs-b200): head-to-head comparison - [MI355X vs B300](https://inferencex.semianalysis.com/chips/mi355x-vs-b300): head-to-head comparison - [MI325X vs H200](https://inferencex.semianalysis.com/chips/mi325x-vs-h200): head-to-head comparison ## GPU Rankings - [GPU Rankings for LLM Inference](https://inferencex.semianalysis.com/rankings): live fastest and cheapest GPU leaderboards per model - [Fastest GPU for DeepSeek V4 Pro](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-deepseek-v4): live benchmark ranking - [Fastest GPU for DeepSeek V4.1 Flash](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-deepseek-v41-flash): live benchmark ranking - [Fastest GPU for DeepSeek R1](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-deepseek-r1): live benchmark ranking - [Fastest GPU for Kimi K3](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-kimi-k3): live benchmark ranking - [Fastest GPU for Kimi K2.6](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-kimi-k26): live benchmark ranking - [Fastest GPU for GLM-5](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-glm-5-1): live benchmark ranking - [Fastest GPU for GLM-5.3](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-glm-5-3): live benchmark ranking - [Fastest GPU for MiniMax M3](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-minimax-m3): live benchmark ranking - [Fastest GPU for MiniMax M2.7](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-minimax-m27): live benchmark ranking - [Fastest GPU for Qwen3.8-Flash-Next](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-qwen-3-8-flash-next): live benchmark ranking - [Fastest GPU for Qwen3.8-27B](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-qwen-3-8-27b): live benchmark ranking - [Fastest GPU for Qwen3.8-27B Eager](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-qwen-3-8-27b-eager): live benchmark ranking - [Fastest GPU for Qwen3.5](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-qwen-3-5): live benchmark ranking - [Fastest GPU for gpt-oss-120b](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-gptoss-120b): live benchmark ranking - [Fastest GPU for Llama 3.3 70B](https://inferencex.semianalysis.com/rankings/fastest-gpu-for-llama-3-3-70b): live benchmark ranking - [Cheapest GPU for DeepSeek V4 Pro](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-deepseek-v4): live benchmark ranking - [Cheapest GPU for DeepSeek V4.1 Flash](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-deepseek-v41-flash): live benchmark ranking - [Cheapest GPU for DeepSeek R1](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-deepseek-r1): live benchmark ranking - [Cheapest GPU for Kimi K3](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-kimi-k3): live benchmark ranking - [Cheapest GPU for Kimi K2.6](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-kimi-k26): live benchmark ranking - [Cheapest GPU for GLM-5](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-glm-5-1): live benchmark ranking - [Cheapest GPU for GLM-5.3](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-glm-5-3): live benchmark ranking - [Cheapest GPU for MiniMax M3](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-minimax-m3): live benchmark ranking - [Cheapest GPU for MiniMax M2.7](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-minimax-m27): live benchmark ranking - [Cheapest GPU for Qwen3.8-Flash-Next](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-qwen-3-8-flash-next): live benchmark ranking - [Cheapest GPU for Qwen3.8-27B](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-qwen-3-8-27b): live benchmark ranking - [Cheapest GPU for Qwen3.8-27B Eager](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-qwen-3-8-27b-eager): live benchmark ranking - [Cheapest GPU for Qwen3.5](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-qwen-3-5): live benchmark ranking - [Cheapest GPU for gpt-oss-120b](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-gptoss-120b): live benchmark ranking - [Cheapest GPU for Llama 3.3 70B](https://inferencex.semianalysis.com/rankings/cheapest-gpu-for-llama-3-3-70b): live benchmark ranking ## Model on GPU Results - [Run Any Model on Any GPU](https://inferencex.semianalysis.com/run): measured throughput and cost for every benchmarked model and GPU pairing ## Articles - [Engrams Embedding Entendre: Codesign for Efficient DRAM/SSD Offloading](https://inferencex.semianalysis.com/blog/engrams-embedding-entendre-codesign): New Model Architecture Implications for TAM of DRAM/NVMe, DeepSeek V4.1 Flash, AgentX, InferenceX, NVMe experiments - [Rubin NVL72 Agentic Inference: 67x better Performance per Dollar](https://inferencex.semianalysis.com/blog/vera-rubin-nvl72-agentic-inference): Jensen Sandbagging Performance Again, 2x more Annual Profit Per GigaWatt, The More you Buy, The More you Earn, AgentX, InferenceX, Extreme Co-Design - [TPU Inference Externalization Full Steam Ahead](https://inferencex.semianalysis.com/blog/tpu-inferencex-full-steam): InferenceX, Up to 50% Better Performance per Dollar, Rapid Externalization of TPU stack, Growing Customer Base, Ironwood, TPUv8i, Reducing CUDA Moat - [DeepSeek V4 Pro on AgentX: B200 vs B300 and the KV Cache Working Set](https://inferencex.semianalysis.com/blog/deepseek-v4-pro-agentx-b200-vs-b300-kv-working-set): 50% more HBM squeezes out extra throughput, and the per-point telemetry shows exactly where it comes from - [DeepSeek V4 Pro on AgentX: GB200 vs GB300 Rack-Scale Disaggregation](https://inferencex.semianalysis.com/blog/deepseek-v4-pro-agentx-gb200-vs-gb300-disagg): Both lean on PD disagg, GB300 adds DEP32 wide-EP decode, and the gap shows up in first-token latency rather than token rate - [DeepSeek V4 Pro on AgentX: MI355X vs B200, and the August 21 Flip](https://inferencex.semianalysis.com/blog/deepseek-v4-pro-agentx-mi355x-vs-b200-august): AMD matched B200 vLLM on performance per dollar for end-to-end latency, then upstream vLLM work moved the line - [GLM 5.3 on AgentX: MI355X ATOM Beats GB300 NVL72 on Part of the Curve](https://inferencex.semianalysis.com/blog/glm-5-3-agentx-mi355x-atom-vs-gb300-nvl72): Where AMD’s vendor engine wins on performance per dollar, and what E2E Normalized Interactivity actually measures - [GLM 5.3 on AgentX: NVIDIA Is Up to 5x Cheaper per Token at 150 tok/s/user](https://inferencex.semianalysis.com/blog/glm-5-3-agentx-nvidia-vs-amd-sglang-150-toks): At this operating point, free AMD silicon would still not close the gap - [Kimi K3 on AgentX: MI355X ATOM Beats GB300 NVL72 on Part of the Curve](https://inferencex.semianalysis.com/blog/kimi-k3-agentx-mi355x-atom-vs-gb300-nvl72): AMD’s vendor engine wins a real slice of the performance per dollar frontier, while Hopper struggles to serve K3 at all - [MiniMax M3 on AgentX: Why B200 and B300 Beat Their Rack-Scale GB200 NVL72 & GB300 NVL72 Counterparts](https://inferencex.semianalysis.com/blog/minimax-m3-agentx-b200-b300-vs-rack-scale): The Dynamo router becomes the bottleneck, no submission runs context parallelism, and AMD leaves KV offload on the table - [MiniMax M3 on AgentX: B300 TRT-LLM TP2 Owns the Crown](https://inferencex.semianalysis.com/blog/minimax-m3-agentx-b300-trtllm-tp2-vs-the-field): NVIDIA sweeps the 432B model, and the missing DP-attention points explain why cache locality became a routing constraint - [OpenAI Jalapeño: Better Than Nvidia Blackwell](https://inferencex.semianalysis.com/blog/openai-jalapeno-better-than-nvidia): OpenAI’s self-designed ASIC compared with Rubin, Jalapeño’s TCO, throughput per MW, and spicy deets - [Qwen3.5 397B on AgentX: B300 FP4 Delivers 12x the Performance per Dollar of H100](https://inferencex.semianalysis.com/blog/qwen3-5-397b-agentx-b300-fp4-vs-h100): What four years of hardware and a 4-bit format buy on a long-context agentic workload - [MI355X versus GB300 NVL72 Inference Performance: 20x Gap on Qwen3.5 SGLang](https://inferencex.semianalysis.com/blog/qwen3-5-397b-agentx-nvidia-vs-amd-sglang): GatedDeltaNet, a 262k native context, and no AMD competition at all on the same engine - [AgentX - InferenceXv3: Does the CUDA Moat Hold Up in Agentic Inferencing?](https://inferencex.semianalysis.com/blog/agentx-inferencexv3-does-cuda-moat): $3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355X, B200 - [Agentic Benchmark for LLM Inference: Metrics and Methodology](https://inferencex.semianalysis.com/blog/agentic-benchmark-agent-benchmark-guide): How an agent benchmark replays long-context, multi-turn workloads to measure latency, throughput, cache behavior, and serving cost - [A Brief Overview of Agentic Workloads](https://inferencex.semianalysis.com/blog/brief-overview-of-agentic-workloads): Multi-turn sessions, long contexts, and near-total prefix reuse make agentic inference a systems problem, and change what a benchmark has to measure - [Ultra-High Interactivity on NVIDIA GPUs? TileRT on InferenceX](https://inferencex.semianalysis.com/blog/ultra-high-interactivity-on-nvidia): Can TileRT software on NVIDIA GPUs compete with Cerebras, Groq LPU, and SambaNova? Batch size 1, disaggregated engine, high-throughput prefill engine, high-interactivity decode engine - [Kimi K3: The Manos, The Mythos, The Legendos](https://inferencex.semianalysis.com/blog/kimi-k3-the-manos-the-mythos-the): Kimi K3's architecture: compressed memory, attention across depth, latent expert routing, and serving performance - [Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis](https://inferencex.semianalysis.com/blog/vera-rubin-nvl72-vs-gb200-nvl72-inference): Rubin LUT Based Tensor Core, Feynman, Rack Scale, Perf Per MegaWatt, Perf Per Dollar, Software Improvements, Public Rubin Software, PyTorch, vLLM, OpenAI Triton - [DeepSeekV4 1.6T Day 0 to Day 43 Performance Over Time — Huawei, GB300 NVL72, MI355X, B200](https://inferencex.semianalysis.com/blog/deepseekv4-16t-day-0-to-day-43-performance): Day 0 Inference Performance, InferenceX, 100x performance improvement in 26 Days, Cost per Million Tokens, Huawei 950DT Inference Trace Analysis - [GB300 NVL72 vs GB200 NVL72 Inference Performance & Perf per Dollar - on DeepSeek-V4-Pro 1.6T: Up to 2.83x Throughput](https://inferencex.semianalysis.com/blog/gb300-nvl72-vs-gb200-nvl72-dsv4-pro-vllm-fp4): DSv4-Pro FP4 8K/1K, Dynamo+vLLM, disaggregated on both racks. GB300's 50% extra HBM (288 vs 192 GB/GPU) unlocks a wider prefill+decode recipe GB200 can't fit — lifting middle-of-curve perf/$ by 2.31x despite a 20% per-GPU TCO premium. - [B200 NVFP4 vs H200 FP8 on GLM-5: Up to 3.65x Better Performance per Dollar with SGLang MTP](https://inferencex.semianalysis.com/blog/b200-glm5-nvfp4-vs-h200-fp8-3-6x-perf-per-dollar): Both SKUs run SGLang EAGLE MTP; the Blackwell generation lifts perf/$ by ~1.2x at the peak and the NVIDIA GLM-5-NVFP4 checkpoint on FlashInfer TRT-LLM sparse MLA stacks another ~2.4–3.0x on 8K/1K - [B200 NVFP4 vs H100 FP8 on MiniMax-M2.5: Up to 8.2x Better Performance per Dollar with vLLM](https://inferencex.semianalysis.com/blog/b200-minimax-m2-5-vllm-nvfp4-vs-h100-fp8-perf-per-dollar): vLLM PR #36307 unlocks the trtllm-gen FP8 MoE kernel for MiniMax on B200; combined with NVFP4, perf/$ scales from 4.0x at 22 tok/s/user to 8.2x at 110 on 8K/1K - [B200 NVFP4 vs H200 INT4 on Kimi K2.5/K2.6: Up to 2.95x Better Performance per Dollar](https://inferencex.semianalysis.com/blog/b200-nvfp4-vs-h200-int4-kimi-k2-vllm-perf-per-dollar): On vLLM 8K/1K the NVFP4 path on B200 is 2.71x–2.95x cheaper per million tokens than H200 INT4 across the entire 30–90 tok/s/user serving band, and 2.45x–2.74x cheaper than B200 INT4 on the same silicon. Both factors decompose cleanly into B200's HBM bandwidth, HBM capacity, and NVFP4 tensor cores - [MI355X DeepSeek-V4-Pro on SGLang: 110.5x Throughput per GPU in 26 Days](https://inferencex.semianalysis.com/blog/mi355x-deepseek-v4-pro-sglang-110x-in-26-days): The amd/deepseek_v4 side branch shipped TileLang attention indexer, Triton sparse MLA, fused RoPE/Hadamard, FlyDSL MoE, and FP4 weights across 31 performance optimizations PRs — lifting first-light 20 tok/s/GPU at 2.4 tok/s/user into 2,256 tok/s/GPU at 9.4 tok/s/user on 8K/1K, with both throughput and interactivity climbing together - [AMD MI355X GLM-5 Inference: Up to 40% Cheaper per Million Tokens than B200 on SGLang FP8](https://inferencex.semianalysis.com/blog/mi355x-glm5-fp8-sglang-40-cheaper-than-b200): 14 weeks after GLM-5 launched, AMD landed both MTP and non-MTP SGLang FP8 recipes on MI355X — fused MLA + FP8 KV cache via TileLang flips the single-node FP8 cost curve in AMD favor across most of the performance Pareto - [AMD MI355X Qwen3.5 397B-A17B Inference: Up to 19x Throughput per GPU in 3 Months on SGLang FP8](https://inferencex.semianalysis.com/blog/mi355x-qwen3-5-sglang-v0-5-12-up-to-17x): From v0.5.8 (Feb) → v0.5.10rc0 (Apr) → v0.5.12 (May), three AITER kernel landings on MI355X plus a TP=8 → TP=2/TP=4 retune push Qwen3.5 8k/1k peak from 1.3k to 6.4k tok/s/GPU and extend the curve out to 75 tok/s/user - [GB200 NVL72 vs B200 on DeepSeek R1 670B: Up to 4.4x Throughput per GPU at 125 tok/s/user](https://inferencex.semianalysis.com/blog/gb200-nvl72-vs-b200-disagg-deepseek-r1-fp4-dynamo-trt): DeepSeek R1 FP4 1k/1k. NVL72's 72-GPU NVLink scale-up fabric lets decode run wide EP up to EP=32, where B200's 8-GPU NVLink island caps out at EP=8 over RoCEv2 - [SGLang 0.5.6 on B200 DeepSeek R1 FP4: Up to 1.8x at Low Concurrency](https://inferencex.semianalysis.com/blog/sglang-0-5-6-b200-deepseek-r1-fp4-up-to-1-8x): Piecewise CUDA graphs for DeepSeek V3, a unified event loop, and JIT kernels push 8k/1k throughput from 508 to 907 tok/s/GPU on the same 16 GPU B200 pool - [GB200 NVL72 vs B200 on Kimi K2.5: 3.1x from Wide EP vLLM](https://inferencex.semianalysis.com/blog/gb200-nvl72-kimi-k2-5-vllm-wide-ep-3x-vs-b200): Rack scale NVLink on NVL72 lets Dynamo vLLM run Kimi K2.5 wide EP up to Decode EP 16, taking peak throughput from 4,021 to 12,587 tok/s/GPU on 8k/1k NVFP4 - [AMD MI355X Kimi K2.5 Inference: 7.7x Throughput, Up To 15x Interactivity in 25 Days on vLLM](https://inferencex.semianalysis.com/blog/mi355x-kimi-k2-5-vllm-aiter-7x-speedup): vLLM PR #35850 Fixed AITER MLA Dispatch on MI355X CDNA4, Unlocking Kimi K2.5 Inference Performance at TP=8, Shipped in vLLM 0.18 - [InferenceX v2: NVIDIA Blackwell Vs AMD vs Hopper - Formerly InferenceMAX](https://inferencex.semianalysis.com/blog/inferencex-v2-nvidia-blackwell-vs-amd-vs-hopper): GB300 NVL72, MI355X, B200, H100, Disaggregated Serving, Wide Expert Parallelism, Large Mixture of Experts, SGLang, vLLM, TRTLLM - [InferenceMAX: Open Source Inference Benchmarking](https://inferencex.semianalysis.com/blog/inferencemax-open-source-inference-benchmarking): NVIDIA GB200 NVL72, AMD MI355X, Throughput Token per GPU, Latency Tok/s/user, Perf per Dollar, Cost per Million Tokens, Tokens per Provisioned Megawatt, DeepSeek R1 670B, GPTOSS 120B, Llama3 70B ## Whitepapers - [AMD Instinct MI355X Kimi K3 Can Generate Up to $32B of Revenue per GigaWatt per Year](https://inferencex.semianalysis.com/whitepaper/amd-mi355x-32b-revenue-per-gigawatt-kimi-k3): Executive Summary - Agentic Inference Serving Economics Analysis (PDF: https://inferencex.semianalysis.com/whitepaper/amd-mi355x-32b-revenue-per-gigawatt-kimi-k3/pdf/SemiAnalysis-InferenceX-Executive-Summary_AMD-MI355X-Revenue-per-Gigawatt-Kimi-K3.pdf)