InferenceX Research

Articles

Benchmark write-ups on agentic inference, AgentX results, and chip and serving-stack economics.

New to the terminology? Browse the AI inference glossary.

37 min read

InferenceMAX: Open Source Inference Benchmarking

NVIDIA GB200 NVL72, AMD MI355X, Throughput Token per GPU, Latency Tok/s/user, Perf per Dollar, Cost per Million Tokens, Tokens per Provisioned Megawatt, DeepSeek R1 670B, GPTOSS 120B, Llama3 70B

  • benchmark
  • gpu
  • inference
  • +1