On the part of the curve between 40 and 60 seconds of end to end latency, MI355X ATOM beats even GB300 NVL72 vLLM on performance per dollar. On a 2.8 trillion parameter model, against a rack-scale NVIDIA system, that is a result AMD earned.

The upstream caveat
It is great that AMD is quickly pushing K3 performance forward with ATOM. The encouragement that goes with that is to prioritize upstreaming these improvements into vLLM. ATOM is currently AMD's best-performing engine, but vLLM remains the more relevant comparison for customers using an upstream open-source serving stack, and a curve that only exists under a vendor runtime does not describe what those customers can deploy.
Hopper cannot really play
Hopper struggles to serve the AgentX workload for Kimi K3. Part of that is simply that Kimi is a massive model, and part of it is that vLLM maintainers and NVIDIA have not been focusing on optimizing Hopper for K3.
The specifics are concrete rather than vague: Hopper, at SM90, requires custom tuned kernels for K3, along with TP32 and EP32 tuned shapes for serving at high interactivity. Neither exists yet, so the hardware is not the binding constraint so much as the absence of anyone working on it.

This is a useful reminder that a generation gap in a benchmark chart is partly a generation gap in engineering attention. Hopper served DeepSeek R1 well for years because that model received sustained kernel work. K3 has not received the same, and the curve reflects the difference.
These results are one slice of AgentX 1.0. The full analysis, the replay methodology, and the 70+ upstream PRs the benchmark drove are in AgentX - InferenceXv3: Does the CUDA Moat Hold Up in Agentic Inferencing?. Every point is explorable on the free dashboard.
All articles and posts are © SemiAnalysis. All rights reserved. The AGPL-3.0 license covering the application source code does not apply to article content.