GLM 5.3 builds on top of GLM 5.2 744B with additional post training. This is a frontier level model.
In terms of open source SGLang performance, this is another model where NVIDIA again beats AMD on realistic agentic inference performance. At 150 tok/s/user p90 interactivity, NVIDIA has up to 5x better cost efficiency.

The free-hardware test
There is a useful way to size a gap like this. With the current state of AMD software, at 150 tok/s/user, NVIDIA's performance advantage is so great that even if the competitor chip hardware was sold for free, with providers still of course paying for datacenter hosting, power, and other operating costs, cost per token would still be cheaper when using NVIDIA.
That framing matters because hardware price is the lever most often reached for in a procurement conversation, and at this operating point it is not a lever that reaches far enough. The constraint is software, not the bill of materials.

This is a snapshot, not a verdict
The gap is measured against the current state of AMD software, and that state has been moving quickly elsewhere in these results. We look forward to AMD's performance optimizations in the upcoming AgentX update article in a couple of weeks, which will also include some other very exciting results.
It is also worth being precise about scope: this is the open source SGLang comparison. AMD's vendor engine tells a different story on parts of this same model, which is covered separately.
These results are one slice of AgentX 1.0. The full analysis, the replay methodology, and the 70+ upstream PRs the benchmark drove are in AgentX - InferenceXv3: Does the CUDA Moat Hold Up in Agentic Inferencing?. Every point is explorable on the free dashboard.
All articles and posts are © SemiAnalysis. All rights reserved. The AGPL-3.0 license covering the application source code does not apply to article content.