UALink
Also known as Ultra Accelerator Link, UALoE
In plain English
UALink is an open industry standard for the fast scale up links between accelerators in a rack, the ecosystem answer to NVLink.
Technical definition
UALink is an open interconnect specification for accelerator to accelerator scale up communication, letting many chips in a rack share memory traffic at switch fabric speeds.
Engineering details
NVLink proved that a rack of chips joined by a low latency memory semantic fabric can behave like one giant accelerator, but it is proprietary. The UALink consortium defines an open equivalent so vendors beyond NVIDIA can build rack scale domains. UALink over Ethernet, shortened to UALoE, runs the protocol over Ethernet switching. AMD Helios generation racks adopt this path to form 72 chip scale up domains comparable in structure to NVL72.
Why it matters
Scale up domain size increasingly decides serving architecture, since wide expert parallelism and disaggregation want dozens of chips within one fast fabric. An open standard determines whether rack scale inference stays a single vendor advantage or becomes an ecosystem capability.
How to read it in InferenceX
InferenceX lists UALoE72 class systems such as AMD MI455X racks alongside NVL72 systems in its hardware coverage, so rack scale fabrics can be compared head to head on identical model workloads as they reach the market.
Source material
See the concept in real benchmarks
Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis
Rubin LUT Based Tensor Core, Feynman, Rack Scale, Perf Per MegaWatt, Perf Per Dollar, Software Improvements, Public Rubin Software, PyTorch, vLLM, OpenAI Triton
AgentX - InferenceXv3: Does the CUDA Moat Hold Up in Agentic Inferencing?
$3 Million USD dataset open sourced, 1 Mil+ Context Length, Multiturn, Sub Agents 95%+ KVCache HitRate, GB300 NVL72, MI355X, B200