Vera Rubin
Also known as Vera Rubin platform, Rubin GPU, Vera CPU, VR NVL72
In plain English
Vera Rubin is NVIDIA’s platform combining Rubin GPUs, Vera CPUs, and the interconnect and networking components around them.
Technical definition
Vera Rubin is the NVIDIA platform that the article describes as co-designed across Rubin GPU, Vera CPU, NVLink 6 Switch, ConnectX-9, BlueField-4, and Spectrum-6.
Engineering details
The platform name covers more than an accelerator alone. The Rubin NVL72 article evaluates a complete serving configuration with early pre-release TensorRT-LLM software. It specifies a production SKU with 2300 W TDP and 1.5 TB of CPU LPDDR5X per compute tray, rather than treating every announced configuration as identical.
Why it matters
Agentic performance reflects the interaction of compute, memory capacity, communication, and serving software. A measured advantage cannot be assigned solely to one component, and a rack-level result does not establish an identical advantage for every model or latency target.
How to read it in InferenceX
Use the article’s model, engine, precision, workload, and interactivity target when comparing Vera Rubin with GB300 or single-node systems. The published results are a software snapshot; the article’s expectations for later gains are projections rather than measurements.
Source material
See the concept in real benchmarks
Rubin NVL72 Agentic Inference: 67x better Performance per Dollar
Jensen Sandbagging Performance Again, 2x more Annual Profit Per GigaWatt, The More you Buy, The More you Earn, AgentX, InferenceX, Extreme Co-Design
Vera Rubin NVL72 vs GB200 NVL72? Inference TCO & Architecture Analysis
Rubin LUT Based Tensor Core, Feynman, Rack Scale, Perf Per MegaWatt, Perf Per Dollar, Software Improvements, Public Rubin Software, PyTorch, vLLM, OpenAI Triton