ubenchX Microbenchmarks (Beta)

Low-level GPU microbenchmarks measuring fundamental hardware characteristics.

Device-Memory Copy Bandwidth

Microbenchmark measuring device-memory copy bandwidth across message sizes on NVIDIA and AMD GPUs.

B300 SXMGB200 NVL72B200 SXMH200 SXMH100 SXMMI355XMI325XMI300X

Memory Bandwidth Utilization (MBU) vs Message Size

MBU is relative to each GPU's own peak HBM bandwidth from GPU_SPECS.

Shift+Scroll to zoom · Drag to pan · Double-click to reset · Click a point to pin tooltip

Methodology

Times b.copy_(a) on float32 tensors using triton.testing.do_bench; bandwidth = 2 * bytes / time (read + write). Power-of-two sizes from 8 B to 16 GiB.

B300 SXM: NVIDIA B300 SXM6 AC | Driver: 580.159.03 | PyTorch: 2.14.0+cu130 | Triton: 3.8.0 | Container: pytorch/pytorch:2.14.0-cuda13.0-cudnn9-runtime | Peak: 8 TB/s (GPU_SPECS)

GB200 NVL72: NVIDIA GB200 | Driver: 580.126.20 | PyTorch: 2.13.0+cu130 | Triton: 3.7.1 | Container: lmsysorg/sglang:nightly-dev-cu13-20260922-582389ce (arm64) | Peak: 8 TB/s (GPU_SPECS)

B200 SXM: NVIDIA B200 | Driver: 580.159.03 | PyTorch: 2.13.0+cu130 | Triton: 3.7.1 | Container: lmsysorg/sglang nightly 2026-09-22 (582389ce) | Peak: 8 TB/s (GPU_SPECS)

H200 SXM: NVIDIA H200 | Driver: 580.173.02 | PyTorch: 2.13.0+cu130 | Triton: 3.7.1 | Container: lmsysorg/sglang:nightly-dev-cu13-20260923-06008c17 | Peak: 4.8 TB/s (GPU_SPECS)

H100 SXM: NVIDIA H100 80GB HBM3 (SXM) | Driver: 580.159.03 | PyTorch: 2.5.1+cu124 | Triton: 3.1.0 | Container: pytorch/pytorch:2.5.1-cuda12.4-cudnn9-runtime | Peak: 3.35 TB/s (GPU_SPECS)

MI355X: AMD Instinct MI355X | Driver: amdgpu 6.16.6 | PyTorch: 2.11.0+rocm7.2 | Triton: 3.7.0 | Container: sglang_local-rocm724-mi35x.sqsh | Peak: 8 TB/s (GPU_SPECS)

MI325X: AMD Instinct MI325X | Driver: amdgpu 6.16.13 | PyTorch: 2.10.0+rocm7.2.4 | Triton: 3.6.0 | Container: rocm/pytorch:rocm7.2.4_ubuntu24.04_py3.12_pytorch_release_2.10.0 | Peak: 6 TB/s (GPU_SPECS)

MI300X: AMD Instinct MI300X | Driver: amdgpu 6.16.13 | PyTorch: 2.10.0+rocm7.2.4 | Triton: 3.6.0 | Container: rocm/pytorch:rocm7.2.4_ubuntu24.04_py3.12_pytorch_release_2.10.0 | Peak: 5.3 TB/s (GPU_SPECS)

Source: https://github.com/SemiAnalysisAI/InferenceX/pull/3879

Source: https://github.com/SemiAnalysisAI/InferenceX/pull/3874

Source: https://github.com/SemiAnalysisAI/InferenceX/pull/3880