Briefed.Briefed.
MonoDark

Tech

Vera Rubin Delivers Up To 3.7x GB300 Inference Throughput

NVIDIA’s Vera Rubin NVL72 posted up to 3.7x GB300 throughput in MLPerf Inference v6.1 preview results.

What happened

NVIDIA’s Vera Rubin NVL72 made its first MLPerf Inference v6.1 preview submission, delivering up to 3.7x the throughput of GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios. On DeepSeek-R1, throughput was up to 2.5x higher, using NVIDIA’s inference software. Separately, a 288-GPU GB300 NVL72 submission spanning four racks achieved 99% scaling efficiency, with throughput rising nearly in line with added hardware. NVIDIA also reported up to 1.6x higher performance from software optimisations versus its v6.0 results. The Vera Rubin results were submitted on Qwen3-VL and DeepSeek-R1, two demanding benchmarks in the suite.

Why it matters

NVIDIA says the higher throughput lets each Vera Rubin rack serve more users, generate more tokens and lower cost per token than a GB300 rack. That directly affects the economics of deploying AI inference infrastructure.

Source: NVIDIA Blog

More briefs on Briefed