What it does

Vera Rubin NVL72 enters MLPerf Inference as a preview rack-scale system. The measured artifact is throughput under Closed Division rules on two heavy benchmarks: DeepSeek-R1 and Qwen3-VL.

Primaries:

Headline deltas NVIDIA attributes to Vera Rubin NVL72 vs GB300 NVL72 (same blog; supplemental NVIDIA statement):

WorkloadStack namedPeak claim vs GB300 NVL72
Qwen3-VLvLLM + NVIDIA Dynamoup to 3.7x throughput
DeepSeek-R1TensorRT-LLMup to 2.5x token throughput

Scenarios named: offline, server, and interactive. Entry IDs cited by NVIDIA: 6.1-0106 and 6.1-0074.

Codesign levers named in the NVIDIA blog: enhanced Tensor Cores and Transformer Engine for prefill and decode; NVFP4 across weights, attention, and KV cache; disaggregated serving (prefill/decode split); large-scale expert parallelism for MoE layers; NVL72 NVLink / NVLink Switch scale-up (NVIDIA claims 10x higher packet rates and 3x lower latency vs off-the-shelf Ethernet in this post).

Companion GB300 results in the same round (baseline context beside the Rubin headline):

  • DeepSeek-R1 scaled from 72 GPUs (one NVL72) to 288 GPUs (four racks) at 99% offline scaling efficiency (entries 6.1-0073 and 6.1-0074 per NVIDIA).
  • Qwen3-VL on GB300 improved up to 1.6x vs v6.0 via lower KV-cache precision, fusion, kernels, and disaggregated serving with vLLM + Dynamo.
  • WAN 2.2 text-to-video on a full GB300 NVL72 rack: 0.65 720p videos/s at 5.7 s/video (9x throughput and 7.5x lower latency vs single-node, per NVIDIA).

Nebius also submitted Vera Rubin NVL72 preview results (NVIDIA blog; Nebius supplemental names VR200 NVL72 built on Vera Rubin NVL72).

MLCommons round context: record 30 submitting organizations; new End-to-End RAG and Edge Agentic tests; speculative decoding support in interactive for selected tasks; NVIDIA Rubin / Vera Rubin NVL72 listed among preview platforms.

Why now

  • First peer-reviewed preview numbers for Vera Rubin NVL72 under MLPerf Inference rules.
  • Separates rack-scale Rubin economics talk from earlier product/production messaging (live Vera Rubin production dispatch is a different artifact).
  • Gives procurement teams comparable Closed Division figures against GB300 NVL72 on the same two model families.
  • Shows the software stack NVIDIA wants associated with those numbers: Dynamo, vLLM, TensorRT-LLM, disaggregated serving, NVFP4.

Maturity

Claim layerStatus
MLPerf Closed Division preview submissionPublished 16 Sep 2026
Vera Rubin NVL72 general availability / list priceUNKNOWN in these primaries
Post-submission NVIDIA optimizations (GPT-OSS-120B, DLRMv3)NVIDIA says further gains; not yet MLCommons-verified
SemiAnalysis AgentX "30x vs GB300" mentionNVIDIA blog preview testing; outside MLPerf Closed Division table
Partner Vera Rubin submissions (e.g. Nebius)Present; detailed public scorecards: see MLCommons results pages

Treat preview as directional under MLPerf harness constraints. It is not a guaranteed production SLA.

Pick vs

Use these MLPerf numbers when:

  • you compare rack-scale inference throughput claims for DeepSeek-R1 and Qwen3-VL under Closed Division rules;
  • you need citation-grade entry IDs (6.1-0106 / 6.1-0074) and the supplemental NVIDIA statement;
  • you evaluate whether Dynamo + vLLM or TensorRT-LLM is the reference path NVIDIA used for a given model.

Do not treat them as:

  • a full TCO or power/token invoice (power and price are UNKNOWN here);
  • proof that every model family will see 2.5-3.7x (only the two named workloads carry those peaks);
  • a substitute for your own serving traces (agentic multi-step latency may need AgentX / future MLPerf Endpoints, which NVIDIA flags as coming).

Prefer GB300 scaling data when:

  • you are buying or operating Blackwell NVL72 today and care about 72→288 GPU efficiency (99% offline on DeepSeek-R1 per NVIDIA).

Failure modes

  1. Preview ≠ fleet. Shipping decisions on preview silicon before GA drivers, firmware, and supply locks risk schedule slips.
  2. Peak ratio shopping. "Up to 3.7x" is a peak across scenarios. Scenario-by-scenario and per-accelerator tables live on MLCommons pages; skim those before RFPs.
  3. Stack lock-in reading. Dynamo/vLLM vs TensorRT-LLM paths differ by model in the submission. Copying the wrong stack for your model leaves the claimed delta on the table.
  4. Disagg complexity. Prefill/decode separation and expert parallelism need fabric and orchestration maturity. Mis-tuned disagg can lose the paper gain.
  5. Confusing AgentX with MLPerf. The 30x AgentX line is separate preview testing in the NVIDIA blog. Keep it out of Closed Division citations.
  6. Ignoring software velocity. GB300's 1.6x vs v6.0 shows software still moves the needle on the prior rack. Rubin deltas will also move after submission day.

Operator read

Pull the Closed Division rows for 6.1-0106 and 6.1-0074 from the MLCommons datacenter results pages. Match your candidate models to the submitted stacks. Keep Vera Rubin NVL72 in the "preview rack" column until GA supply and power numbers are public. Use the companion Record piece (mlperf-v6-1-vera-rubin-record) for the evidence table posture.