Artifact set

ArtifactDateRole
NVIDIA blog: Vera Rubin NVL72 MLPerf debut16 Sep 2026Vendor narrative + entry IDs + ratio claims
MLCommons v6.1 results announcement16 Sep 2026Round rules, new tests, participation record
Supplemental PDF (submitter statements)Embargo lifted 16 Sep 2026 8:00 AM PTNVIDIA ~300-word statement + partner statements
MLCommons visualizer / Datacenter & Edge results pagesLivingMachine-readable Closed Division rows

Division for the NVIDIA Vera Rubin claims cited below: MLPerf Inference v6.1, Closed Division, per NVIDIA blog footnote.

NVIDIA Vera Rubin NVL72 preview claims (as stated)

From the NVIDIA blog and the NVIDIA block in the supplemental PDF:

WorkloadSoftware pathClaim vs GB300 NVL72Scenarios named
Qwen3-VLvLLM with NVIDIA Dynamoup to 3.7x higher throughputoffline, server, interactive
DeepSeek-R1TensorRT-LLMup to 2.5x higher token throughputoffline, server, interactive

Entry IDs cited by NVIDIA for the platform comparison figures: 6.1-0106 and 6.1-0074.

Hardware / technique claims attached to these submissions in the NVIDIA blog (not separately measured here):

  • Enhanced Tensor Cores and Transformer Engine on prefill and decode.
  • NVFP4 precision across model weights, attention, and KV cache.
  • Disaggregated serving (prefill/decode separation).
  • Large-scale expert parallelism for MoE layers.
  • NVL72 scale-up via sixth-generation NVLink and NVLink Switch (NVIDIA states 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet in this post).

Partner note: Nebius also submitted Vera Rubin NVL72 preview results (NVIDIA blog; Nebius supplemental names Nebius VR200 NVL72 built on NVIDIA Vera Rubin NVL72).

Companion GB300 NVL72 figures in the same round

Cited so the Record does not orphan the baseline rack:

ClaimDetailEntries / notes
DeepSeek-R1 scale-out72 GPUs → 288 GPUs (four GB300 NVL72 racks); 99% offline scaling efficiency vs single-rack baselineNVIDIA cites 6.1-0073 and 6.1-0074
Qwen3-VL software gainup to 1.6x vs MLPerf Inference v6.0 on GB300lower KV-cache precision, fusion, kernels, disaggregated serving with vLLM + Dynamo
WAN 2.2 text-to-video0.65 720p videos/s; 5.7 s/video on full-rack GB300 NVL72; 9x throughput and 7.5x lower latency vs single-nodeNVIDIA blog / supplemental
Edge-AgenticQwen3.6-27B on Jetson AGX Thor with TensorRT Edge-LLMnew v6.1 edge test

Post-submission (NVIDIA blog): further unverified gains on GPT-OSS-120B and DLRMv3. Status: not MLCommons-verified as of the blog.

SemiAnalysis AgentX preview line in the NVIDIA blog (30x vs GB300): outside the Closed Division table. Do not merge into MLPerf citations.

MLCommons round facts (announcement)

  • Results date: 16 September 2026.
  • Participation: record 30 organizations (six first-time submitters named by David Kanter).
  • New tests: End-to-End RAG; Edge Agentic Inference.
  • Speculative decoding supported in interactive for selected benchmarks (including GPT-OSS task).
  • Preview / new accelerators named include NVIDIA Rubin and NVIDIA Vera Rubin NVL72.
  • Largest system submitted: 512 accelerators (AMD/Crusoe path in supplemental narrative).
  • 50% of submitters used the new API-centric harness (foundation for MLPerf Endpoints).

  • Per-accelerator server-scenario improvements cited by MLCommons: VLM best result 2.99x vs v6.0 (six months); DeepSeek R1 best result 5.7x vs v5.1 (one year). These are round-level bests. They are separate from Vera Rubin-specific ratios.

How to cite

  1. For Rubin vs GB300 ratios, cite NVIDIA blog + entry IDs 6.1-0106 / 6.1-0074, and confirm the numeric cells on the MLCommons datacenter results page or visualizer.
  2. For round rules and new tests, cite the MLCommons announcement.
  3. For submitter self-description, quote the Supplemental PDF NVIDIA block (and Nebius if partner preview is in scope).
  4. Mark any power, price, or availability figure UNKNOWN unless it appears in these artifacts (it does not, for Vera Rubin list price / TDP in the cited passages).

Gaps (UNKNOWN)

  • Full per-scenario numeric table reproduced in this Record body (retrieve live from MLCommons pages; cells can be updated after errata).
  • Vera Rubin NVL72 public list price, TDP, and GA date.
  • Independent re-implementation of the Dynamo/vLLM or TensorRT-LLM recipes outside submitters.
  • Whether post-submission GPT-OSS / DLRMv3 gains will appear in a future verified round.