Artifact set
| Artifact | Date | Role |
|---|---|---|
| NVIDIA blog: Vera Rubin NVL72 MLPerf debut | 16 Sep 2026 | Vendor narrative + entry IDs + ratio claims |
| MLCommons v6.1 results announcement | 16 Sep 2026 | Round rules, new tests, participation record |
| Supplemental PDF (submitter statements) | Embargo lifted 16 Sep 2026 8:00 AM PT | NVIDIA ~300-word statement + partner statements |
| MLCommons visualizer / Datacenter & Edge results pages | Living | Machine-readable Closed Division rows |
Division for the NVIDIA Vera Rubin claims cited below: MLPerf Inference v6.1, Closed Division, per NVIDIA blog footnote.
NVIDIA Vera Rubin NVL72 preview claims (as stated)
From the NVIDIA blog and the NVIDIA block in the supplemental PDF:
| Workload | Software path | Claim vs GB300 NVL72 | Scenarios named |
|---|---|---|---|
| Qwen3-VL | vLLM with NVIDIA Dynamo | up to 3.7x higher throughput | offline, server, interactive |
| DeepSeek-R1 | TensorRT-LLM | up to 2.5x higher token throughput | offline, server, interactive |
Entry IDs cited by NVIDIA for the platform comparison figures: 6.1-0106 and 6.1-0074.
Hardware / technique claims attached to these submissions in the NVIDIA blog (not separately measured here):
- Enhanced Tensor Cores and Transformer Engine on prefill and decode.
- NVFP4 precision across model weights, attention, and KV cache.
- Disaggregated serving (prefill/decode separation).
- Large-scale expert parallelism for MoE layers.
- NVL72 scale-up via sixth-generation NVLink and NVLink Switch (NVIDIA states 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet in this post).
Partner note: Nebius also submitted Vera Rubin NVL72 preview results (NVIDIA blog; Nebius supplemental names Nebius VR200 NVL72 built on NVIDIA Vera Rubin NVL72).
Companion GB300 NVL72 figures in the same round
Cited so the Record does not orphan the baseline rack:
| Claim | Detail | Entries / notes |
|---|---|---|
| DeepSeek-R1 scale-out | 72 GPUs → 288 GPUs (four GB300 NVL72 racks); 99% offline scaling efficiency vs single-rack baseline | NVIDIA cites 6.1-0073 and 6.1-0074 |
| Qwen3-VL software gain | up to 1.6x vs MLPerf Inference v6.0 on GB300 | lower KV-cache precision, fusion, kernels, disaggregated serving with vLLM + Dynamo |
| WAN 2.2 text-to-video | 0.65 720p videos/s; 5.7 s/video on full-rack GB300 NVL72; 9x throughput and 7.5x lower latency vs single-node | NVIDIA blog / supplemental |
| Edge-Agentic | Qwen3.6-27B on Jetson AGX Thor with TensorRT Edge-LLM | new v6.1 edge test |
Post-submission (NVIDIA blog): further unverified gains on GPT-OSS-120B and DLRMv3. Status: not MLCommons-verified as of the blog.
SemiAnalysis AgentX preview line in the NVIDIA blog (30x vs GB300): outside the Closed Division table. Do not merge into MLPerf citations.
MLCommons round facts (announcement)
- Results date: 16 September 2026.
- Participation: record 30 organizations (six first-time submitters named by David Kanter).
- New tests: End-to-End RAG; Edge Agentic Inference.
- Speculative decoding supported in interactive for selected benchmarks (including GPT-OSS task).
- Preview / new accelerators named include NVIDIA Rubin and NVIDIA Vera Rubin NVL72.
- Largest system submitted: 512 accelerators (AMD/Crusoe path in supplemental narrative).
-
50% of submitters used the new API-centric harness (foundation for MLPerf Endpoints).
- Per-accelerator server-scenario improvements cited by MLCommons: VLM best result 2.99x vs v6.0 (six months); DeepSeek R1 best result 5.7x vs v5.1 (one year). These are round-level bests. They are separate from Vera Rubin-specific ratios.
How to cite
- For Rubin vs GB300 ratios, cite NVIDIA blog + entry IDs 6.1-0106 / 6.1-0074, and confirm the numeric cells on the MLCommons datacenter results page or visualizer.
- For round rules and new tests, cite the MLCommons announcement.
- For submitter self-description, quote the Supplemental PDF NVIDIA block (and Nebius if partner preview is in scope).
- Mark any power, price, or availability figure UNKNOWN unless it appears in these artifacts (it does not, for Vera Rubin list price / TDP in the cited passages).
Gaps (UNKNOWN)
- Full per-scenario numeric table reproduced in this Record body (retrieve live from MLCommons pages; cells can be updated after errata).
- Vera Rubin NVL72 public list price, TDP, and GA date.
- Independent re-implementation of the Dynamo/vLLM or TensorRT-LLM recipes outside submitters.
- Whether post-submission GPT-OSS / DLRMv3 gains will appear in a future verified round.
