A preview enters a measured contest

NVIDIA has put a preview of its Vera Rubin NVL72 system into MLPerf Inference v6.1, using published results to make a case for faster and more efficient AI serving. In the NVIDIA announcement, dated 16 September, it reports up to 3.7 times the throughput of GB300 NVL72 on Qwen3-VL and up to 2.5 times on DeepSeek-R1. Those headline ratios are meaningful only alongside the workloads, software and scenarios in which they were measured.

The Rubin result is a preview submission, not a claim that every buyer can obtain identical production performance today. NVIDIA also reports a separate 288-GPU GB300 NVL72 submission spanning four racks, with 99 per cent scaling efficiency in an offline scenario. That result describes how closely throughput tracked added hardware in this test, not the efficiency of every deployment or request pattern.

What the comparison actually measures

MLPerf Inference uses prescribed tasks and measurement rules, making it more informative than an unqualified vendor speed claim. NVIDIA says the Rubin preview covered Qwen3-VL in offline, server and interactive scenarios with vLLM and its Dynamo inference framework. DeepSeek-R1 used TensorRT-LLM. A purchasing team should therefore compare the benchmark configuration with its own model, request mix and latency objective before transferring the ratios to a capacity plan.

Qwen3-VL combines language and vision, while DeepSeek-R1 is a reasoning workload. Both can stress different parts of a serving system. Performance on them says more than a single small-model test would, but it still does not describe every model that an organisation may operate. Context length, batch size, concurrency and service-level targets can change the economics substantially.

NVIDIA cites MLCommons submission identifiers in its report, which gives technically minded readers a route to the underlying entries. Its own interpretation remains a vendor account of those results. The sensible reading is that Rubin performed strongly in the specified preview comparisons, while independent assessment of a planned deployment still needs local workload testing.

Hardware and software move together

The company attributes the Rubin numbers to full-stack design, including improved Tensor Cores, the Transformer Engine, NVFP4 precision and sixth-generation NVLink. It also describes separating the prefill and decode stages of inference and using expert parallelism for mixture-of-experts models. Those techniques affect memory use and communication as well as raw arithmetic speed, which helps explain why one chip specification cannot stand in for a whole-system result.

The release reports software-only progress on existing GB300 NVL72 infrastructure too. In its v6.1 submissions, Qwen3-VL performance improved by up to 1.6 times compared with v6.0, which NVIDIA attributes to changes such as cache precision, kernel fusion and disaggregated serving. This is relevant to customers who will not immediately replace hardware: ongoing software work can alter the useful capacity of equipment already installed.

NVIDIA also mentions improvements measured after the submission deadline on other workloads. It explicitly says those later figures have not been verified by MLCommons. They should not be folded into the official v6.1 result or presented as an equivalent certified comparison.

The rack-scale question

The four-rack Blackwell result asks a practical question: does adding GPUs actually yield proportionate throughput? NVIDIA says its DeepSeek-R1 offline submission expanded from one 72-GPU rack to four racks and retained 99 per cent scaling efficiency. Near-linear growth is valuable where operators need to enlarge a service without losing much capacity to coordination overhead.

However, offline throughput is not the whole experience of an interactive application. An operator may need tight response-time guarantees during demand spikes, isolation between customers or a mix of workloads that leaves some processors waiting. Benchmark scaling under one scenario cannot settle those operational questions. It does indicate a system worth evaluating under the buyer’s own conditions.

The post also reports a preview result from a Nebius system and broad participation by partners. That ecosystem matters for availability and integration, but a partner submission should still be checked against the exact configuration offered for purchase.

How to evaluate the result

A useful procurement test starts with a representative model and traffic trace. Teams should measure throughput at the latency their users can accept, then include energy, networking, storage and software operation in cost per useful response. The fastest published number is not necessarily the least expensive sustainable service.

For Australian operators, deployment location and access through a particular cloud or supplier will also determine whether the hardware claim is actionable. The announcement does not provide a local availability date or an Australian price. Those facts need confirmation from a provider, rather than extrapolation from a US benchmark post.

NVIDIA’s v6.1 release is substantial because it presents a new-system preview, multi-rack scaling and a separate software improvement in one benchmark cycle. Its strongest reading is precise rather than sweeping: a promising preview on named tests, plus evidence that software optimisation continues to matter, with real-world capacity to be proved workload by workload.