NVIDIA says its Vera Rubin rack-scale AI platform is moving into full production, with systems running at CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure. The company describes a supply chain spanning more than 350 AI-factory sites in 30 countries and a partner ecosystem exceeding 300 organisations.
The production ramp matters because Vera Rubin is not presented as a single accelerator. It combines seven chips and five rack trays into a co-designed system covering compute, scale-up and scale-out networking, data processing and software. NVIDIA is positioning the platform as the next foundation for large training clusters and continuously operating inference factories.
Efficiency is the central claim
NVIDIA says an early CoreWeave benchmark running DeepSeek-R1 produced ten times more throughput per megawatt than Grace Blackwell NVL72. It also claims one-tenth the cost per million tokens in a Microsoft and Mistral deployment comparison with GB200 NVL72.
Those figures reflect an industry shift from peak chip performance towards useful output within power and cooling limits. A data centre may have access to more accelerators than its electrical infrastructure can support, so tokens per megawatt and completed work per unit of energy increasingly determine effective capacity.
The Vera CPU uses NVIDIA’s custom Olympus core, while NVLink 6 handles scale-up communication and Spectrum-6 with ConnectX-9 supports scale-out Ethernet. NVIDIA says the platform’s co-design delivers more than twice the throughput on complex NVLink workloads, lower latency and higher packet rates than general-purpose alternatives.
Deployment and infrastructure changes
The rack design removes cables, fans and hoses from the compute tray, which NVIDIA says cuts tray assembly time from hours to about one minute. A 45-degree Celsius liquid-cooling inlet is designed to support chiller-free dry cooling, with the company claiming large water savings for new factories using closed-loop systems.
Cloud and infrastructure partners will determine how quickly those technical changes become accessible to ordinary customers. CoreWeave, Google Cloud, Microsoft Azure and Oracle are named as current deployment partners, but the announcement does not provide a uniform availability date, instance catalogue or price across their services.
Mistral is also adding thousands of Vera Rubin GPUs to expand European compute capacity under a new agreement with Microsoft. NVIDIA says the hardware will support the next generation of Mistral Compute and Microsoft’s European AI infrastructure, spanning public cloud, connected private environments and fully disconnected deployments.
How to assess the claims
The reported performance figures come from NVIDIA and named partners, not a broad independent comparison across production workloads. Buyers should examine the exact model, precision, batch size, latency target, software version and utilisation behind any tokens-per-megawatt result.
Total economics also include networking, storage, power delivery, cooling, reserved capacity, software licensing and operational support. A tenfold gain on one benchmark does not automatically translate into the same improvement for every training, fine-tuning or inference workload.
Capacity planning should also separate training and inference needs. Large training runs favour sustained communication across many accelerators, while production inference may prioritise latency, predictable availability and the ability to serve many models efficiently. Vera Rubin’s integrated design is intended to address both, but the best configuration and commercial model may differ substantially between them.
Adopting a new rack-scale generation can create migration work beyond replacing hardware. Teams may need to qualify model kernels, containers, drivers, networking libraries and monitoring systems, then confirm that saved energy or faster throughput outweighs transition costs. Cloud services can hide some of this complexity, although customers still need to test application behaviour and quota availability.
Vera Rubin’s full-production status is nonetheless a significant milestone. It moves the platform from roadmap discussion into partner deployment and gives large AI operators a clearer basis for capacity planning. The scale of the announced ecosystem also suggests NVIDIA is using coordinated hardware, networking and software delivery to reduce the risk of adopting a new architecture.
For customers, the next useful information will come from cloud availability pages, independently reproducible benchmarks and service pricing. Until those details are published, NVIDIA’s announcement establishes the direction and initial scale of the ramp rather than a single globally consistent buying option.