A new system reaches the cloud
CoreWeave has announced availability of NVIDIA Vera Rubin NVL72 systems on its cloud, with Cognition among the first customers running production work. NVIDIA’s 30 September post frames the move around agentic AI: large numbers of model calls, long contexts, isolated execution environments and repeated evaluation. The announcement is materially different from a future roadmap or a benchmark alone. Early customers can now use the platform, although capacity and access should be checked with CoreWeave rather than assumed to be generally available everywhere.
Cognition uses the infrastructure for work related to Devin, its AI software engineer. NVIDIA says the company scaled to thousands of GPUs on CoreWeave over nine months. That context helps explain why throughput and sandbox-start time matter: a coding agent may spend much of its life waiting on tools, tests and intermediate reasoning. Improving one hardware metric can help, but the customer’s full task completion time is the number that ultimately matters.
What the early performance claim means
Cognition compared Vera Rubin NVL72 with a GB200 NVL72 baseline using a sample of software-engineering tasks from FrontierCode. NVIDIA reports up to 4.8 times total token throughput for SWE-2 inference workloads in those early tests. The words up to and early matter. They describe a selected workload and configuration, not a guaranteed gain across every model, sequence length, batch size or production environment. Buyers should ask for the comparison conditions and reproduce tests using their own traffic.
Higher aggregate throughput could make a shared AI service more responsive when many agents are active. It does not by itself establish lower cost per accepted change or better code quality. If an agent spends time on repository access, tool execution or human approval, those stages may dominate the end-to-end journey. A practical evaluation should record queue time, tokens served, power and rental cost, tool latency and final acceptance rate. The launch creates a new option for that evaluation rather than settling it.
Vera CPU and isolated execution
CoreWeave also plans to offer NVIDIA Vera CPU for agent workloads. NVIDIA says a rack configuration brings 128 CPUs and 11,264 cores, supporting a large number of one-core environments in principle. In its tests, CoreWeave reported sandbox starts more than three times faster and a 1.7-times performance gain on passing Terminal-Bench tasks compared with its baseline. The results point to the importance of CPU-side execution for agents, which repeatedly start tools, run code and create temporary environments around GPU inference.
The company says CoreWeave Sandboxes are hardware-isolated and can operate beside training jobs, using NVIDIA networking and data-processing components. Isolation is a useful feature, but security claims must be tested against realistic permissions, data paths and failure modes. An agent that can run code needs limits on network access, credentials and persistence, regardless of the speed of the sandbox. Operators should verify those controls before connecting the environment to sensitive repositories or customer data.
Forge connects use and improvement
CoreWeave Forge is another component of the announcement. It brings together Weights & Biases, OpenPipe expertise and the marimo notebook project in a connected environment for training, evaluating and improving models and agents. NVIDIA’s post says CoreWeave Sandboxes and ARIA are generally available, while some reinforcement-learning rollouts remain in private preview. These differences in availability are important when assessing a platform pitch that spans several products.
The proposed loop is straightforward: production traces reveal failures, evaluation identifies patterns and post-training attempts to improve a model. In practice, each transition needs data governance and a quality check. Production examples may contain sensitive information; a better benchmark score may not translate into fewer user-visible errors; and a new checkpoint can regress elsewhere. A useful deployment keeps versions, evaluation sets and rollback criteria explicit as it moves from experiments to live agent services.
Questions for customers
The NVIDIA account provides a clear 30 September publication date and concrete availability statements, but its benchmarks and efficiency figures are vendor-reported. Potential buyers should compare current CoreWeave capacity with alternatives using an agreed workload and a transparent total-cost model. They should also separate present services from previews and ask how data moves between inference, sandboxes and post-training tools. The systems may offer material gains, yet integration complexity can absorb them if the workflow is not designed carefully.
The broader news is that agent infrastructure is being sold as a connected production loop, not only as a faster accelerator. CoreWeave and NVIDIA are trying to make inference, isolated tool execution and model improvement work together. Whether that combination delivers durable value will be determined by completed-task economics and operational control, rather than a single peak token-throughput result. Customers should request sustained results across busy periods as well as isolated tests, because contention and queueing can erase an apparent hardware advantage. They should verify that claimed gains persist when safety checks and normal observability are enabled.