NVIDIA expands NVLink Fusion with NVHBM for custom AI infrastructure
NVIDIA has added NVHBM to its NVLink Fusion platform, putting the memory controller inside the HBM stack. The company says the design can improve bandwidth and power efficiency while giving cloud and AI-chip builders a more standard route to custom, rack-scale systems for enterprise AI teams.
NVIDIA puts Groq 3 LPX into production for agent inference
NVIDIA says its Groq 3 LPX inference system is now in full production alongside Vera Rubin NVL72. The company positions the LPU-based platform for faster token generation and long-context responsiveness in agentic applications, citing early deployments at Nebius and CoreWeave and a tighter integration with its wider AI-factory stack.
Firebird has opened an NVIDIA-powered AI factory in Armenia that the companies describe as the CIS region’s largest. The project combines Dell systems, NVIDIA DSX infrastructure and planned Rubin and Blackwell GPU deployments, signalling a push to build AI capacity closer to regional research, enterprise and public-sector users.
Baseten has added NVIDIA Nemotron 3.5 ASR Streaming models to its model library, offering English and 40-locale multilingual transcription through NVIDIA NIM. Baseten reports support for 100 concurrent real-time WebSocket streams on one H100, with finalisation latency below 140 milliseconds in its test. Production results will vary by workload.
NVIDIA Helps Launch Open Alliance for Secure AI Agents
NVIDIA and roughly 40 technology organisations have formed the Open Secure AI Alliance to develop shared, open defences for AI agents and software. NVIDIA is contributing models, data and its new NOOA agent-harness research project, although delivery and governance details remain limited.
NVIDIA Deploys Vera CPUs to Speed Future Chip Design
NVIDIA is deploying its Vera CPU in the electronic design automation workflows used to build future processors. Early tests on selected Cadence Jasper and Synopsys VCS workloads reached up to 1.5 times higher performance, but full benchmark and external availability details were not disclosed.
NVIDIA Opens GPU Medical Simulation Framework for Robots
NVIDIA has open-sourced a GPU-accelerated Medical Physics Simulation framework within Isaac for Healthcare. It combines anatomy and device physics, simulated sensors and robot learning so developers can generate difficult clinical scenarios and evaluate medical robotics policies before costly hardware and laboratory testing.
NVIDIA Spectrum-6 Targets Gigascale AI Factory Networking
NVIDIA has introduced Spectrum-6, a 102.4-terabit-per-second Ethernet switch system designed for Vera Rubin AI factories. Early deployments are planned by major cloud and infrastructure operators, but NVIDIA’s performance, efficiency and reliability figures remain vendor claims that buyers should validate against their workloads.
NVIDIA Vera Rubin Enters Production Across Global AI Partners
NVIDIA says its Vera Rubin rack-scale AI platform is entering full production across a global partner network. The company reports large gains in throughput per megawatt and token cost, alongside new deployments by cloud and model providers, though independent workload testing is still needed.
NVIDIA Releases Cosmos 3 Edge for On-Device Robotics
NVIDIA has released the 4-billion-parameter Cosmos 3 Edge world model through Hugging Face. The open model combines visual reasoning, prediction and action generation for robots and vision agents, with real-time operation demonstrated on Jetson Thor and scripts for post-training specialised policies.
NVIDIA Agent Toolkit now includes open Omniverse libraries for sensor simulation, GPU-accelerated physics and validation of simulation-ready 3D assets. The company is also releasing a Blender blueprint, while SideFX, PTC and several startups begin integrating the components into existing design and simulation workflows.
Bristol Myers Squibb is deploying a second NVIDIA DGX SuperPOD built from eight Vera Rubin NVL72 systems. The pharmaceutical group plans to make the unified environment available across its research organisation for model training, predictions and agentic drug-discovery workflows using NVIDIA BioNeMo tooling.
Hugging Face and NVIDIA Scale Diffusion Model Fine-Tuning
Hugging Face and NVIDIA have integrated Diffusers-format image and video models with NeMo Automodel for distributed fine-tuning. The Apache 2.0 tooling supports direct Hub checkpoints, configurable parallelism and ready-made recipes for models including Wan, FLUX, HunyuanVideo and Qwen-Image.
Baseten has made NVIDIA's Nemotron 3 Embed 8B and 1B models available for dedicated inference. The pair targets enterprise and code retrieval with different balances of accuracy, throughput and indexing cost, giving developers a managed deployment option while published performance figures remain vendor-reported.