NVIDIA and Microsoft have outlined a new generation of Windows systems designed to run AI agents locally, led by RTX Spark laptops and compact desktops. Announced on 7 October, the hardware puts NVIDIA’s CUDA stack, Blackwell graphics and Grace processing into personal computers intended for continuous agent workloads as well as development, creation and gaming.

Laptop preorders opened with availability scheduled for 16 October, while compact desktops are due in November. Systems are planned from Acer, ASUS, Dell, HP, Lenovo, Microsoft, MSI and Gigabyte. Microsoft’s Surface Laptop Ultra is among the first devices built around the platform.

Large local models on a personal machine

RTX Spark combines an NVIDIA Blackwell RTX GPU with up to 6,144 cores and a Grace CPU with as many as 20 cores, connected at 600 GB per second. Configurations offer up to 128GB of unified memory and one petaflop of FP4 AI performance, allowing models that would not fit on a conventional laptop to run locally.

NVIDIA says the system can run Qwen 3.8 Flash Next, a large model with 125 billion parameters, without metered cloud inference or sending data off the device. Local operation can improve privacy, responsiveness and cost predictability, although power, thermal limits and model optimisation still shape practical performance.

Always-on agents change the PC workload

The compact desktop is designed for 24-hour operation, giving local agents a persistent machine that does not depend on a laptop remaining open. That supports background tasks, scheduled automation and development workflows that continue over long periods. It also creates security and administration demands more like a server than a traditional personal computer.

Microsoft is addressing that challenge with Execution Containers, Windows security and Agent 365 controls. The aim is to contain agent access, distinguish agent activity from human activity and give administrators visibility. Hardware capability alone is insufficient if a persistent agent inherits unrestricted access to files, networks and applications.

CUDA continuity from laptop to workstation

RTX Spark runs the full CUDA platform, allowing developers to move models and workflows across NVIDIA systems with less rewriting. The same stack extends from portable devices to the compact desktop and the larger DGX Station. That continuity may help teams prototype locally before moving demanding jobs to stronger infrastructure.

Creators receive fifth-generation Tensor Cores, NVFP4 support, hardware-accelerated AV1 and 4:2:2 video processing, while gamers retain familiar RTX features. Combining those workloads in one device may appeal to technical professionals, but buyers should assess whether shared memory and cooling remain adequate under sustained agent and creative use.

DGX Station arrives on Windows

NVIDIA also previewed DGX Station for Windows, its first deskside AI supercomputer built for the Windows enterprise desktop. The system uses a GB300 Grace Blackwell Ultra Desktop Superchip, 748GB of coherent memory and up to 20 petaFLOPS of FP4 performance, targeting models at trillion-parameter scale.

Until now, DGX Station used Linux, while many enterprise developers worked primarily in Windows. The new version is intended to remove that split and still preserve Linux tools through Windows Subsystem for Linux. Researchers can fine-tune and run large models alongside familiar Windows applications and organisational infrastructure.

Local does not automatically mean safe

Keeping inference on the device can reduce data egress, but it does not solve access control, malware or supply-chain risk. An agent able to read sensitive files and use applications needs a separate identity, narrow permissions and comprehensive logs. Organisations should also manage model files, extensions and updates with the same discipline applied to other software.

Procurement decisions should consider memory configuration, sustained performance, serviceability and energy use rather than peak AI figures alone. Teams should test the exact models and agent harnesses they plan to operate, including recovery from crashes and the effect of multiple concurrent workflows.

A new category between laptop and cluster

RTX Spark and DGX Station for Windows aim to bring workloads formerly associated with cloud GPUs or Linux workstations onto the enterprise desktop. That can shorten development cycles and give agents lower-latency access to local applications while preserving a path to larger NVIDIA infrastructure.

The category will succeed if local capability is paired with manageable containment and clear operational economics. The October launch provides the hardware foundation; enterprises will now need evidence that always-on agents can be governed as reliably as the conventional applications they are expected to assist.

Vendors will also need to publish comparable measurements for battery life, sustained inference, memory bandwidth and performance at realistic model precision. Peak petaFLOPS alone does not describe an agent that retrieves files, waits on tools and runs for hours. Independent tests should show whether local execution remains responsive alongside normal Windows work and how performance changes when several agents share the system.

Those results will help enterprise buyers decide which workloads belong locally and which still justify dedicated external cloud infrastructure.