NVIDIA and partners simplify local AI for users

NVIDIA and its ecosystem partners are rolling out new tools and software updates that make it much easier for enterprises to deploy and run AI agents locally on RTX-powered PCs.

Three popular agent applications — Perplexity Portable Computer, Hermes Agent and OpenClaw — are introducing one-click local model setup on Windows, simplifying the complex manual configuration that has hindered local AI adoption.

Previously, deploying local agents required IT teams to manually select models, configure inference servers, tune quantisation settings, and manage ongoing updates It is a time-consuming process that has created friction for enterprise adoption. These new streamlined workflows automatically detect NVIDIA GPUs, select appropriate models and configurations, and deploy optimised inference backends without manual intervention.

Perplexity’s Portable Computer, available on Linux systems including NVIDIA DGX Spark with Windows support coming soon, enables enteprises to run complete AI workflows locally without consuming cloud credits. The system intelligently escalates only specific tasks to frontier cloud models when necessary, with explicit user authorisation, helping enterprises maintain control over sensitive data while leveraging cloud capabilities when beneficial.

Hermes Agent from Nous Research will provide automated local model deployment across RTX and DGX systems on Windows, integrating llama.cpp with NVIDIA inference optimisations to deliver production-ready performance without requiring specialised AI infrastructure expertise.

OpenClaw, the largest AI project on GitHub with more than 380,000 stars, is launching a Windows App that simplifies optimised local model setup on any RTX GPU with at least 24GB of VRAM to make enterprise-grade AI accessible to knowledge workers without dedicated ML teams.

Driving business value

Local AI becomes viable for enterprise workloads when inference performance meets production requirements. NVIDIA’s ongoing optimisations to llama.cpp and vLLM are delivering substantial throughput improvements on RTX hardware that enable local deployment of sophisticated AI agents which previously required cloud infrastructure.

llama.cpp achieves up to 1.9x higher throughput on GeForce RTX 5090 hardware through kernel optimisations, enhanced speculative decoding techniques and faster prefill processing. These improvements make local agents more responsive for real-time business applications.

vLLM delivers 1.2x performance gains on RTX PRO 6000 Blackwell Workstation Edition and up to 1.4x on dual DGX Spark clusters, with new XQA attention kernels in FlashInfer and backend optimisations accelerating inference across both platforms.

Maximising existing investment

With many households and enterprises operating two or more PCs, significant computing capacity often sits idle during business hours. NVIDIA’s open source Personal AI Router (PAIR) enables enterprises to leverage this underutilised infrastructure for local AI workloads without additional capital expenditure.

PAIR automatically discovers compatible PCs across the local network and routes independent inference requests to systems with available capacity, working seamlessly with Ollama and LM Studio. Enterprises can distribute AI workloads across multiple machines, keeping primary workstations available for user tasks while secondary systems handle agent processing in the background.

For instance, an AI agent tasked with email triage and prioritisation can split work across multiple subagents, with PAIR distributing those jobs across available PCs rather than queuing everything on a single GPU. This approach maximises ROI on existing hardware while enabling more sophisticated AI workflows without cloud dependency.

PAIR beta is available for Windows, macOS and Linux, supporting GeForce RTX 20 Series and newer, RTX PRO workstation GPUs (Turing and newer), DGX Spark, and Apple M4 or newer silicon, providing flexibility for heterogeneous enterprise environments.

Staying on-premise

For creative professionals and marketing teams, local AI deployment addresses two critical business concerns: data privacy and cost predictability.

CyberLink’s PhotoDirector AI PC Mode integrates diffusion models directly into the editing workflow, optimised for the upcoming NVIDIA RTX Spark PCs to enable teams to experiment with generative AI without token-based pricing models or uploading proprietary creative assets to the cloud.

When RTX Spark launches in October, PhotoDirector 365 users will gain access to AI-powered tools for generative editing, enhancement, object and distraction removal, background replacement, portrait refinement, and creating entirely new visuals. Users have the option to run locally or in the cloud based on workload requirements. On NVIDIA GPUs, PhotoDirector uses TensorRT-RTX and FP8 to accelerate local AI to deliver production-quality results without cloud latency or data transfer concerns.

Arriving in October

NVIDIA RTX Spark Windows PCs are scheduled to ship in October 2026, with new designs from Acer and Lenovo joining existing OEM partners. The platform features a one-petaflop RTX Blackwell GPU, up to 128GB of unified memory and a 20-core Grace CPU to support always-on AI agents and demanding creative workloads in a single system.

Major enterprise software publishers including Electronic Arts, Embark and Ubisoft have announced support for RTX Spark, alongside earlier commitments from Krafton, NetEase, Riot Games, and Xbox.