OFICIAL NVIDIA Newsroom

Sparks Fly: NVIDIA Accelerates Local AI at IFA 2026

What happened
Based on NVIDIA Newsroom · Sep 03, 2026

NVIDIA, Microsoft and partners announced at IFA 2026 tools to simplify local AI agent deployment on RTX and DGX systems, including new RTX Spark Windows PCs and streamlined setup for popular agent apps.

Video

Video available

Key points
·
At IFA 2026, NVIDIA, Microsoft and its partners are teaming up to provide faster inference and new tools that make agents easier to set up and run locally on NVIDIA hardware.
·
New compact NVIDIA RTX Spark Windows PCs are also coming in October to give AI enthusiasts, developers and creators more ways to run capable agents locally and securely.
·
Getting a local agent up and running with local models required some effort — choosing a model, finding a compatible inference server, dialing in quantization settings and keeping everything updated.
·
Three of the most widely used agent apps will offer simplified local model setup on Windows, each built on llama.cpp and incorporating NVIDIA’s latest inference optimizations.
Key numbers
·
Perplexity’s Portable Computer agent, currently available on Linux systems with NVIDIA DGX Spark, will expand to Windows RTX GPUs with at least 24GB VRAM.
·
9x higher throughput on a GeForce RTX 5090, while vLLM delivers 1.
·
2x on RTX PRO 6000 Blackwell and up to 1.

NVIDIA, Microsoft and partners unveiled at IFA 2026 new tools to reduce the complexity of running AI agents locally on NVIDIA hardware. Compact NVIDIA RTX Spark Windows PCs will launch in October, offering developers and creators a local, secure way to operate capable agents. The initiative aims to eliminate manual steps such as model selection, inference server configuration and quantization tuning by integrating optimized setups directly into widely used agent applications. NVIDIA states these changes are designed to lower barriers to entry for local AI deployment on supported systems.

Three leading agent applications will introduce simplified local model setup on Windows, each leveraging llama.cpp and NVIDIA’s inference optimizations. Perplexity’s Portable Computer agent, currently available on Linux systems with NVIDIA DGX Spark, will expand to Windows RTX GPUs with at least 24GB VRAM. The app packages models, orchestration and tools into a single experience, allowing users to run workflows locally without cloud credits while selectively escalating tasks to cloud models when needed.

NVIDIA reports performance gains in inference acceleration through collaborations with the open-source llama.cpp and vLLM communities. llama.cpp achieves up to 1.9x higher throughput on a GeForce RTX 5090, while vLLM delivers 1.2x on RTX PRO 6000 Blackwell and up to 1.4x on two DGX Spark clusters, attributed to kernel optimizations and backend improvements. These enhancements are accessible via llama.cpp and vLLM backends and integrated into applications such as LM Studio and Ollama.

NVIDIA introduced the Personal AI Router (PAIR), a free, open-source tool that coordinates AI workloads across multiple local PCs to improve parallel processing for agentic tasks. PAIR automatically detects compatible systems on a network and distributes inference requests to available GPUs, reducing bottlenecks. The beta supports Windows, macOS and Linux on NVIDIA GeForce RTX 20 Series and newer, RTX PRO workstations, DGX Spark and Apple M4 silicon, working with Ollama and LM Studio.

Original source → Deals on Clipraptor.com →