NVIDIA and Local AI Community Fuel Open Source Models and Intelligent Agents
NVIDIA and open source communities released new local AI models and tools in August, enabling developers to run advanced agents on consumer and edge hardware. The updates include robotics, video generation, coding and multimodal models optimized for NVIDIA GPUs.
Video
Video available
NVIDIA highlighted August as a month for local AI progress, showcasing open models, software and tools designed to run on consumer and edge devices. The company emphasized accelerated computing resources, libraries and educational materials to help developers build and customize AI agents locally. Updates were shared through a special-edition blog series, with new entries planned over several weeks.
New models released include Cosmos 3 Edge, a 4-billion-parameter robotics and vision model; MiniMax-H3, a 33-billion-parameter video generation model; Poolside AI’s Laguna S 2.1, an 118-billion-parameter coding agent; and DeepSeek-V4-Flash, a 284-billion-parameter mixture-of-experts model with a 1 million-token context window. Each model is optimized for NVIDIA GPUs, including DGX Spark, Jetson and DGX Station systems.
Thinking Machines Lab introduced Inkling-Small, a 276-billion-parameter multimodal model with native reasoning across text, images and audio. Unsloth launched Unsloth Desktop, an open source desktop application for local model training and inference, integrating diffusion, fine-tuning and agent workflows. Alibaba released Wan-Animate-2, a 14-billion-parameter motion transfer model, and LTX-2.5, a state-of-the-art video generation model with multishot support and improved prompt adherence.
Meta unveiled Muse Glimmer, a 30-billion-parameter dense model optimized for coding and local agentic AI. The model supports over 200 tokens per second on NVIDIA RTX 5090 and is designed for always-on local agents, enabling private, multistep workflows without cloud dependency. Developers can fine-tune it using NVIDIA NeMo Automodel or run it with frameworks like vLLM and llama.cpp.