NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI
NVIDIA introduces Nemotron 3.5 Lightning, a 30-billion-parameter model optimized for agentic AI tasks, and NeMo Switchyard, an open-source routing library to direct tasks to the most suitable model.
Video
Video available
NVIDIA has expanded its Nemotron 3 model family with Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model designed for long-running agentic AI workloads. The model is optimized for specialized tasks within multi-agent systems, offering up to 4x faster output speed and 30% faster task completion compared to similar models. It supports local deployment on NVIDIA RTX PCs, DGX systems, and Jetson devices, as well as cloud and data center environments.
NeMo Switchyard, an open-source library, enables intelligent routing of AI agent requests to the most capable model based on task requirements. It allows enterprises to use a mix of open, proprietary, and NVIDIA models without rewriting applications. The library supports customization of routing algorithms to prioritize quality, latency, or cost, improving token efficiency in model ensembles.
Nemotron 3.5 Lightning is fully customizable and can be post-trained on domain-specific data using NVIDIA NeMo to enhance accuracy for specialized tasks. Early adopters, including CrowdStrike, Harvey with Trajectory, and CodeRabbit with Baseten, are using the model to improve performance in cybersecurity, legal services, and code review, respectively.
The Nemotron 3.5 Lightning model and NeMo Switchyard are available on Hugging Face, ModelScope, OpenRouter, and NVIDIA’s build.nvidia.com as an NIM microservice, with broader ecosystem support through cloud partners and inference platforms. NeMo Switchyard is accessible on GitHub, with integration planned for partner platforms.