Interestana
Home/News/NVIDIA Ships Nemotron 3.5 Lightning for AI Agents
MarkTechPost3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

NVIDIA Ships Nemotron 3.5 Lightning for AI Agents

NVIDIA has introduced new open technologies designed to facilitate the creation of always-on AI agents composed of specialized models. These technologies include Nemotron 3.5 Lightning, a customizable open model optimized for high-volume agentic tasks, and NeMo Switchyard, an open-source routing library that intelligently directs workflow steps to the most suitable and efficient model. The core problem these innovations address is the inefficiency of sending every task, particularly routine ones like tool calls, result validation, and subagent delegation, to large, frontier reasoning models, which incurs significant cost and latency. Nemotron 3.5 Lightning is a 30 billion parameter Mixture-of-Experts (MoE) model, featuring 3 billion active parameters. Its architecture is a hybrid of Mamba-2, MoE, and Attention mechanisms, and it supports an extensive 1 million token context window. NVIDIA reports that this model achieves up to four times faster output speeds compared to similarly sized models. Furthermore, it demonstrated 30% faster completion of 10,000 PinchBench tasks than Qwen3.6 35B, while maintaining comparable accuracy. The model is designed for broad adoption and is generally available under the permissive OpenMDW-1.1 license, which includes open weights, training data, and recipes, making it ready for commercial use. NVIDIA highlights that Nemotron 3.5 Lightning is deployable on a single modern GPU, such as a 1x DGX Spark (GB10) or a 1x H100, enabling solo developers and startups to utilize it alongside large enterprises. Mid-market teams can leverage cloud platforms like Baseten, Together AI, or Nebius for serving the model, while regulated enterprises have the option to maintain it entirely on-premises. Several industry leaders, including CrowdStrike, Harvey, CodeRabbit, Fastino Labs, and Lila Sciences, are already customizing Nemotron 3.5 Lightning for specific workloads in sectors such as cybersecurity, legal services, coding, finance, and healthcare. The applications for this technology span tool calling, result validation, subagent delegation, code review routing, log triage, contract parsing, and long-context retrieval across its 1 million token context window. The model's versatility and accessibility aim to accelerate the development and deployment of sophisticated AI agents across a wide range of industries and use cases.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next