NVIDIA AI Releases Nemotron 3.5 Lightning: A 30B Open MoE with 3B Active Parameters, and NeMo Switchyard Model Router
NVIDIA has released Nemotron 3.5 Lightning, a 30-billion-parameter Mixture-of-Experts (MoE) model with only 3 billion active parameters, designed for efficient agent execution. Alongside it, the NeMo Switchyard Model Router dynamically routes each reasoning step to the most cost-effective capable model, reducing compute overhead without sacrificing performance. This release emphasizes open access and practical deployment, targeting developers building AI agents that need to balance capability, speed, and cost. The model is optimized for NVIDIA hardware and integrates with the NeMo framework, enabling scalable, low-latency AI workflows.
What this means for you
Developers and AI teams should evaluate Nemotron 3.5 Lightning for agent-based applications where cost-efficient inference and modular routing can significantly reduce operational expenses.
Put this to work
Try the tools in this story inside the AI Tool Lab, browse roles in the AI Job Portal, test your skills with Aptitude Tests, or read our in-depth guides.
Keep reading
More AI news
Anthropic CEO says it’s time to pump the brakes on AI
Anthropic CEO Dario Amodei has proposed a strategic shift in AI development, advocating for a measured pace rather than an unchecked race for intelligence. By inviting independent third-party assessors like METR to audit their systems, the company aims to establish a new standard for transparency and safety. This move signals a significant departure from the 'growth-at-all-costs' mentality, suggesting that the industry must prioritize verifiable safety protocols before deploying more powerful models to the public.
Read the briefAnthropic CEO outlines plan to ‘pace the frontier’
Anthropic CEO Dario Amodei has proposed a strategic shift aimed at decelerating the breakneck speed of artificial intelligence development. By advocating for a more deliberate 'pacing' of innovation, the company suggests that industry leaders should prioritize safety evaluations and societal impact assessments over mere technical capability benchmarks. This shift marks a significant departure from the competitive race between major AI labs, highlighting growing concerns that rapid deployment without adequate guardrails could pose existential or systemic risks to global infrastructure.
Read the briefY Combinator’s Garry Tan wants U.S. open-weight AI labs to ‘distill’ frontier models, too
Y Combinator leader Garry Tan is advocating for a shift in how leading artificial intelligence developers share their technology. He contends that since foundational models are built upon vast repositories of human-generated information, the resulting capabilities should be accessible as a public benefit. By encouraging labs to create smaller, distilled versions of frontier models, Tan hopes to democratize access to high-performance AI, moving away from closed-off ecosystems toward a more equitable distribution of innovative tools.
Read the brief