Nvidia Nemotron 3.5 Lightning and NeMo Switchyard

246 points · 123 comments on HN · read original →

Points and comments are a snapshot, not live.

Nvidia releases Nemotron 3.5 Lightning and NeMo Switchyard for efficient agentic AI routing.

Nvidia expanded its Nemotron 3 family with Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model for specialized agentic tasks. It claims up to 4x faster output and 30% faster task completion vs. other models in its class. Separately, Nvidia released NeMo Switchyard, an open-source library for routing prompts to the most suitable model. Partners like Boomi, Cadence, Cognition, and LangChain tested Switchyard, reporting cost reductions of 21-58% while maintaining accuracy. Nemotron 3.5 Lightning is available on Hugging Face and other platforms; Switchyard is on GitHub.

What commenters are saying

Commenters focused on technical questions about model routing and caching. The top thread asked how NeMo Switchyard handles prompt caching when routing between models, noting that KV caches are model-specific and routing could break cache efficiency. Some argued smart routing is snake oil for this reason. Another comment noted the README labels Switchyard "Experimental software. Not for production use," creating confusion with the press release. Separate threads compared Nemotron 3.5 Lightning's benchmarks to Meta's Muse Glimmer 30B, with one commenter noting Lightning is a sparse MoE model with far fewer active parameters than Glimmer, making direct speed comparisons misleading.