Nvidia Releases Nemotron 3.5 Lightning and NeMo Switchyard to Attack the Cost of AI Agents
Nvidia introduced Nemotron 3.5 Lightning and NeMo Switchyard, an experimental routing layer designed to send AI requests to different models based on quality, latency and cost.
Why it matters: Model routing is emerging as a control plane for AI-agent cost and quality, giving Nvidia another strategic layer above accelerated compute.
Nvidia is expanding its AI software strategy with Nemotron 3.5 Lightning and NeMo Switchyard. Switchyard acts as a routing decision layer that can direct requests to different model backends according to factors such as quality, latency and cost. The architecture could let enterprises reserve expensive frontier models for harder tasks while routing simpler requests to cheaper models. Nvidia's current documentation labels Switchyard experimental, so it should not yet be treated as a mature production standard.