Nemotron Lightning and Switchyard separate routing from provider approval
NVIDIA introduced a 30B mixture-of-experts model and an open routing library. Routing can improve task selection, but it does not choose an approved provider or data boundary.
Historical reporting: NVIDIA published this model and routing announcement on August 11, 2026; Rangoon reviewed it on October 4, 2026.

Key takeaways
- NVIDIA presented Lightning as a 30B mixture-of-experts model for specialized agent work.
- Switchyard is an open routing library across a developer’s own mix of models.
- A routing score cannot replace an approved-provider, egress, or data-residency decision.
What NVIDIA announced[source 1]
On August 11, NVIDIA announced Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model positioned for specialized tasks within larger agent systems. The same announcement introduced NeMo Switchyard, an open-source library intended to route requests across a developer’s own mix of open, proprietary, and NVIDIA models.
NVIDIA framed the pairing as a way to select models for different jobs across PCs, workstations, data centers, and cloud environments. This retrospective reports that release description. It does not infer that any route automatically meets a customer’s provider, location, retention, or approval requirements.
Editorial analysis: useful routing starts after policy
A router can answer a narrow technical question: which already-eligible model should receive this request? That is different from deciding whether a model endpoint is allowed for the request at all. A route optimized for latency, estimated quality, capacity, or price can still be unsuitable if the selected provider lacks approval for the data class or if its egress path is outside the workload’s declared boundary.
The clean design has two stages. First, policy constructs an eligible set from the request’s data classification, project, region, vendor approval, tool needs, and retention rules. Second, a router chooses within that set using task signals. The router never receives a candidate that policy has excluded. This avoids turning a technical fallback into an unreviewed transfer to another provider.
A bounded engineering example records the eligible endpoint set and policy revision before inference. If the router picks an endpoint, the run receipt records the selected target, selection reason, and effective configuration. If no endpoint qualifies, the request stops with a reason instead of treating absence of a route as permission to broaden the set.
- Construct eligible providers before evaluating routing preferences.
- Preserve endpoint selection and policy revision in the run record.
- Require an explicit change when a provider or location enters the candidate set.
Integration and evidence boundary
Rangoon’s launch architecture separates compatibility records, runtime configuration, and external action authority. A future router adapter could make an approved candidate set inspectable, but the public architecture does not claim a shipped Switchyard adapter, model route, automatic data residency, or inferred provider approval.
LNSAT’s reference execution path addresses a different question: whether one exact external action may proceed. It can bind requested target and configuration evidence into an action packet, then retain approval and receipt data separately from a model-selection heuristic. An inference route and connector side effect are both boundaries that need their own evidence.
The vendor announcement supplies one routing design reference; an organization still owns the policy that defines permissible destinations and the evidence that proves a selected configuration was used. Routing telemetry should be retained only at the level needed to explain selection, test failures, and reviewed policy changes.
Sources
- NVIDIA: Nemotron 3.5 Lightning and NeMo Switchyard (August 11, 2026)

