The model picker is not the traffic path
Kong and NVIDIA keep the model picker off the traffic path. OpenThoughts-TBLite reports the cost of that split, not the placement.

The picker and the traffic path are different jobs. Kong’s August 11, 2026 note shows why that split belongs in production routing.
The traffic path is the door that sends the request, holds keys, and records who did what and when. A picker names which model may run.
Kong kept the picker off the traffic path, while NVIDIA NeMo Switchyard picks the model as a side service. Kong then sends the request, holds keys, applies limits, and writes the audit log, because a picker on the traffic path sees prompts and holds keys.
Kong, NVIDIA, Microsoft, Slack, and Kai Waehner are public evidence. None is a Muniment customer or endorser.
Kong moved the costly share with one threshold
A confidence threshold is the score a picker must beat before it sends work to the costly model. Kong raised the stage_router confidence threshold from 0.3 to 0.5 on OpenThoughts-TBLite, a set of 20 difficulty-calibrated agent tasks.
Kong reports these OpenThoughts-TBLite figures from its August 11, 2026 note.
| Kong measurement | Result |
|---|---|
| Confidence threshold | Raised from 0.3 to 0.5 |
| Escalations to the costly model | Fell from 85 percent to 17 percent |
| Cost per completed task | 43.7 percent lower |
| Cheap model | GLM-5.2 |
| Costly model | Claude Opus 4.8 |
| Benchmark | OpenThoughts-TBLite |
Terminal-Bench 2 remains an open check, not a measurement in this table.
Kong ran GLM-5.2 as the cheap model and Claude Opus 4.8 as the costly model, while escalations to the costly model fell from 85 percent to 17 percent. Cost per completed task fell 43.7 percent.
Switchyard can sit beside the host path
NVIDIA published NeMo Switchyard on August 11, 2026. Its library can sit beside a host traffic path. NVIDIA names LangChain, Cognition, and Kong in that post.
A first-party NVIDIA video walks the live-signal design. NVIDIA’s post presents that walkthrough as live-signal routing across a configured model pool.
GitHub hosts the code under Apache 2.0. NVIDIA still marks Switchyard experimental and not for production use.
The process gate names the allowed set
Kai Waehner wrote on August 17, 2026, and says the process gate should name the allowed set. The traffic path should route inside that set, an approach he names decide high, route low.
Jurisdiction is the legal place whose rules apply. Waehner puts jurisdiction in the process gate, not in the router. He says a Microsoft router reports which model it chose but not why.
Slack Engineering tests a new model after security and compliance verifications. Its routing layer diverts traffic to a healthy alternative when an endpoint degrades. Waehner reads this as a production version of the same split, with failover inside that cleared set.
Records still need both jobs
Earlier routing notes cover cost per completed task. They do not record this layer split or these Kong figures. A classifier can fail closed downstream even when its own boundary holds. Keys and records belong on a control plane.
Muniment’s position is that the picker names the model and the traffic path holds the keys. Keep those jobs apart.
If your records need to show who named the model and who sent the request, join the waitlist.
Sources
- Kong: Intelligent Model Routing: Kong AI Gateway Applies NVIDIA NeMo Switchyard Across Model Traffic konghq.com
- NVIDIA: Route AI Agents Across Models with NVIDIA NeMo Switchyard developer.nvidia.com
- NVIDIA-NeMo on GitHub: Switchyard github.com
- NVIDIA: Why AI Agents Need More Than One Model youtu.be
- Kai Waehner: Multi-Model AI Orchestration: Which Layer Should Pick the Model? www.kai-waehner.de
- Slack Engineering: Slack AI: The Path to Multi-Cloud slack.engineering