Model routing algorithm
Choose an eligible model for the task. Count the cost of the decision, the answer, and the repair.
This page describes the routing implementation under development. Installed releases can differ until this change ships.
Measured test results
What the tests measured.
Expected policy choices matched 0 → 56 requests
- Jev 1.13.056 / 56
- Kev 4B50 / 56
- Nimble 9B50 / 56
- Decider 2B49 / 56
- Rizzo Flow 1.7B48 / 56
28 distinct cases × 2 option orders. Reversed orders are not independent tasks.
Synthetic policy adherence, not completed tasks. These five models ran both sets. Different runtimes limit direct comparison.
Jev: hosted. Kev: BF16 MLX. Nimble: NF4 CUDA adapter. Decider: local. Rizzo: 4-bit Metal.
Expected policy choices matched 0 → 24 requests
- Jev 1.13.024 / 24
- Kev 4B24 / 24
- Nimble 9B22 / 24
- Decider 2B23 / 24
- Rizzo Flow 1.7B23 / 24
12 distinct cases × 2 option orders. Reversed orders are not independent tasks.
Synthetic policy adherence, not completed tasks. These five models ran both sets. Different runtimes limit direct comparison.
Jev: hosted. Kev: BF16 MLX. Nimble: NF4 CUDA adapter. Decider: local. Rizzo: 4-bit Metal.
Expected policy choices matched 0 → 32 requests
- Jev 1.13.032 / 32
- Kev 4B26 / 32
- Nimble 9B28 / 32
- Decider 2B26 / 32
- Rizzo Flow 1.7B25 / 32
16 harder cases × 2 option orders. Reversed orders are not independent tasks.
Synthetic policy adherence, not completed tasks. These five models ran both sets. Different runtimes limit direct comparison.
Jev: hosted. Kev: BF16 MLX. Nimble: NF4 CUDA adapter. Decider: local. Rizzo: 4-bit Metal.
Total model tokens 0 → 34,470 tokens
- Full context34,47018 / 40 correct checkpoints · 0 summary updates
- Summary after each reply17,04624 / 40 correct checkpoints · 40 summary updates
- Summary at known changes5,02937 / 40 correct checkpoints · 8 summary updates
Qwen3-0.6B (4-bit), four synthetic conversations. Totals include summary calls. Lower token counts alone do not establish better routing.
Known changes used fixture annotations. We did not test an automatic change detector. This is not a product savings claim.
Animation duration does not represent latency.
Rules before classification
muniment filters answer accounts before it sends routing context to a classifier. Automatic routing requires known context limits and required tool and image support.
Loopback endpoints only excludes remote answer and classifier URLs. A local proxy can still forward requests. This setting cannot enforce a proxy’s network policy.
An optional estimated task budget requires a thread ID and known prices. The router reserves an uncached estimate before each attempt. An interrupted attempt keeps its reservation. A completed answer replaces that reservation with measured usage.
Budgeted tasks use deterministic selection and skip the classifier. Unknown classifier billing cannot spend the remaining budget. Provider charges, hidden tokens, and stale prices can differ from the estimate.
Keep context bounded
The classifier receives an extract of the first user request, recent user requests, recent observations, and current routing state. Tool output remains observation data.
A memory cache tracks a task revision and a digest of user requests. A background refresh commits only if both still match. An old refresh cannot replace context for a newer request.
This extract uses no model tokens to build. It can omit older requirements. The answer model still receives the original request. This is not a complete task memory or an automatic semantic change detector.
Validate each classifier
Jev is the hosted top choice. Kev 4B and Nimble are preferred candidates for further local evaluation. These labels reflect a small policy test, not proven task completion rates.
Each saved connection has its own minimum score and score type. A native confidence and an option probability have different meanings. Neither proves accuracy without evaluation on your tasks.
The router rejects foreign choices, abstentions, malformed scores, and inconsistent probability maps. A rejected result uses an eligible fallback. No classifier can add a filtered model.
Connect a classifier in Settings → Models & routing → Accounts → Connect accounts. Each preferred connection has setup instructions and upstream links.
Estimate the whole route
Selection accounts for input, output, verified cache reuse, expected repair cost, and model success estimates. A cheaper token price alone does not prove a cheaper completed task.
Cost optimization uses models with supplied success estimates above the policy floor. If estimates are absent, the router can retain the eligible current model or use the eligible selection. It does not invent a success rate.
Minimum residence and a savings threshold reduce unnecessary switches. Structured validation failures can trigger escalation. An account outage can change the account without declaring the model incapable.
Research and evidence
MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings motivates task-level utility and cost. Its trained estimator is not part of this implementation.
Pull: Lazy Materialization of Working Memory for Stateful LLM Conversations motivates selective access to history. This bounded extract does not reproduce Pull’s reversible memory system.
Read the classifier research and its limits. Download the fixtures, results, and runtime manifest.