Model routing algorithm

Choose an eligible model for the task. Count the cost of the decision, the answer, and the repair.

This page describes the routing implementation under development. Installed releases can differ until this change ships.

Measured test results

muniment / ROUTING RESEARCH

What the tests measured.

Expected policy choices matched 0 → 56 requests

  1. Jev 1.13.056 / 56
  2. Kev 4B50 / 56
  3. Nimble 9B50 / 56
  4. Decider 2B49 / 56
  5. Rizzo Flow 1.7B48 / 56

28 distinct cases × 2 option orders. Reversed orders are not independent tasks.

Synthetic policy adherence, not completed tasks. These five models ran both sets. Different runtimes limit direct comparison.

Jev: hosted. Kev: BF16 MLX. Nimble: NF4 CUDA adapter. Decider: local. Rizzo: 4-bit Metal.

Animation duration does not represent latency.

Download the results animation · Context token totals

Rules before classification

muniment filters answer accounts before it sends routing context to a classifier. Automatic routing requires known context limits and required tool and image support.

Loopback endpoints only excludes remote answer and classifier URLs. A local proxy can still forward requests. This setting cannot enforce a proxy’s network policy.

An optional estimated task budget requires a thread ID and known prices. The router reserves an uncached estimate before each attempt. An interrupted attempt keeps its reservation. A completed answer replaces that reservation with measured usage.

Budgeted tasks use deterministic selection and skip the classifier. Unknown classifier billing cannot spend the remaining budget. Provider charges, hidden tokens, and stale prices can differ from the estimate.

Keep context bounded

The classifier receives an extract of the first user request, recent user requests, recent observations, and current routing state. Tool output remains observation data.

A memory cache tracks a task revision and a digest of user requests. A background refresh commits only if both still match. An old refresh cannot replace context for a newer request.

This extract uses no model tokens to build. It can omit older requirements. The answer model still receives the original request. This is not a complete task memory or an automatic semantic change detector.

Validate each classifier

Jev is the hosted top choice. Kev 4B and Nimble are preferred candidates for further local evaluation. These labels reflect a small policy test, not proven task completion rates.

Each saved connection has its own minimum score and score type. A native confidence and an option probability have different meanings. Neither proves accuracy without evaluation on your tasks.

The router rejects foreign choices, abstentions, malformed scores, and inconsistent probability maps. A rejected result uses an eligible fallback. No classifier can add a filtered model.

Connect a classifier in Settings → Models & routing → Accounts → Connect accounts. Each preferred connection has setup instructions and upstream links.

Estimate the whole route

Selection accounts for input, output, verified cache reuse, expected repair cost, and model success estimates. A cheaper token price alone does not prove a cheaper completed task.

Cost optimization uses models with supplied success estimates above the policy floor. If estimates are absent, the router can retain the eligible current model or use the eligible selection. It does not invent a success rate.

Minimum residence and a savings threshold reduce unnecessary switches. Structured validation failures can trigger escalation. An account outage can change the account without declaring the model incapable.

Research and evidence

MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings motivates task-level utility and cost. Its trained estimator is not part of this implementation.

Pull: Lazy Materialization of Working Memory for Stateful LLM Conversations motivates selective access to history. This bounded extract does not reproduce Pull’s reversible memory system.

Read the classifier research and its limits. Download the fixtures, results, and runtime manifest.

Search docs

Join the waitlist

Get desktop release updates.

We will email you about desktop releases and new features. muniment is a desktop workspace for your models, tools, and files.