Routing needs more than a classifier
A small classifier test, two papers, and a routing design that checks constraints before it asks a model.

Jev matched 56 / 56 expected choices in our small routing test. It is our hosted starting preference, not proof of the cheapest model that completes a task.
We tested routing rules, not the selected answer models. The design checks rules first and treats classifier answers as proposals.
What the tests measured.
Expected policy choices matched 0 → 56 requests
- Jev 1.13.056 / 56
- Kev 4B50 / 56
- Nimble 9B50 / 56
- Decider 2B49 / 56
- Rizzo Flow 1.7B48 / 56
28 distinct cases × 2 option orders. Reversed orders are not independent tasks.
Synthetic policy adherence, not completed tasks. These five models ran both sets. Different runtimes limit direct comparison.
Jev: hosted. Kev: BF16 MLX. Nimble: NF4 CUDA adapter. Decider: local. Rizzo: 4-bit Metal.
Expected policy choices matched 0 → 24 requests
- Jev 1.13.024 / 24
- Kev 4B24 / 24
- Nimble 9B22 / 24
- Decider 2B23 / 24
- Rizzo Flow 1.7B23 / 24
12 distinct cases × 2 option orders. Reversed orders are not independent tasks.
Synthetic policy adherence, not completed tasks. These five models ran both sets. Different runtimes limit direct comparison.
Jev: hosted. Kev: BF16 MLX. Nimble: NF4 CUDA adapter. Decider: local. Rizzo: 4-bit Metal.
Expected policy choices matched 0 → 32 requests
- Jev 1.13.032 / 32
- Kev 4B26 / 32
- Nimble 9B28 / 32
- Decider 2B26 / 32
- Rizzo Flow 1.7B25 / 32
16 harder cases × 2 option orders. Reversed orders are not independent tasks.
Synthetic policy adherence, not completed tasks. These five models ran both sets. Different runtimes limit direct comparison.
Jev: hosted. Kev: BF16 MLX. Nimble: NF4 CUDA adapter. Decider: local. Rizzo: 4-bit Metal.
Total model tokens 0 → 34,470 tokens
- Full context34,47018 / 40 correct checkpoints · 0 summary updates
- Summary after each reply17,04624 / 40 correct checkpoints · 40 summary updates
- Summary at known changes5,02937 / 40 correct checkpoints · 8 summary updates
Qwen3-0.6B (4-bit), four synthetic conversations. Totals include summary calls. Lower token counts alone do not establish better routing.
Known changes used fixture annotations. We did not test an automatic change detector. This is not a product savings claim.
Animation duration does not represent latency.
Two papers changed the unit of the decision
MTRouter combines history and candidate model embeddings across turns. Its outcome estimator learns from logged trajectories to connect routes with task utility and cost.
Pull keeps an addressable history directory. Models retrieve needed turns. Other turns stay collapsed but accessible, unlike a summary that permanently drops details.
Neither probe reproduces the papers’ benchmarks or savings. The MTRouter probe tests an embedding model, not its trained routing estimator. The Pull probe tests a small local pipeline.
Jev led this policy test
We tested 12 initial and 16 harder synthetic cases in two option orders: 28 distinct cases, 56 requests. Reversed options are not independent tasks.
Cases cover routine requests, deeper reasoning, privacy, quoted instructions, and task changes. The harder set exposed more failures.
| Classifier | Initial requests | Harder requests | Runtime tested |
|---|---|---|---|
| Jev 1.13.0 | 24 / 24 | 32 / 32 | Hosted TypeSafe API |
| Kev 4B | 24 / 24 | 26 / 32 | BF16, MLX |
| Nimble 9B | 22 / 24 | 28 / 32 | NF4, CUDA, unmerged adapter |
| Decider 2B | 23 / 24 | 26 / 32 | Local runtime |
| Rizzo Flow 1.7B | 23 / 24 | 25 / 32 | Q4_K_M, llama.cpp Metal |
| SemIf 4B | 15 / 24 | Not run | 4-bit, pinned MLX runtime |
| NanoJev | 10 / 24 | Not run | CUDA |
| Laya | 9 / 24 | Not run | English checkpoint |
| Von | 8 / 24 | Not run | Local runtime |
| Needle 3 | 3 / 24 | Not run | Per-option tool adapter |
Needle produced one valid choice on 13 requests. The other 11 produced no single choice. Its SDK’s default macOS runtime returned 404, so we used the published 3.0.1 runtime. An enum adapter scored zero. These results test the adapters as well as the model.
A typed Laya request scored 11 / 24. Kev 0.8B scored 16 / 24. An Ollama Qwen3.5 2B baseline scored 19 / 24. These smaller probes did not receive the harder set.
We used a shared M2 Max with 32 GB and NVIDIA RTX 4000 SFF Ada with 20 GB. Latency on these machines cannot rank production speed.
Kev 4B and Nimble each matched 50 / 56 requests across both sets. Their errors differed, and Nimble missed quoted-document cases. Our Nimble quantization and adapter setup also differ from the upstream full-precision serving path.
Muniment marks Jev as Top choice, with Kev 4B and Nimble as Preferred candidates. The sample does not establish a statistically reliable local winner. Download the fixtures, per-request results, and runtime manifest to inspect each test.
No organization named here is presented as a Muniment customer or endorser.
A summary after every reply has a cost
The context experiment used Qwen3-0.6B on four synthetic conversations with ten checkpoints each. Token totals include selector and summary work.
| Context method | Correct checkpoints | Total tokens | Summary updates |
|---|---|---|---|
| Full context | 18 / 40 | 34,470 | 0 |
| Summary after each reply | 24 / 40 | 17,046 | 40 |
| Summary at known change points | 37 / 40 | 5,029 | 8 |
Fixture annotations supplied the change points. We tested no automatic detector. These results are not product savings claims.
A final-query Pull pipeline probe answered 3 / 4 cases with 8,032 total tokens. Full context answered 1 / 4 with 6,747 tokens. Selector overhead erased token savings in that probe. An embedding component found the needed evidence in 4 / 4 cases, but that alone says nothing about trained routing utility.
Constraints stay outside the classifier
The implementation filters endpoints, known capabilities, context limits, and estimated task budget before classification. It validates the returned choice against that filtered set. A score cannot override those rules.
Each connection has its own score type and threshold. Nimble’s option probabilities and Jev’s native confidence are not interchangeable accuracy estimates. An abstention, invalid choice, malformed score, or inconsistent probability map sends the request to an eligible fallback.
The router uses a bounded extract of user requests and recent observations. A background refresh checks the task revision before it commits. An old refresh cannot replace a newer task’s context. This deterministic extract needs no summary-model call, but it can omit older requirements.
Cost selection includes known input and output prices, verified cache reuse, repair estimates, and supplied success estimates. It does not create success estimates from this classifier test. Switching also has a minimum residence and a savings threshold. Budgeted tasks skip classifier calls with unbounded billing.
This implementation is under development. Installed releases can differ until it ships. The model routing documentation states the current limits and connection steps.
Next, test a recorded task that forces a repair. Count every call and verify task completion.