# Routing needs more than a classifier

[Journal](/journal/)

September 26, 2026 · [evidence](/journal/#evidence)



A small classifier test, two papers, and a routing design that checks constraints before it asks a model.

![A request line passes three rule gates that turn other branches away, then a stack of context cards and a classifier dial, and only one verdigris route reaches a model socket.](/_astro/routing-needs-more-than-a-classifier.jtO0zlks_ZdeDWg.avif)

Jev matched 56 / 56 expected choices in our small routing test. It is our hosted starting preference, not proof of the cheapest model that completes a task.

We tested routing rules, not the selected answer models. The design checks rules first and treats classifier answers as proposals.

muniment / ROUTING RESEARCH

What the tests measured.

Expected policy choices matched 0 → 56 requests

1.  **Jev 1.13.0**56 / 56
    
2.  **Kev 4B**50 / 56
    
3.  **Nimble 9B**50 / 56
    
4.  **Decider 2B**49 / 56
    
5.  **Rizzo Flow 1.7B**48 / 56
    

28 distinct cases × 2 option orders. Reversed orders are not independent tasks.

Synthetic policy adherence, not completed tasks. These five models ran both sets. Different runtimes limit direct comparison.

Jev: hosted. Kev: BF16 MLX. Nimble: NF4 CUDA adapter. Decider: local. Rizzo: 4-bit Metal.

Expected policy choices matched 0 → 24 requests

1.  **Jev 1.13.0**24 / 24
    
2.  **Kev 4B**24 / 24
    
3.  **Nimble 9B**22 / 24
    
4.  **Decider 2B**23 / 24
    
5.  **Rizzo Flow 1.7B**23 / 24
    

12 distinct cases × 2 option orders. Reversed orders are not independent tasks.

Synthetic policy adherence, not completed tasks. These five models ran both sets. Different runtimes limit direct comparison.

Jev: hosted. Kev: BF16 MLX. Nimble: NF4 CUDA adapter. Decider: local. Rizzo: 4-bit Metal.

Expected policy choices matched 0 → 32 requests

1.  **Jev 1.13.0**32 / 32
    
2.  **Kev 4B**26 / 32
    
3.  **Nimble 9B**28 / 32
    
4.  **Decider 2B**26 / 32
    
5.  **Rizzo Flow 1.7B**25 / 32
    

16 harder cases × 2 option orders. Reversed orders are not independent tasks.

Synthetic policy adherence, not completed tasks. These five models ran both sets. Different runtimes limit direct comparison.

Jev: hosted. Kev: BF16 MLX. Nimble: NF4 CUDA adapter. Decider: local. Rizzo: 4-bit Metal.

Total model tokens 0 → 34,470 tokens

1.  **Full context**34,470
    
    18 / 40 correct checkpoints · 0 summary updates
2.  **Summary after each reply**17,046
    
    24 / 40 correct checkpoints · 40 summary updates
3.  **Summary at known changes**5,029
    
    37 / 40 correct checkpoints · 8 summary updates

Qwen3-0.6B (4-bit), four synthetic conversations. Totals include summary calls. Lower token counts alone do not establish better routing.

**Known changes used fixture annotations.** We did not test an automatic change detector. This is not a product savings claim.

[Inspect the test evidence](/research/model-routing/evidence.json)

Animation duration does not represent latency.

[Download the results animation](/research/model-routing/routing.mp4) · [Context token totals](/research/model-routing/context-summary.csv)

## Two papers changed the unit of the decision

[MTRouter](https://arxiv.org/abs/2604.23530) combines history and candidate model embeddings across turns. Its outcome estimator learns from logged trajectories to connect routes with task utility and cost.

[Pull](https://arxiv.org/abs/2609.14773) keeps an addressable history directory. Models retrieve needed turns. Other turns stay collapsed but accessible, unlike a summary that permanently drops details.

Neither probe reproduces the papers’ benchmarks or savings. The MTRouter probe tests an embedding model, not its trained routing estimator. The Pull probe tests a small local pipeline.

## Jev led this policy test

We tested 12 initial and 16 harder synthetic cases in two option orders: 28 distinct cases, 56 requests. Reversed options are not independent tasks.

Cases cover routine requests, deeper reasoning, privacy, quoted instructions, and task changes. The harder set exposed more failures.

| Classifier | Initial requests | Harder requests | Runtime tested |
| --- | --- | --- | --- |
| Jev 1.13.0 | 24 / 24 | 32 / 32 | Hosted TypeSafe API |
| Kev 4B | 24 / 24 | 26 / 32 | BF16, MLX |
| Nimble 9B | 22 / 24 | 28 / 32 | NF4, CUDA, unmerged adapter |
| Decider 2B | 23 / 24 | 26 / 32 | Local runtime |
| Rizzo Flow 1.7B | 23 / 24 | 25 / 32 | Q4\_K\_M, llama.cpp Metal |
| SemIf 4B | 15 / 24 | Not run | 4-bit, pinned MLX runtime |
| NanoJev | 10 / 24 | Not run | CUDA |
| Laya | 9 / 24 | Not run | English checkpoint |
| Von | 8 / 24 | Not run | Local runtime |
| Needle 3 | 3 / 24 | Not run | Per-option tool adapter |

Needle produced one valid choice on 13 requests. The other 11 produced no single choice. Its SDK’s default macOS runtime returned 404, so we used the published 3.0.1 runtime. An enum adapter scored zero. These results test the adapters as well as the model.

A typed Laya request scored 11 / 24. Kev 0.8B scored 16 / 24. An Ollama Qwen3.5 2B baseline scored 19 / 24. These smaller probes did not receive the harder set.

We used a shared M2 Max with 32 GB and NVIDIA RTX 4000 SFF Ada with 20 GB. Latency on these machines cannot rank production speed.

Kev 4B and Nimble each matched 50 / 56 requests across both sets. Their errors differed, and Nimble missed quoted-document cases. Our Nimble quantization and adapter setup also differ from the upstream full-precision serving path.

Muniment marks Jev as **Top choice**, with Kev 4B and Nimble as **Preferred** candidates. The sample does not establish a statistically reliable local winner. Download the [fixtures, per-request results, and runtime manifest](/research/model-routing/evidence.json) to inspect each test.

No organization named here is presented as a Muniment customer or endorser.

## A summary after every reply has a cost

The context experiment used Qwen3-0.6B on four synthetic conversations with ten checkpoints each. Token totals include selector and summary work.

| Context method | Correct checkpoints | Total tokens | Summary updates |
| --- | --- | --- | --- |
| Full context | 18 / 40 | 34,470 | 0 |
| Summary after each reply | 24 / 40 | 17,046 | 40 |
| Summary at known change points | 37 / 40 | 5,029 | 8 |

Fixture annotations supplied the change points. We tested no automatic detector. These results are not product savings claims.

A final-query Pull pipeline probe answered 3 / 4 cases with 8,032 total tokens. Full context answered 1 / 4 with 6,747 tokens. Selector overhead erased token savings in that probe. An embedding component found the needed evidence in 4 / 4 cases, but that alone says nothing about trained routing utility.

## Constraints stay outside the classifier

The implementation filters endpoints, known capabilities, context limits, and estimated task budget before classification. It validates the returned choice against that filtered set. A score cannot override those rules.

Each connection has its own score type and threshold. Nimble’s option probabilities and Jev’s native confidence are not interchangeable accuracy estimates. An abstention, invalid choice, malformed score, or inconsistent probability map sends the request to an eligible fallback.

The router uses a bounded extract of user requests and recent observations. A background refresh checks the task revision before it commits. An old refresh cannot replace a newer task’s context. This deterministic extract needs no summary-model call, but it can omit older requirements.

Cost selection includes known input and output prices, verified cache reuse, repair estimates, and supplied success estimates. It does not create success estimates from this classifier test. Switching also has a minimum residence and a savings threshold. Budgeted tasks skip classifier calls with unbounded billing.

This implementation is under development. Installed releases can differ until it ships. The [model routing documentation](/docs/model-routing/) states the current limits and connection steps.

Next, test a recorded task that forces a repair. Count every call and verify task completion.

## Sources

1.  [MTRouter: Cost-Aware Multi-Turn LLM Routing with History-Model Joint Embeddings](https://arxiv.org/abs/2604.23530) arxiv.org
2.  [Pull: Lazy Materialization of Working Memory for Stateful LLM Conversations](https://arxiv.org/abs/2609.14773) arxiv.org
3.  [Kev source and setup](https://github.com/jaredpalmer/kev) github.com
4.  [Nimble source and model contract](https://github.com/bespokelabsai/nimble) github.com
