Journal

· evidence

What it took to fine-tune a small router

A pre-registered bar, crossed test sets, label checks, balanced data, and quantization probes mattered more than the first high score.

A compact routing instrument passes recorded signals through crossed evaluation benches while a broken side path remains visible.

A router earns its place through tests designed to break a flattering result. Our search chose a model for further validation. We did not ship it to customers.

The router makes a small fixed set of routing decisions. That narrow job still exposed authorship bias, interlocked classes, and a quantization failure that looked healthy from the outside.

We registered the gate first

Before any candidate ran, we registered the acceptance bar: macro accuracy at least 0.85 and every class at least 0.85. Warm p50 had to stay at or below 150 ms, with p95 at or below 400 ms.

That order kept a promising score from rewriting the definition of success.

Seven off-the-shelf small models received the same constrained task. Each had to answer with one class from the valid options. They scored 0.46 to 0.70 macro.

The models were Qwen3.5-0.8B (Apache 2.0), Qwen3.5-2B (Apache 2.0), Qwen3-0.6B (Apache 2.0), and Qwen3-1.7B (Apache 2.0). We also tested SmolLM2-360M-Instruct (Apache 2.0), SmolLM2-1.7B-Instruct (Apache 2.0), and Granite-3.3-2B-Instruct (Apache 2.0).

Qwen3.5-4B (Apache 2.0), built for the job, reached 0.9156 macro. It still missed two per-class bars and needed 1,297 ms at the median.

The useful comparison was not simply large against small. It was whether a candidate cleared every registered accuracy and latency gate.

The base-model authors are neither Muniment customers nor endorsers.

Candidate Measured macro accuracy Measured latency Gate result
Seven named off-the-shelf models, each Apache 2.0 0.46 to 0.70 Not measured here Missed the macro bar
Qwen3.5-4B, Apache 2.0 0.9156 1,297 ms at the median Missed two per-class bars and the latency bar
Fine-tuned 149-million-parameter encoder, Apache 2.0 0.9417 original, 0.9667 held-out 32.7 ms on CPU Selected for more validation

The fine-tuned 149-million-parameter encoder reached 0.9417 macro on the original set. It reached 0.9667 on a held-out set written afterward, at 32.7 ms on CPU.

These measurements were English only, and multilingual validation is under way. The selection remains open while that work runs.

A fresh test set found the flattering result

One component looked like 0.94 during the search. A held-out set written after the search returned 0.85, with two classes at 0.70.

That reversal changed how we read every candidate score. A test set can preserve the habits of its author even when no row repeats.

We crossed independently authored data against independently trained versions. Two versions each passed the test set whose data shared its authorship and failed the other.

That is fitting an author’s phrasing, not the task.

Independent judge models then relabeled the data. Zero rows had judges agreeing on a label other than the recorded one.

Those labels were sound. The model was the limit.

The classes moved together

The first repair targeted the two weakest classes with more training rows. Yet the whole result got worse because the classes interlock.

Adding rows to every class instead moved the weakest class from 0.80 to 0.85. Local repair had changed the boundaries around the entire decision space.

This is why a per-class gate mattered. Macro accuracy alone could conceal which boundary moved in the wrong direction.

A valid file can contain a broken model

Quantizing one encoder to 8-bit silently destroyed it. Cosine similarity to the original fell to 0.09, and accuracy collapsed to 0.3333.

The file still loaded and ran.

A load check therefore proved almost nothing. We needed an output-level comparison and a task-level accuracy check after conversion.

What we would change

We would write the independently authored held-out set before the candidate search. We would cross each version against every author’s test set from the start.

We would add training rows across every class before targeting the weakest classes alone. We would test similarity and accuracy after every quantization.

The registered bar stays where it belongs: before the first candidate gets a chance to flatter us.

Continue reading

All publications

Join the waitlist

Get desktop release updates.

We will email you about desktop releases and new features. muniment is a desktop workspace for your models, tools, and files.