Local inference is becoming a real routing tier
Large open-weight models can now run on local machines, which gives governed routing policies a credible on-device tier with hard operational limits.

Local execution has become credible as the lowest tier in a governed routing policy. When an organization’s policy or contract requires work to remain on-device, a router can now send some jobs to very large open-weight models without pretending that every workload belongs there.
The evidence is narrow. DeepSeek publishes V4 Flash as a 284-billion-parameter mixture-of-experts model with 13 billion parameters active and MIT-licensed weights. Z.ai publishes GLM-5.2 weights under the MIT license. Two project authors have built local runtimes around those weights. Their measurements show feasibility on personal machines, not equivalence to closed frontier systems and not independent evaluation.
The floor now has evidence
DS4 uses an asymmetric 2-bit scheme that quantizes only routed mixture-of-experts weights while leaving shared experts, projections, and routing untouched. Its current documentation says the 2-bit Flash quant is 81 GB. That footprint fits a high-memory laptop class without reducing the claim to a hardware shopping list.
Colibri takes another route: it keeps GLM-5.2’s dense portion resident and streams routed experts from a roughly 370 GB int4 model on local storage. The consequence is a much lower demonstrated RAM floor and a much slower cold decode.
Meta describes Muse Glimmer as a 30-billion-parameter model for local agent workflows. Its Apache 2.0 weights are downloadable now. Meta said optimized integrations with llama.cpp, MLX, ExecuTorch, and other tools would arrive later.
Meta reports that quantization compresses the language model to under 20 GB by storing its weights with approximately four bits each. Meta says this footprint leaves room within 24 GB or 32 GB for the model’s working memory, called a KV cache, plus two companion components. One component uses speculative decoding, where a small drafter proposes blocks of text before the main model checks them. Meta reports the table’s speeds on M4 Max, M5 Max, and RTX 5090 hardware at batch size one with greedy decoding.
FreeToken’s authors treat a personal machine as one inference platform. They split each model’s work across the graphics processor, the main processor, and system memory. A mixture-of-experts design stores many specialist networks and uses only a few of them for each small unit of generated text, called a token. They state that partial results from the graphics processor and the main processor merge exactly, with no algorithmic approximation.
They report Qwen3.6-35B-A3B under NVFP4 at 39.3 tokens per second on a coding workload on an RTX 4060 Laptop. That laptop has 8 GB of graphics memory and 32 GB of LPDDR5 system memory. They report that DeepSeek-V4-Flash at 284 billion parameters runs on an RTX 5090 Desktop. That desktop has 32 GB of graphics memory and 192 GB of DDR5 system memory.
On an RTX PRO 6000 workstation they report GLM-5.2 at 753 billion parameters at 14.9 tokens per second on a math workload. They report 7.3 tokens per second for llama.cpp on the same workstation. That workstation has 96 GB of graphics memory and 512 GB of DDR5 system memory. Colibri’s table row keeps GLM-5.2 at 744 billion parameters. Each count belongs to its own source. The paper reports no answer-quality benchmark and no hardware price.
| Model and source | Scale | Active | Machine | Footprint | Speed | Status |
|---|---|---|---|---|---|---|
| DeepSeek V4 Flash via DS4 | 284B | 13B | M3 Max MacBook Pro, 128 GB | 81 GB q2 | 26.68 tok/s generation | measured |
| GLM-5.2 via Colibri | 744B | roughly 40B | WSL2, 12 cores, 25 GB RAM, NVMe via VHDX | roughly 370 GB int4; roughly 20 GB peak RSS | roughly 0.05 to 0.1 tok/s cold | measured |
| GLM-5.2 via Colibri | 744B | roughly 40B | Native Linux, 32 GB RAM, PCIe 4 NVMe | roughly 370 GB int4 | roughly 0.5 to 1 tok/s | project estimate |
| Nemotron-H 120B via RotaryQuant | 120B | not stated | Apple M4 Max, 128 GB | 17.2 GB peak within a 32 GB target budget | 14.85 tok/s | preprint measurement |
| Muse Glimmer via Meta builds | 30B | not stated | MacBook M4 Max MacBook M5 Max Nvidia RTX 5090 |
under 20 GB in a 24 GB or 32 GB envelope | 37.8 tok/s 50.2 tok/s 233.4 tok/s |
Meta report |
| Qwen3.6-35B-A3B via FreeToken | 35B | 3B | RTX 4060 Laptop 8 GB graphics memory 32 GB LPDDR5 system memory |
NVFP4 | 39.3 tok/s | preprint measurement |
| DeepSeek-V4-Flash via FreeToken | 284B | 13B | RTX 5090 Desktop 32 GB graphics memory 192 GB DDR5 system memory |
MXFP4 experts | not stated | preprint measurement |
| GLM-5.2 via FreeToken | 753B | 40B | RTX PRO 6000 96 GB graphics memory 512 GB DDR5 system memory |
NVFP4, 433 GB checkpoint | 14.9 tok/s | preprint measurement |
DS4 labels the 26.68 tok/s figure as a single-run Metal CLI measurement on a 128 GB M3 Max MacBook Pro using q2, a short prompt, a 32,768-token context setting, non-thinking mode, greedy decoding, and 256 generated tokens. It is a project-author benchmark. It should travel with all of those conditions.
Colibri’s project-author baseline used WSL2, 12 cores, 25 GB RAM, and NVMe through VHDX; the project reports roughly 20 GB peak RSS and roughly 0.05 to 0.1 tok/s for cold decode. That documentation puts native Linux with 32 GB RAM and PCIe 4 NVMe at roughly 0.5 to 1 tok/s, then explicitly calls that row an estimate rather than a measurement. This estimate belongs in capacity planning as a hypothesis to test, never as a proven baseline.
Runtime choice moves each phase differently
Generation speed covers only the phase that writes the answer. Prompt processing reads the request and prepares its context before the first answer token appears.
BaseRT’s authors tested fifteen model setups on an Apple M5 Pro, from fewer than 1 billion stored values to 35 billion. They report up to 6.4 times llama.cpp’s prompt-processing throughput and 3.9 times MLX’s throughput. For answer generation, they report smaller leads: up to 1.75 times llama.cpp’s throughput and 1.33 times MLX’s throughput.
Those split results expose a blind spot in a generation-only table. Runtime choice can move request preparation and answer writing by very different amounts. Mobile measurements show the same need for phase-aware scheduling.
FreeToken’s authors report worst-case time to first token below 44 seconds on an RTX 5090. They report up to 232 seconds for llama.cpp, up to 179 seconds for Ollama, and up to 946 seconds for KTransformers.
BaseRT’s command-line tool and bindings carry the Apache-2.0 license. Its prebuilt engine binary carries a separate proprietary license. That boundary does not make the engine available for inspection.
The RotaryQuant preprint tests compression, which stores the same model information in less memory. A parameter is one stored value that the model learned. Its authors measured Nemotron-H 120B at a 17.2 GB peak against a 32 GB target budget. Every experiment ran on one Apple M4 Max with 128 GB of unified memory. That machine therefore held four times the target budget.
Table 4 reports 12.85 tokens per second for Gemma 4-26B-A4B, 14.85 for Nemotron-H 120B, and 15.6 for Qwen3.6-35B-A3B. A token per second counts how many pieces of text the model writes each second. The abstract gives a 9 to 19 tokens-per-second range, but the full text maps neither endpoint to a model.
A quality gate used 12 prompts. The paper measured perplexity with a 2,048-token context. Perplexity measures how surprised a model is by the next piece of text. Section 4.4 reports a 5.3 times throughput cost when the model already fits in memory. Section 5 calls the method counterproductive when the model fits comfortably.
The preprint reports no production adoption. It says production serving would need a purpose-built inference server. RotaryQuant’s authors and their institutions are neither Muniment customers nor endorsers. Those institutions are Cognizant, Southern Methodist University, and the University of Maryland, Baltimore County.
A routing tier has a narrow job
These results change the routing menu. A policy can reserve a local tier for data that must stay on-device, then keep larger or higher-throughput workloads in the data center. Local becomes an enforceable placement choice, not a claim that one machine should absorb the fleet.
Quality claims need the same discipline. DS4’s author writes that V4 Flash “feels quasi-frontier”. That is the author’s subjective impression. It is neither a benchmark result nor evidence that the model matches a closed frontier model.
The operational limits are substantial. Colibri’s proven baseline is slow because cold inference streams expert weights from storage, and both approaches demand large model footprints and meaningful RAM. DS4 currently describes its engine and model files as beta quality and its agent as alpha quality; Colibri describes a roughly 370 GB int4 model with roughly 20 GB peak RSS on its baseline. Neither project documentation resolves the training-data provenance of the underlying weights. Data-center inference keeps the larger and higher-throughput work.
DeepSeek, Z.ai, DS4, Colibri, BaseRT, BaseRT’s authors, Meta, FreeToken, and FreeToken’s authors are neither Muniment customers nor endorsers. This evidence does not show that Apple uses or endorses BaseRT. Their public records establish something useful and smaller: governed routing now has a local floor worth testing.
Sources
- DeepSeek on Hugging Face: DeepSeek-V4-Flash huggingface.co
- antirez on GitHub: ds4: DeepSeek 4 Flash and PRO local inference engine github.com
- JustVugg on GitHub: colibri: Run GLM-5.2 on a 25GB-RAM consumer machine github.com
- Z.ai on Hugging Face: GLM-5.2 huggingface.co
- BaseRT: Advancing Best-in-Class LLM Inference with Apple M5 Neural Accelerators arxiv.org
- basecompute on GitHub: baseRT github.com
- RotaryQuant: Fitting 120B MoE Models on Consumer Hardware via Fused Compressed-Space Attention arxiv.org
- Meta: Introducing Muse Glimmer research.meta.ai
- Meta on Hugging Face: Muse Glimmer 30B huggingface.co
- FreeToken: Efficient Edge-Native MoE Serving with Bandwidth-Adaptive Execution arxiv.org