# A confident answer is not a checked answer

[Journal](/journal/)

August 29, 2026 · [evidence](/journal/#evidence)



One on-device model’s wrong answers looked like its right answers. Repeated checks recovered accuracy at a computation cost.

![An inspection instrument compares two matching outputs while a verdigris loop sends one output through repeated checks.](/_astro/a-confident-answer-is-not-a-checked-answer.DnkA4UiT_Z1pCio5.avif)

A confident answer is not a checked answer. One preprint audit of one on-device model [found that wrong and right outputs looked alike to users](https://arxiv.org/abs/2608.23663).

On-device means the model runs on hardware the user already owns. That placement changes where operators can inspect failures. A reader cannot grade an answer by its confident tone when visible evidence carries no useful separation signal.

[Shashwat Pandey, Satwik Pandey, and Suresh Raghu](https://arxiv.org/abs/2608.23663) are neither Muniment customers nor endorsers.

## The visible answer carried almost no warning

On false-premise questions, the model [gave false answers 69 percent of the time](https://arxiv.org/abs/2608.23663). On entirely benign inputs, it [refused 18 percent of prompts](https://arxiv.org/abs/2608.23663).

AUROC is a separation score where 0.5 equals chance. A classifier built from 15 visible features [scored 0.55 AUROC when separating correct from wrong answers](https://arxiv.org/abs/2608.23663). The model’s own confidence signal [scored 0.47 AUROC](https://arxiv.org/abs/2608.23663).

Those scores close off a tempting shortcut. Operators cannot treat polished wording, length, or the model’s stated confidence as a reliable release check.

## Repeated checks recovered accuracy at a cost

The authors [added a black-box consistency wrapper, which compares repeated answers without access to the model itself](https://arxiv.org/abs/2608.23663). It [raised selective accuracy from 43 percent to 83 percent](https://arxiv.org/abs/2608.23663). Selective accuracy means accuracy across the answers the system chooses to return.

The wrapper [recovered reliability at a tunable cost](https://arxiv.org/abs/2608.23663). Repeated answers spend more computation than one answer. A local model tier therefore needs a verification route with an explicit computation budget.

Muniment’s operating conclusion is narrow. Route low-risk work to an on-device model only when repeated checks fit the budget and latency limit. Otherwise, send the work to a tier with a stronger verification path.

This audit does not prove that local models fail in general. It shows that one audited model produced failures its visible outputs could not reveal. Confidence came free, while verification sent a computation bill.

## Sources

1.  [Confidently Wrong, Silently So: Auditing Undetectable Failures of a Deployed On-Device Language Model](https://arxiv.org/abs/2608.23663) arxiv.org
