# Audit the backend from the text alone

[Journal](/journal/)

August 6, 2026 · [evidence](/journal/#evidence)



IRIS tests which model a gateway served from returned text alone, turning the model name on a receipt into a claim operators can check.

![A routing manifold feeds a paper trace into an inspection chamber that compares it with three backend fingerprint cylinders.](/_astro/audit-the-backend-from-the-text-alone.SdUk8SQx_Z12gBhj.avif)

A receipt that names a model states a checkable claim. Returned text can carry enough evidence to test which backend answered, without gateway logs.

The backend is the model that actually answered. Model substitution means a gateway serves a different model from the one it advertised. Partial rerouting means only a fraction of requests goes to the substitute.

Yuewei Zhang, Zhi-Hai Zhang, and Hanzhang Qin call [their audit IRIS](https://arxiv.org/abs/2607.20860). It [asks models for random strings, then fingerprints their visible patterns](https://arxiv.org/abs/2607.20860).

## Four results make the claim measurable

AUROC scores how well an audit separates honest responses from substituted ones. A score of [1 means perfect separation, while 0.5 means chance performance](https://arxiv.org/abs/2607.20860).

On the authors’ six-model, intra-family Qwen3 ladder, [IRIS verified the backend at 0.99 AUROC](https://arxiv.org/abs/2607.20860). That result measures claimed and substitute pairs inside one model family.

Their commercial sample contained 17 OpenRouter model interfaces. On margin-qualified pairs, [IRIS caught a 0.3 routing fraction at 0.85 mean power, with a 0.017 false-positive rate](https://arxiv.org/abs/2607.20860).

For enrolled substitutes in that library, [the audit recovered the routing fraction to within 0.04](https://arxiv.org/abs/2607.20860). Enrolled means the candidate library already contained the substitute’s fingerprint.

The live sample compared 15 pairs that offered the same model through different providers. IRIS [flagged 14 pairs for backend differences tied to quantization and inference kernels](https://arxiv.org/abs/2607.20860).

On claimed and substitute pairs from the six-model ladder, [adaptive allocation raised the matched-budget target-hit rate from 73% to 87%](https://arxiv.org/abs/2607.20860). The comparison held total query spending equal while assigning queries by measured difficulty.

Those samples matter as much as the scores. The authors measured their audit on their model sets, including commercial interfaces, rather than proving a claim about every gateway.

## The query budget is a commitment, not a cost guarantee

IRIS runs [a cheap pilot on labeled reference responses](https://arxiv.org/abs/2607.20860). It fits error decay, then [fixes the budget before querying a suspect endpoint](https://arxiv.org/abs/2607.20860).

That order prevents a suspicious result from quietly buying more evidence until it passes. It remains the authors’ reported design choice, not a guarantee that an audit will cost little.

Low-margin pairs can remain hard to separate. [The paper returns some infeasible pairs as indeterminate](https://arxiv.org/abs/2607.20860), which keeps a weak fingerprint from becoming a confident accusation.

## A named route needs an independent check

Muniment issues a per-request record that names the route and model. That is our own first-party claim about the product, and the IRIS paper does not validate it.

The useful connection is narrower. An [audit harness preserves what the system claimed](/journal/audit-guarantees-belong-in-the-harness/), while an output audit tests whether returned text matches that claim.

Operators need both halves when the gateway sits outside their control. The receipt fixes the claim in time, and the audit supplies independent evidence about the backend.

No paper author, gateway, or provider is a Muniment customer or endorser.

[Join the waitlist](/waitlist/?utm_campaign=audit-the-backend-from-the-text-alone) if per-request route records belong in your operating layer. The next uncomfortable question is how often a production fleet should challenge its own receipts.

## Sources

1.  [Which Model Is Actually Serving You? IRIS: Budgeted Black-Box Auditing of Model Substitution and Routing Dilution in LLM Gateways](https://arxiv.org/abs/2607.20860) arxiv.org
