# An average score hides the governance check

[Journal](/journal/)

August 11, 2026 · [evidence](/journal/#evidence)



Capability cannot certify release readiness. A release test must keep governance evidence visible beside behavior, context, and compliance.

![Four separate instrument gauges feed a release rail with visible checkpoints before a final mechanical gate.](/_astro/an-average-score-hides-the-governance-check.BuZomIZT_117IQB.avif)

A capability score cannot certify release readiness. Capability describes what an agent can do. A release test must keep governance evidence separate and visible.

Fouad Bousetouane [proposes the ProofAgent Index as a governance readiness index](https://arxiv.org/abs/2607.27677). It puts four checks beside one another:

1.  Evaluation measures [observed behavior](https://arxiv.org/abs/2607.27677).
2.  Context measures [the operating environment that shapes that behavior](https://arxiv.org/abs/2607.27677).
3.  Compliance measures [adherence to applicable rules and controls](https://arxiv.org/abs/2607.27677).
4.  Governance measures whether an organization can [authorize, monitor, audit, and control the agent during operation](https://arxiv.org/abs/2607.27677).

These checks answer different release questions. Averaging them discards the distinction that makes the governance check useful.

## Governance needs its own receipt

The abstract says [capability can improve behavior without determining readiness](https://arxiv.org/abs/2607.27677). Authority, monitoring, audit, and control remain separate questions.

That separation matters before release and during operation. [An activity log cannot replace a release control](/journal/an-activity-log-is-not-a-release-control/) because a later record cannot govern an earlier decision.

Oversight also depends on records that survive the transfer of authority. [Agent oversight loosens](/journal/agent-oversight-loosens-only-when-the-records-hold/) only when the records hold.

The proposal’s sharpest requirement is that [governance evidence must remain visible rather than average away](https://arxiv.org/abs/2607.27677). A composite score can hide a weak governance check behind stronger behavior.

## A proposal needs a boundary

The abstract [reports validation in healthcare and finance](https://arxiv.org/abs/2607.27677). It [says the index separates higher-risk from lower-risk configurations](https://arxiv.org/abs/2607.27677).

This evidence supports [one author’s proposal](https://arxiv.org/abs/2607.27677). It does not establish an accepted standard. [No production adoption appears](https://arxiv.org/abs/2607.27677) in the abstract.

No paper author is a Muniment customer or endorser.

Muniment’s position is that a release record should show each check. One average cannot stand in for visible governance evidence.

[Join the waitlist](/waitlist/?utm_campaign=an-average-score-hides-the-governance-check) if your release test needs separate receipts for capability and governance.

## Sources

1.  [Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness](https://arxiv.org/abs/2607.27677) arxiv.org
