Journal

· evidence

An average score hides the governance check

Capability cannot certify release readiness. A release test must keep governance evidence visible beside behavior, context, and compliance.

Four separate instrument gauges feed a release rail with visible checkpoints before a final mechanical gate.

A capability score cannot certify release readiness. Capability describes what an agent can do. A release test must keep governance evidence separate and visible.

Fouad Bousetouane proposes the ProofAgent Index as a governance readiness index. It puts four checks beside one another:

  1. Evaluation measures observed behavior.
  2. Context measures the operating environment that shapes that behavior.
  3. Compliance measures adherence to applicable rules and controls.
  4. Governance measures whether an organization can authorize, monitor, audit, and control the agent during operation.

These checks answer different release questions. Averaging them discards the distinction that makes the governance check useful.

Governance needs its own receipt

The abstract says capability can improve behavior without determining readiness. Authority, monitoring, audit, and control remain separate questions.

That separation matters before release and during operation. An activity log cannot replace a release control because a later record cannot govern an earlier decision.

Oversight also depends on records that survive the transfer of authority. Agent oversight loosens only when the records hold.

The proposal’s sharpest requirement is that governance evidence must remain visible rather than average away. A composite score can hide a weak governance check behind stronger behavior.

A proposal needs a boundary

The abstract reports validation in healthcare and finance. It says the index separates higher-risk from lower-risk configurations.

This evidence supports one author’s proposal. It does not establish an accepted standard. No production adoption appears in the abstract.

No paper author is a Muniment customer or endorser.

Muniment’s position is that a release record should show each check. One average cannot stand in for visible governance evidence.

Join the waitlist if your release test needs separate receipts for capability and governance.

Sources

  1. Stop Shipping AI Agents on Faith: Capability Is Not Production Readiness arxiv.org

Continue reading

All publications

Join the waitlist

Get desktop release updates.

We will email you about desktop releases and new features. muniment is a desktop workspace for your models, tools, and files.