Each governance layer needs its own test
A control that governs one agent may fail across a group. Each governance layer needs tests and records at its own scale.

A control that governs one agent may fail when agents work as a group. Each governance layer needs its own tests and records.
An agent is software that completes steps toward a goal. Srinivas Telukunta, Georgios Nektarios Lilis, and Lucio Baron propose four governance layers: individual agents, agent collectives, human-agent teams, and agent fleets. A fleet is the set of agents an organization runs.
Their CASE framework is a proposal in a new paper. Its layer model and reported studies do not prove that the framework improves production outcomes.
Individual agents need bounded tests
Test an individual agent against its declared goal, allowed tools, budgets, and intervention path. Keep the inputs, policy version, observed actions, and intervention result.
That record answers a narrow question: did this agent stay inside its operating bounds during this test? It says little about shared state or coordinated action.
Agent collectives need interaction tests
Agent collectives introduce connections that no isolated test exercises. A shared memory entry, tool result, or delegated task can carry one agent’s error into another agent’s work.
Test the full interaction graph under staged failures and poisoned shared state. Record the participating agents, declared connections, shared-state changes, cascade path, breaker response, and final outcome.
A passing isolated test cannot substitute for that record. The collective test asks whether interaction changes behavior or spreads a failure.
Human-agent teams need oversight drills
Human-agent teams add a capacity problem. A person can hold formal override authority and still receive too many cases or too little context to act.
Run drills at expected peak volume. Record the escalation threshold, evidence packet, responsible role, response time, decision, override result, and unresolved work.
An approval record proves that someone approved one state. It does not prove that the team can detect and stop a later change.
Agent fleets need operating drills
Agent fleets need tests across the organization’s release and recovery path. Stage a failed rollout, stale policy, broken dependency, and partial rollback.
Record the fleet inventory, release artifact, policy version, affected scope, error budget, rollback result, and work that remained incomplete. A zero-incident interval still needs the inventory and test record.
Fleet health cannot come from averaging away a failed layer. Release evidence should show a current pass for each layer and preserve every partial failure.
The reported counts describe separate studies
The authors report a corpus of 62 coded incidents. They say 51 of those incidents, or 82 percent, carried at least one secondary layer.
The paper separately reports a review of 22 ecosystem tools and 35 scored public deployments. Those counts cover different studies. They do not form one incident total.
These are the authors’ results from coded public evidence. They are not controlled production tests, an industry census, or proof of effectiveness.
The authors and their institutions are neither Muniment customers nor endorsers.
The useful release question is no longer whether the agent passed. Ask which layer passed, against which state, with which record.