Journal

· Amended · evidence

Guarantees belong in the audit harness

An audit harness controls the sources, schema, route, and record around a replaceable model. Every contract it enforced held in a public-data evaluation.

A removable model module sits inside a technical harness linked to a manifest, schema, validation gate, and recorded trace.

An audit harness guarantees four things that a prompt can only request. It decides which sources an answer may cite, which schema that answer must fit, which route the request takes, and which record it leaves. A harness is the surrounding software that supplies tools and rules, records what happened, and checks the result. An audit harness is that software built to hold the guarantees, in versioned code around a replaceable model boundary. When the model breaks one of its rules, deterministic code catches the violation and leaves an artifact. That difference gives operators something durable to review.

A contract pass is not a perfect model output

The authors tested a public-data slice spanning five Korean corporate groups and 25 listed companies. Across 270 composition-boundary runs using three hosted models, every harness-enforced contract held under model substitution.

That result does not mean every model output passed. The paper reports model-composed failures, which validators caught and recorded before the output crossed the contract boundary. That guarantee belongs to the enforcement path: deterministic code rejects or transforms a violation and preserves evidence of what happened. A model-side pass is one possible event inside that path, not the definition of success.

For an operator, the durable boundary is concrete: code defines the checks, manifests bind allowed sources and entities, and schemas define acceptable output. An audit trace records the sources, decisions, and validation results behind an answer. An independent output audit can test the model named in that trace. Any hosted model can then be replaced without asking a new prompt to inherit the old model’s institutional memory.

Prompting asks and the audit harness decides

The ablation holds the model fixed and varies the enforcement layer across 120 runs per condition: 40 scenarios repeated three times. Utility counts runs answered rather than blocked. It is separate from contract compliance, which is why prompt-only output can retain utility while admitting violations. These values belong to the fixed public-data study rather than to production traffic.

The external guardrail shows that blocking is only half the job. Its 88/120 utility score trails the harness’s 120/120 because enforcement detached from application semantics can refuse useful work. Code owned with the application can validate the specific contract and retain the response that satisfies it.

The boundary now exists as shipped code

On its fixed data, the paper establishes the pattern. On August 20, 2026, Temporal released the same outer boundary as product code a team can adopt instead of building every part. A harness is software that runs the model and tools while holding saved state and controls around them.

Temporal says its Agent Harness wraps Gemini, the OpenAI Agents SDK, and PydanticAI instead of replacing their agent loops. Temporal names saved execution state, tool approval policies, strongly typed operations, and structured AgentEvent records. Temporal describes durable execution as saved progress that survives a worker crash or deployment, waits for approval, and resumes without repeating completed work.

Available code and future direction need separate labels. Temporal publishes the code and examples in its project repository, which marks the harness and its integrations as experimental. Separately, Temporal states that the release sits earlier than public preview and that its interfaces will change. Its announcement reports product capabilities, not measured outcomes. Temporal is neither a Muniment customer nor an endorser.

September 1, 2026 amendment: Orkes ships the boundary in Conductor

Orkes presents these agent capabilities as shipping now on Conductor, its existing workflow engine. That status differs from Temporal’s release, which sits earlier than public preview and whose repository marks it experimental.

Orkes says model responses and tool calls become workflow tasks. Orkes says the conversation sits in the engine rather than the agent process. Orkes also says Conductor rejects a tool that the agent never declared before it schedules anything. Orkes says Conductor records the prompt, model response, tool inputs and outputs, and approval with its approver.

Here, durable means that task state remains after a process stops. Orkes says Conductor stores every transition before it schedules the next task. Orkes reports that crashes, deployments, and approval waits leave a run ready to resume from its stopped task.

These are vendor-reported product capabilities, not measured outcomes. Orkes is neither a Muniment customer nor an endorser.

Keep the evidence boundary narrow

This preprint is the authors’ public-data evaluation, not independent proof of production performance or universal generalization. Its fixed slice does not establish that the same contracts, utility, or failure rates will transfer to another workload. Passing the specified contracts does not establish the correctness of the underlying investment claims or the investment quality of the briefings. The study does establish a testable engineering pattern: keep deterministic guarantees and audit artifacts in the versioned audit harness, keep source-backed claims authoritative, and keep the model replaceable.

The paper’s authors, the model providers, and the studied companies are neither Muniment customers nor endorsers.

Sources

  1. arXiv: From Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents arxiv.org
  2. Temporal: Temporal Agent Harness: An early look at durable agent infrastructure temporal.io
  3. GitHub: Temporal Agent Harness github.com
  4. Orkes: Conductor Now Runs Agents Alongside Your Workflows, on the Same Durable Engine orkes.io
  5. Orkes: Agents on Conductor: Architecture for Production AI orkes.io

Continue reading

All publications

Join the waitlist

Get desktop release updates.

We will email you about desktop releases and new features. muniment is a desktop workspace for your models, tools, and files.