Journal

· evidence

The driver states the limit beside the command

Anthropic’s hardware preview puts device limits beside commands. Operators still need the script, result, and review in one record.

A command line meets a mechanical stop inside a device driver before reaching a laboratory laser.

A written instruction asks. A driver-enforced limit refuses. Anthropic’s hardware preview puts both at the device interface, where an operator can inspect the difference.

A driver is software that translates between a computer and a device. A command requests an action. A limit blocks an action outside an allowed range.

Put the refusal beside the device

Anthropic reports that its Model Hardware Standard, or MHS, uses a common driver for programmable laboratory and manufacturing equipment.

Each driver produces a reference file that states what the device measures, what callers may adjust, and what safety limits the driver enforces.

That placement matters. A prompt can describe a boundary. Software at the device interface can reject a command that crosses it.

Muniment’s analysis is that instructions and enforced limits belong in separate records. A polite request cannot prove that the device would refuse an unsafe value.

Anthropic also reports orchestration across several devices. It says an agent can fix a learned sequence into a script that runs without more reasoning.

Fixing the sequence narrows one source of variation, but it does not record the limits applied during execution.

The measurements need their test stage

Anthropic reports the following QuEra laser recovery measurements. The six-second development result and the final success rate came from separate test stages.

Measure Before After
Recovery success 58% 99.3%
Recovery time About 150 seconds About 6 seconds during development

Anthropic says the baseline script succeeded about 58 percent of the time and took around 150 seconds.

Anthropic reports about six seconds and 96 percent success after overnight development. It later reports 695 recoveries across 700 blind trials, or 99.3 percent.

Anthropic reports Genentech’s system chose about 140 microliters per second for water.

Anthropic reports 10 microliters per second for viscous BSA samples.

Anthropic reports that Carnegie Mellon ran serial dilution experiments about three times faster.

These figures cover different devices, tasks, and test stages. They do not establish one general performance rate for the standard.

A preview still needs an operator

Anthropic describes MHS as a research preview with early projects and proofs of concept, not production adoption.

It names Genentech, the University of Washington Baker and Pinglay labs, Carnegie Mellon University, HHMI Janelia Research Campus, and QuEra Computing as launch partners.

Anthropic says models still lack physical intuition and require expert oversight. It reports that Genentech researchers identified foaming as a physical failure.

Anthropic says it plans to publish the standard under a public license after more safety tests.

An operator needs the commands sent, the limits in force, the fixed scripts, the results, and the human review. Without that set, refusal remains an interface claim. We argue the same point for model outputs in Guarantees belong in the audit harness.

Anthropic and the named launch partners are neither Muniment customers nor endorsers.

A command records what the caller wanted. The driver must record what the device refused.

Sources

  1. Anthropic: Previewing the Model Hardware Standard www.anthropic.com

Continue reading

All publications

Join the waitlist

Get desktop release updates.

We will email you about desktop releases and new features. muniment is a desktop workspace for your models, tools, and files.