Extract workflow heredoc parsers into scripts
Our iOS failure-digest parsers became testable and reviewable after they moved from workflow heredocs into standalone scripts.

Code no test can load is code no review can trust, even when it sits in plain sight. Our iOS failure-digest parsers proved the point.
About 300 lines of Python and awk lived as heredocs inside one workflow file. Running that workflow was the only way to exercise them.
When our iOS runner went offline for days, we could not even check the parsers by hand.
Two platforms exposed the testing gap
Our mobile nightly turns failed device journeys into short reports. It parses the captured view hierarchy, driver command log, and app system log.
Android already ran its digest from a standalone shell script with a self-test harness. A guard against truncated captures shipped beside a regression scenario.
iOS had a twin guard, but its parsing logic remained inside workflow YAML. That guard shipped without a test because no harness could load it.
This asymmetry made the cost concrete. Equivalent safeguards carried different proof because only one parser had crossed the workflow boundary.
Extraction turned workflow text into testable code
The fix was mechanical. We moved each heredoc into a version-controlled script with an argv interface and an exit code of 0 or 1.
Each self-test creates small fixture captures in a temporary directory and compares the exact output. Pull requests now run those tests without any device or runner.
The workflow step became one script call with a fallback. Its responsibility stayed orchestration, while the parser gained an interface that a test could call directly.
Review found what the workflow had hidden
Extraction made the parsers reviewable. Review then found two defects in the old inline logic.
A node counter under-reported dropped evidence after the parser reached its print cap. One malformed input record made a loop discard an entire computed digest.
Both defects had sat in visible workflow text. Visibility did not make that text reviewable because reviewers lacked a focused test and an exact expected result.
A small script changed that condition. The parser could fail against a fixture before a workflow or device entered the picture.
Fixtures cannot detect capture-format drift
This test boundary has a limit. A fixture proves behavior against its recorded shape, while a live device tool can change that shape between versions.
Only a real run can catch such upstream drift. Our self-tests bound regressions in the parsing logic, not changes in the live capture format.
That narrower guarantee still matters. Parser changes no longer depend on an online iOS runner before anyone can test or review them.
This evidence comes from our Muniment Mobile note pinned to commit b179b22baabadcf3e1f0f3e02d703a0bc6fa433e.