105e9f5bae45f03ea123e37c37106e2621d1305f
S1-S3 and S5-S8 pass; S4 is prepared but deliberately unfiled pending user sign-off (tasks.md 7.4 stays unchecked). Two bugs surfaced by running the composed pipeline against real data rather than synthetic fixtures, both in load-bearing places: audit-snapshot: cost_per_invocation folded to 0 when a candidate had zero invocations in a window, so a step that simply did not run reported -100% and direction() called it `shrunk`. That is the quiet-window artefact FR-5.4 exists to prevent, one level up from absolute tokens, and it contradicted direction()'s own docstring. It is now None, metric_delta marks the pair `unmeasured`, direction returns a new `unmeasured` status, and the pages render "not run". tests/calibration/mask-worktree-audit.py: recomputed script_loc from scripts alone, silently dropping all 1051 LOC of worktree-discipline's hooks. Masking one 272-LOC script appeared to remove 1323, taking script_coverage 0.353 -> 0.083 instead of 0.297, which let S3 pass partly on an artefact of the mask. The mask now matches build_aggregate. S3 still flags the known cut `high` / llm-over-script-digest on the corrected, harder input. Both fixes carry regression tests; suite is 156 assertions, exit 0.
inference-arbitrage
Audits a Claude Code plugin — its skill/agent definitions and its real usage transcripts — to find steps that are being done by raw LLM inference but pass every test of a deterministic script, and files the well-evidenced ones as issues on the target plugin's own repo. Runs on demand; each run accumulates into a snapshot history so cost and candidate status can be tracked over time.
The name echoes builder-arbitrage: route each unit of work to the cheapest
executor that can do it correctly — here, "script vs. model" instead of
"which build node."
Status: scaffolding, pre-implementation. See design/ for the full
specification:
design/spec.md— why, scope, functional requirementsdesign/plan.md— architecture and file layoutdesign/rubric.md— the script-vs-inference boundary rubricdesign/tasks.md— build order
Tracked as issues on this repo under the v0.1.0 milestone.
Languages
Python
66%
Shell
34%