dcec9cabfb
Phase 7 filed anxious/agent-wip/release-policy-derivation as `measured` at 11.94% of audited spend on a 30-day window. The same candidate over 7 days came back `thin` at 0.00%. Only the date range differed. A human reading either issue body alone cannot know the other exists, which is rubric.md §4a's failure mode reached through the plugin's own headline number. That was caught by a human running a second skeptical review — luck, not a control. It is now mechanical, in the same register as FR-4.2's overrule-case gate: - bin/stability-classify matches each `file` candidate against the most recent prior snapshot recording it (FR-5.2 identity) and downgrades it to a boundary question if `measurement_strength` changed or the share crossed the 2% filing threshold between windows. Movement within a label is a trend, not an instability; no prior window is `unchecked`, not unstable (FR-2.4). - The downgrade is a finding against THIS repo, not the target's — the target did not change, the auditor described it two ways. Carries an `ia-stability` marker so a re-run comments rather than duplicating. - bin/ia_store.py factors the store and identity resolution out of audit-snapshot so both scripts hash signatures identically. A divergence there would make the gate match nothing and fail open. Documented as spec FR-4.5 and rubric §5b (shipped verbatim in references/). Wired into skills/offload-audit as step 4b, before filing-plan; steps 5-6 now consume its output rather than classified.json. S4 is retired in its old form — a human triage catching a bad candidate is not a repeatable test — and re-passed as: the gate catches the real release-policy-derivation case with no human in the loop. Asserted in tests/stability.test.sh case (f) against the verbatim two-window Phase 7 output. Closes kotkan/claude-plugin-inference-arbitrage#11 Suite: 178 assertions, exit 0.
26 lines
1.4 KiB
Markdown
26 lines
1.4 KiB
Markdown
---
|
|
description: Audit a Claude Code plugin for steps that should be a deterministic script instead of raw inference — measure, classify, file the well-evidenced candidates, and write the run summary. Usage — /offload-audit <plugin|path> [--days N]
|
|
argument-hint: <plugin name or path> [--days N]
|
|
allowed-tools: ["Bash", "Read", "Grep", "Glob", "Skill", "Agent"]
|
|
---
|
|
|
|
# /offload-audit
|
|
|
|
Run the full audit on **$ARGUMENTS**.
|
|
|
|
1. Parse `$ARGUMENTS` into a target (a plugin name or a directory path) and an
|
|
optional `--days N` window, defaulting to 30.
|
|
2. If no target was given, ask which plugin to audit. Do not guess one, and do
|
|
not default to the current directory — auditing the wrong plugin wastes a
|
|
full transcript scan.
|
|
3. Invoke the `offload-audit` skill and follow it exactly, including both filing
|
|
gates (the overrule case, FR-4.2; the evidence-stability check against the
|
|
previous window, FR-4.5) and the marker search. For a heavy target, or when
|
|
the report should not consume this session's context, hand the work to
|
|
`Agent(subagent_type="inference-arbitrage:offload-analyst")` instead — one
|
|
agent for the whole audit, never one per skill (spec §7).
|
|
|
|
Filing is implicit: a candidate that clears the rubric gate and the value
|
|
threshold is filed without asking. Everything below the gate is reported as a
|
|
boundary question and filed nowhere.
|