Root cause: bin/filing-plan and references/issue-template.md always render the ia-candidate marker HTML-comment-wrapped. A direct issue_write probe confirmed the Gitea MCP write path preserves a literal <!-- --> unchanged, so the loss on oleks/claude-plugin-cluster#56 and #57 happened on the Agent(anxious:issuer-agent) relay hop: an HTML comment is invisible when its surrounding text is read as markdown, so an agent relaying "what the body should say" instead of pasting it verbatim can silently drop it. Two fixes: - bin/filing-plan: find_existing()'s marker_re() now also matches the bare, unwrapped "ia-candidate: <id>" line (end-of-line anchored, not \b, since candidate ids contain "-" and a word-boundary check would false-match a longer id sharing a prefix) so already-filed issues like cluster#56/#57 are still recognized and not duplicate-filed. - skills/offload-audit/SKILL.md: the issuer delegation prompt now fences the rendered body in its own code block and explicitly warns that the leading marker is an invisible-by-design HTML comment, not stray syntax to clean up, instead of inlining it as bare markdown the relaying agent could read right through. references/issue-template.md documents both, and tests/filing.test.sh adds coverage for the bare-marker match plus a same-prefix-different-id non-match case. Bumps plugin.json to 0.7.1.
inference-arbitrage
Audits a Claude Code plugin — its skill/agent definitions and its real usage transcripts — to find steps that are being done by raw LLM inference but pass every test of a deterministic script, and files the well-evidenced ones as issues on the target plugin's own repo. Runs on demand; each run persists a snapshot, so cost, candidate status, and this plugin's own calibration judgment all accumulate over time instead of resetting on every audit.
The name echoes builder-arbitrage: route each unit of work to the cheapest
executor that can do it correctly — here, "script vs. model" instead of
"which build node."
Use it
/offload-audit <plugin name or path> [--days N]
or ask directly: "audit cluster for token optimization", "is this plugin
wasting inference?". Read-only with respect to whatever it audits — it never
edits the plugin it's looking at and never implements its own recommendations.
What it actually does
- Static pass (
bin/plugin-inventory) — verb ratios, duplicated command blocks, rule tables, script coverage, per plugin definition. - Dynamic pass (
bin/offload-scan) — real cost attribution from this environment's transcripts: Mechanical Turn Ratio, recurring tool-call n-grams, read amplification, retry density, judgment density, fan-out. - Boundary classification — every candidate gets graded against the
rubric below (
design/rubric.md) and gets aboundary_confidence. - Two hard gates, enforced mechanically, never by agent judgment:
- FR-4.2 — no falsifiable "case where a human would overrule the script" → downgraded to a boundary question, never filed.
- FR-4.5 — the evidence-strength label (
measured/thin/unmeasured) or the filing-threshold side must be stable across measurement windows, checked against this plugin's own snapshot history. Ameasured 12%that becomesthin 0%on a different date range is a defect in the auditor, not a fact about the target, and it gets self-reported on this repo, not the target's.
- Filing — a
high-confidence candidate above the 2% audited-spend threshold is filed as a tracked issue on the target's own repo, idempotent via a marker comment so a re-run comments instead of duplicating. Everything below the bar is reported as a boundary question, never filed. - Run summary — a durable page on this plugin's own wiki, plus (FR-9) drawers and knowledge-graph facts in mempalace: one set scoped to the audited plugin's own wing (so a human working on that plugin later finds the findings while searching their own project's memory), one set in this plugin's own wing accumulating cross-target calibration judgment. Recalled memory is advisory input only — it can never override the two gates above.
The rubric, in one table
Every candidate is exactly one of four positions:
| Position | What it means | Example |
|---|---|---|
1 — pure-script |
Deterministic end to end | parsing, diffing, formatting a report |
2 — script-with-compiled-judgment |
Judgment made once at authoring time, frozen into a rule | a hunt's require/exclude rules |
3 — llm-over-script-digest |
Script gathers/filters/ranks; model judges only the shortlist — where most real value is | worktree-audit's provably-safe classification, model handles only the must-reconcile pile |
4 — pure-inference |
Irreducible — say so explicitly | intent, naming, novel synthesis |
The gate to reach position 1 or 3 is five determinism tests (closed input
domain, a stable output oracle, invariance, no semantic gap, bounded branch
factor) — all five, no averaging — plus a falsifiability triple: a typed
function signature, three input→output pairs, and the case where a human
would overrule the script. Missing the third element is the hard gate above.
Full writeup: design/rubric.md.
Worked examples, from real audits
oleks/claude-plugin-cluster#56—cluster:gitea-agent's sequential issue-read batching,llm-over-script-digest, high confidence, measured at 7.26% of audited spend over 887 invocations.oleks/claude-plugin-anxious#29—pr-rot's script-verb-dominant body (verb ratio 1.000, zero backing script), filed on the static case alone per FR-2.4 since it had zero invocations in the audited window (the plugin disables its own model invocation).kotkan/claude-plugin-inference-arbitrage#15— a real catch by the FR-4.5 gate:anxious/agent-wip/release-policy-derivationmeasuredthinat 0.00% of spend on one window andmeasuredat 11.85% on another. Correctness case was complete (a duplicate hand-synced implementation already exists inbin/wip), but the window-dependent value claim meant it was refused rather than filed — self-reported here, not onanxious's repo, because the instability is this plugin's own defect.
Design docs
design/spec.md— scope, functional requirements (the numberedFR-*this README references)design/plan.md— architecture and file layoutdesign/rubric.md— the full script-vs-inference boundary rubricdesign/tasks.md— build order and phase history
Tracked as issues on this repo, milestone v0.1.0.