kotkan/claude-plugin-inference-arbitrage#10 — a hook file is only counted if it's referenced by .claude-plugin/hooks.json, executable, or has a shebang. hooks/hooks.json itself is always excluded regardless, since it is configuration rather than logic. mask-worktree-audit.py needed no change: it derives script_loc from the already-filtered hooks[] output. Bumps plugin.json to 0.6.2 (fix).
inference-arbitrage
Audits a Claude Code plugin — its skill/agent definitions and its real usage transcripts — to find steps that are being done by raw LLM inference but pass every test of a deterministic script, and files the well-evidenced ones as issues on the target plugin's own repo. Runs on demand; each run persists a snapshot, so cost, candidate status, and this plugin's own calibration judgment all accumulate over time instead of resetting on every audit.
The name echoes builder-arbitrage: route each unit of work to the cheapest
executor that can do it correctly — here, "script vs. model" instead of
"which build node."
Use it
/offload-audit <plugin name or path> [--days N]
or ask directly: "audit cluster for token optimization", "is this plugin
wasting inference?". Read-only with respect to whatever it audits — it never
edits the plugin it's looking at and never implements its own recommendations.
What it actually does
- Static pass (
bin/plugin-inventory) — verb ratios, duplicated command blocks, rule tables, script coverage, per plugin definition. - Dynamic pass (
bin/offload-scan) — real cost attribution from this environment's transcripts: Mechanical Turn Ratio, recurring tool-call n-grams, read amplification, retry density, judgment density, fan-out. - Boundary classification — every candidate gets graded against the
rubric below (
design/rubric.md) and gets aboundary_confidence. - Two hard gates, enforced mechanically, never by agent judgment:
- FR-4.2 — no falsifiable "case where a human would overrule the script" → downgraded to a boundary question, never filed.
- FR-4.5 — the evidence-strength label (
measured/thin/unmeasured) or the filing-threshold side must be stable across measurement windows, checked against this plugin's own snapshot history. Ameasured 12%that becomesthin 0%on a different date range is a defect in the auditor, not a fact about the target, and it gets self-reported on this repo, not the target's.
- Filing — a
high-confidence candidate above the 2% audited-spend threshold is filed as a tracked issue on the target's own repo, idempotent via a marker comment so a re-run comments instead of duplicating. Everything below the bar is reported as a boundary question, never filed. - Run summary — a durable page on this plugin's own wiki, plus (FR-9) drawers and knowledge-graph facts in mempalace: one set scoped to the audited plugin's own wing (so a human working on that plugin later finds the findings while searching their own project's memory), one set in this plugin's own wing accumulating cross-target calibration judgment. Recalled memory is advisory input only — it can never override the two gates above.
The rubric, in one table
Every candidate is exactly one of four positions:
| Position | What it means | Example |
|---|---|---|
1 — pure-script |
Deterministic end to end | parsing, diffing, formatting a report |
2 — script-with-compiled-judgment |
Judgment made once at authoring time, frozen into a rule | a hunt's require/exclude rules |
3 — llm-over-script-digest |
Script gathers/filters/ranks; model judges only the shortlist — where most real value is | worktree-audit's provably-safe classification, model handles only the must-reconcile pile |
4 — pure-inference |
Irreducible — say so explicitly | intent, naming, novel synthesis |
The gate to reach position 1 or 3 is five determinism tests (closed input
domain, a stable output oracle, invariance, no semantic gap, bounded branch
factor) — all five, no averaging — plus a falsifiability triple: a typed
function signature, three input→output pairs, and the case where a human
would overrule the script. Missing the third element is the hard gate above.
Full writeup: design/rubric.md.
Worked examples, from real audits
oleks/claude-plugin-cluster#56—cluster:gitea-agent's sequential issue-read batching,llm-over-script-digest, high confidence, measured at 7.26% of audited spend over 887 invocations.oleks/claude-plugin-anxious#29—pr-rot's script-verb-dominant body (verb ratio 1.000, zero backing script), filed on the static case alone per FR-2.4 since it had zero invocations in the audited window (the plugin disables its own model invocation).kotkan/claude-plugin-inference-arbitrage#15— a real catch by the FR-4.5 gate:anxious/agent-wip/release-policy-derivationmeasuredthinat 0.00% of spend on one window andmeasuredat 11.85% on another. Correctness case was complete (a duplicate hand-synced implementation already exists inbin/wip), but the window-dependent value claim meant it was refused rather than filed — self-reported here, not onanxious's repo, because the instability is this plugin's own defect.
Design docs
design/spec.md— scope, functional requirements (the numberedFR-*this README references)design/plan.md— architecture and file layoutdesign/rubric.md— the full script-vs-inference boundary rubricdesign/tasks.md— build order and phase history
Tracked as issues on this repo, milestone v0.1.0.