From ba445cf4c9f33460f4fef81591c40f237f00c515 Mon Sep 17 00:00:00 2001 From: Oleks Date: Thu, 30 Jul 2026 04:43:02 +0300 Subject: [PATCH] =?UTF-8?q?Phase=208:=20publish=20=E2=80=94=20README=20rew?= =?UTF-8?q?rite,=20local-only=20scope?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit README replaced (was still 'scaffolding, pre-implementation' at v0.5.1 with two real audits already filed): usage, the six pipeline steps, the rubric's four positions in one table, and three worked examples from real runs (oleks/claude-plugin-cluster#56, oleks/claude-plugin-anxious#29, kotkan/claude-plugin-inference-arbitrage#15). Publish scoped to local use per explicit user decision: already registered in oleks-local (1.2c, since Phase 1); deliberately not added to the public claude-plugin-index or verified on claude-plugins.oleks.space. tasks.md 8.2 records that as a decision, not an oversight, so a later session doesn't 'fix' it by publishing publicly without asking. 8.3 (optional cc-tokens upstream PR) stays open, not required for this release. 8.4 (memory checkpoint) already satisfied — FR-9 writes on every run plus this session's manual checkpoints. --- .claude-plugin/plugin.json | 2 +- README.md | 94 ++++++++++++++++++++++++++++++++++---- design/tasks.md | 12 +++-- 3 files changed, 94 insertions(+), 14 deletions(-) diff --git a/.claude-plugin/plugin.json b/.claude-plugin/plugin.json index 830975c..59c80c1 100644 --- a/.claude-plugin/plugin.json +++ b/.claude-plugin/plugin.json @@ -1,6 +1,6 @@ { "name": "inference-arbitrage", - "version": "0.5.1", + "version": "0.6.0", "description": "Audits a Claude Code plugin's definitions and usage transcripts to find steps that should be a deterministic script instead of raw LLM inference, and files the well-evidenced ones as issues on the target repo.", "author": { "name": "oleks" diff --git a/README.md b/README.md index ce7da69..d37a8c2 100644 --- a/README.md +++ b/README.md @@ -3,19 +3,97 @@ Audits a Claude Code plugin — its skill/agent definitions and its real usage transcripts — to find steps that are being done by raw LLM inference but pass every test of a deterministic script, and files the well-evidenced ones as -issues on the target plugin's own repo. Runs on demand; each run accumulates -into a snapshot history so cost and candidate status can be tracked over time. +issues on the target plugin's own repo. Runs on demand; each run persists a +snapshot, so cost, candidate status, and this plugin's own calibration +judgment all accumulate over time instead of resetting on every audit. The name echoes `builder-arbitrage`: route each unit of work to the cheapest executor that can do it correctly — here, "script vs. model" instead of "which build node." -Status: **scaffolding, pre-implementation.** See `design/` for the full -specification: +## Use it -- [`design/spec.md`](design/spec.md) — why, scope, functional requirements +```text +/offload-audit [--days N] +``` + +or ask directly: *"audit `cluster` for token optimization"*, *"is this plugin +wasting inference?"*. Read-only with respect to whatever it audits — it never +edits the plugin it's looking at and never implements its own recommendations. + +## What it actually does + +1. **Static pass** (`bin/plugin-inventory`) — verb ratios, duplicated command + blocks, rule tables, script coverage, per plugin definition. +2. **Dynamic pass** (`bin/offload-scan`) — real cost attribution from this + environment's transcripts: Mechanical Turn Ratio, recurring tool-call + n-grams, read amplification, retry density, judgment density, fan-out. +3. **Boundary classification** — every candidate gets graded against the + rubric below (`design/rubric.md`) and gets a `boundary_confidence`. +4. **Two hard gates**, enforced mechanically, never by agent judgment: + - **FR-4.2** — no falsifiable "case where a human would overrule the + script" → downgraded to a boundary question, never filed. + - **FR-4.5** — the evidence-strength label (`measured` / `thin` / + `unmeasured`) or the filing-threshold side must be *stable* across + measurement windows, checked against this plugin's own snapshot + history. A `measured 12%` that becomes `thin 0%` on a different date + range is a defect in the auditor, not a fact about the target, and it + gets self-reported on *this* repo, not the target's. +5. **Filing** — a `high`-confidence candidate above the 2% audited-spend + threshold is filed as a tracked issue on the **target's own repo**, + idempotent via a marker comment so a re-run comments instead of + duplicating. Everything below the bar is reported as a boundary question, + never filed. +6. **Run summary** — a durable page on this plugin's own wiki, plus (FR-9) + drawers and knowledge-graph facts in mempalace: one set scoped to the + audited plugin's own wing (so a human working on that plugin later finds + the findings while searching their own project's memory), one set in this + plugin's own wing accumulating cross-target calibration judgment. Recalled + memory is advisory input only — it can never override the two gates above. + +## The rubric, in one table + +Every candidate is exactly one of four positions: + +| Position | What it means | Example | +| --- | --- | --- | +| 1 — `pure-script` | Deterministic end to end | parsing, diffing, formatting a report | +| 2 — `script-with-compiled-judgment` | Judgment made once at authoring time, frozen into a rule | a hunt's require/exclude rules | +| 3 — `llm-over-script-digest` | Script gathers/filters/ranks; model judges only the shortlist — **where most real value is** | `worktree-audit`'s provably-safe classification, model handles only the must-reconcile pile | +| 4 — `pure-inference` | Irreducible — say so explicitly | intent, naming, novel synthesis | + +The gate to reach position 1 or 3 is five determinism tests (closed input +domain, a stable output oracle, invariance, no semantic gap, bounded branch +factor) — all five, no averaging — plus a **falsifiability triple**: a typed +function signature, three input→output pairs, and the case where a human +would overrule the script. Missing the third element is the hard gate above. +Full writeup: [`design/rubric.md`](design/rubric.md). + +## Worked examples, from real audits + +- **`oleks/claude-plugin-cluster#56`** — `cluster:gitea-agent`'s sequential + issue-read batching, `llm-over-script-digest`, high confidence, measured at + 7.26% of audited spend over 887 invocations. +- **`oleks/claude-plugin-anxious#29`** — `pr-rot`'s script-verb-dominant body + (verb ratio 1.000, zero backing script), filed on the static case alone per + FR-2.4 since it had zero invocations in the audited window (the plugin + disables its own model invocation). +- **`kotkan/claude-plugin-inference-arbitrage#15`** — a real catch by the + FR-4.5 gate: `anxious/agent-wip/release-policy-derivation` measured + `thin` at 0.00% of spend on one window and `measured` at 11.85% on + another. Correctness case was complete (a duplicate hand-synced + implementation already exists in `bin/wip`), but the window-dependent + value claim meant it was refused rather than filed — self-reported here, + not on `anxious`'s repo, because the instability is this plugin's own + defect. + +## Design docs + +- [`design/spec.md`](design/spec.md) — scope, functional requirements (the + numbered `FR-*` this README references) - [`design/plan.md`](design/plan.md) — architecture and file layout -- [`design/rubric.md`](design/rubric.md) — the script-vs-inference boundary rubric -- [`design/tasks.md`](design/tasks.md) — build order +- [`design/rubric.md`](design/rubric.md) — the full script-vs-inference + boundary rubric +- [`design/tasks.md`](design/tasks.md) — build order and phase history -Tracked as issues on this repo under the `v0.1.0` milestone. +Tracked as issues on this repo, milestone `v0.1.0`. diff --git a/design/tasks.md b/design/tasks.md index 23e5c95..c90ec73 100644 --- a/design/tasks.md +++ b/design/tasks.md @@ -190,13 +190,15 @@ that ships without it would publish an audit loop that forgets every run. ## Phase 8 — Publish -- [ ] **8.1** README with the thesis, the rubric summary, and worked examples. -- [ ] **8.2** Publish per `plugin-publishing`: push, index entry in - `~/projects/claude-plugin-index/plugins.json`, verify - `claude-plugins.oleks.space/api/plugins`, `claude plugin add`. +- [x] **8.1** README with the thesis, the rubric summary, and worked examples. +- [x] **8.2** Publish per `plugin-publishing`, scoped to local use: pushed to + `main`, already registered in `oleks-local` (1.2c). **Deliberately not** + added to `~/projects/claude-plugin-index/plugins.json` or verified on + `claude-plugins.oleks.space` — user's explicit call, since this plugin is + used only through `oleks-local`. Revisit if that changes. - [ ] **8.3** Optional, from decision 0.1: PR `cc-tokens attribute --by skill|plugin` to `token-budget`. -- [ ] **8.4** Memory checkpoint — the rubric and the calibration findings into +- [x] **8.4** Memory checkpoint — the rubric and the calibration findings into `claude-plugins`, room `inference-arbitrage-lessons`. Written against `wing_claude_memory` / `working-practice` before FR-9 existed; that predates the wing/room split FR-9.3 actually shipped with, and having two