Scaffold plugin: manifest, license, design docs

This commit is contained in:
Oleks
2026-07-29 15:32:34 +03:00
commit d34118ec9a
8 changed files with 1582 additions and 0 deletions
+14
View File
@@ -0,0 +1,14 @@
{
"name": "inference-arbitrage",
"version": "0.1.0",
"description": "Audits a Claude Code plugin's definitions and usage transcripts to find steps that should be a deterministic script instead of raw LLM inference, and files the well-evidenced ones as issues on the target repo.",
"author": {
"name": "oleks"
},
"keywords": [
"token-optimization",
"plugin-audit",
"script-offload",
"efficiency"
]
}
+3
View File
@@ -0,0 +1,3 @@
.claude/
.cache/
*.log
+21
View File
@@ -0,0 +1,21 @@
MIT License
Copyright (c) 2026 Oleks Kuksenko
Permission is hereby granted, free of charge, to any person obtaining a copy
of this software and associated documentation files (the "Software"), to deal
in the Software without restriction, including without limitation the rights
to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
copies of the Software, and to permit persons to whom the Software is
furnished to do so, subject to the following conditions:
The above copyright notice and this permission notice shall be included in all
copies or substantial portions of the Software.
THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
SOFTWARE.
+21
View File
@@ -0,0 +1,21 @@
# inference-arbitrage
Audits a Claude Code plugin — its skill/agent definitions and its real usage
transcripts — to find steps that are being done by raw LLM inference but pass
every test of a deterministic script, and files the well-evidenced ones as
issues on the target plugin's own repo. Runs on demand; each run accumulates
into a snapshot history so cost and candidate status can be tracked over time.
The name echoes `builder-arbitrage`: route each unit of work to the cheapest
executor that can do it correctly — here, "script vs. model" instead of
"which build node."
Status: **scaffolding, pre-implementation.** See `design/` for the full
specification:
- [`design/spec.md`](design/spec.md) — why, scope, functional requirements
- [`design/plan.md`](design/plan.md) — architecture and file layout
- [`design/rubric.md`](design/rubric.md) — the script-vs-inference boundary rubric
- [`design/tasks.md`](design/tasks.md) — build order
Tracked as issues on this repo under the `v0.1.0` milestone.
+476
View File
@@ -0,0 +1,476 @@
# plan.md — `inference-arbitrage` architecture
Implements `spec.md`, classifying by `rubric.md`. Structure follows the
`plugin-publishing` skill's layout and the `worktree-discipline` / `anxious`
precedent (skills + agent + `bin/` scripts + `references/` + `tests/` + a
Woodpecker pipeline).
---
## 1. Naming
**`inference-arbitrage`.**
The name is deliberate and echoes `builder-arbitrage`, which already exists in
this environment and means: *route each unit of work to the cheapest executor
capable of doing it correctly.* That is exactly this plugin's thesis, with the
executor choice being "script vs. model" instead of "which build node." Reusing
the vocabulary means the concept arrives already understood.
Rejected alternatives: `script-offload` (describes the mechanism, not the
judgment, and the judgment is the hard part), `token-optimizer` (collides
conceptually with `token-budget` and implies it optimizes tokens directly, which
it does not — it optimizes *where work executes*). The placeholder directory name
`token-offload-audit` remains available if the user prefers it; nothing in the
design depends on the choice.
### 1.1 Workspace conventions this must satisfy
From `~/projects/claude-plugins/CLAUDE.md` — these are mechanical rules, not
preferences:
- **Directory name must equal the manifest `name`.** True for every plugin in the
workspace. If `inference-arbitrage` is accepted, the checkout is renamed from
`token-offload-audit` at scaffold time (TASKS 1.1).
- **Gitea repo name is exactly `claude-plugin-<manifest-name>`** — no exceptions,
including stuttering cases.
- **The org is NOT derivable from the name.** Both `oleks` and `kotkan` hold
plugins. `kotkan` already holds `token-budget`, `worktree-discipline`, and
`anti-patterns` — the three plugins this one is most closely related to — which
is the argument for putting it there. Decision 0.5 in `tasks.md`.
- **Remote protocol is mixed** (SSH vs HTTPS) and signals nothing; copy whatever
the sibling repos use rather than normalizing.
- **Branch `main`; `.gitignore` = `.cache/`, `.claude/`, `*.log`.**
- **`bin/` scripts must keep the executable bit** (`git ls-files -s``100755`).
This plugin ships three of them, so it is a live concern.
- **A global pre-push hook auto-formats** markdown/JSON and may append a
`style: auto-format from pre-push hooks` commit during push — re-check
`git status` and ahead/behind afterwards and push again if it committed.
- **Version bumps are mandatory** for any behavioural change; Claude Code caches
by `version`.
Install (once published): `claude plugin add https://claude-plugins.oleks.space/plugins/inference-arbitrage.git`
Note that publishing is **two independent actions**: adding to the local
`oleks-local` dev marketplace (`~/projects/claude-plugins/.claude-plugin/marketplace.json`,
`"source": "./inference-arbitrage"`) and publishing to the public index
(`~/projects/claude-plugin-index/plugins.json`, a separate repo). A plugin can be
in one and not the other, and several here are deliberately local-only. Local-dev
registration first is the right order — it allows `@oleks-local` testing through
Phase 7 before anything becomes publicly installable.
---
## 2. Directory layout
```
inference-arbitrage/
├── .claude-plugin/
│ └── plugin.json
├── agents/
│ └── offload-analyst.md # inference-arbitrage:offload-analyst-agent
├── skills/
│ ├── offload-audit/SKILL.md # run a full audit on a target plugin
│ ├── boundary-rubric/SKILL.md # classify ONE step; usable standalone
│ └── offload-trend/SKILL.md # diff vs history; chase stale recommendations
├── commands/
│ └── offload-audit.md # /offload-audit <plugin> [--days N]
├── bin/
│ ├── plugin-inventory # static pass → JSON
│ ├── offload-scan # dynamic pass → JSON
│ └── audit-snapshot # persist / load / diff snapshots
├── references/
│ ├── boundary-rubric.md # rubric.md, verbatim
│ ├── signals-catalog.md # spec.md §5, expanded with thresholds
│ ├── snapshot-schema.md # §6 below
│ └── issue-template.md # the issue body contract
├── tests/
│ ├── run-all.sh
│ ├── inventory.test.sh
│ ├── scan.test.sh # against synthetic fixture transcripts
│ ├── snapshot.test.sh
│ └── fixtures/
├── .woodpecker/test.yaml
├── README.md LICENSE .gitignore .markdownlint.json .shellcheckrc
```
### 2.1 The division of labour, stated as a constraint
Per FR-8, the plugin must obey its own thesis. The split is:
| Layer | Owns | Never does |
|---|---|---|
| `bin/*` | All parsing, aggregation, metric computation, snapshot diffing, signature normalization | Classify a candidate; decide what to file |
| `agents/offload-analyst.md` | Rubric application, boundary questions, ranking narrative, deciding what crosses the filing threshold | Parse a transcript; count anything by hand |
| `skills/*` | Procedure, thresholds, reporting contract, delegation protocol | Restate logic that lives in a script |
If a future change puts a loop or an aggregation into skill prose, that is a
regression, and `tests/` should be able to catch the coarse version of it (a
skill body containing a fenced multi-command pipeline that duplicates a `bin/`
capability).
---
## 3. `bin/plugin-inventory` — the static pass
**Language:** Python 3, stdlib only, offline. (Consistent with `cc-tokens`; no
npm, no PyYAML — a minimal frontmatter parser for the `---` block is ~30 lines
and avoids a dependency.)
```
plugin-inventory <path|name> [--json]
```
**Resolution:** a path is used directly; a bare name globs
`~/.claude/plugins/cache/*/<name>/*/`, filters out directories containing
`.orphaned_at`, and picks the highest semver.
**Output:** one JSON document.
```jsonc
{
"plugin": {"name": "...", "version": "...", "source": "oleks-local", "path": "..."},
"skills": [{
"name": "sweep-worktrees",
"description_words": 61,
"body_words": 1840,
"allowed_tools": ["Bash", "Read"],
"procedural_density": 0.42, // steps-with-commands / total blocks
"script_verbs": 14, "judgment_verbs": 9, "verb_ratio": 0.61,
"command_blocks": [{"sig": "git worktree list --porcelain", "count": 3}],
"rule_tables": 2,
"references": ["extraction-notes.md"]
}],
"agents": [{"name": "...", "model": "sonnet", "tools": [...],
"body_words": 900, "mechanical_tool_share": 0.8}],
"hooks": [{"file": "...", "event": "PreToolUse", "loc": 120}],
"scripts": [{"file": "bin/worktree-audit", "loc": 430, "lang": "bash"}],
"aggregate": {
"skill_body_words": 6200, "script_loc": 980,
"script_coverage": 0.158,
"duplicated_command_blocks": [{"sig": "...", "skills": ["a", "b"]}],
"hook_prose_drift": [...]
}
}
```
**Verb lexicons** live in `references/signals-catalog.md`, not hardcoded, so they
can be tuned without a code change — and so a tuning is a reviewable diff.
Notably, `plugin-inventory` is itself the proof of FR-8: reading a plugin's
structure is a `pure-script` step by every one of the five determinism tests, and
having a model do it by reading files would be the exact anti-pattern the plugin
reports on.
---
## 4. `bin/offload-scan` — the dynamic pass
```
offload-scan --plugin <name> [--days N | --since D --until D] [--json]
[--root ~/.claude/projects]
```
### 4.1 Pipeline
1. **Stream** `~/.claude/projects/**/*.jsonl` line by line. Never load a file
whole; histories are multi-gigabyte and this runs on a 16 GiB laptop.
2. **Filter** to assistant lines whose `attributionPlugin == <name>`, plus
sidechain lines whose `agentName` matches one of the target's agents.
3. **De-duplicate** turns by `(message.id, requestId)`, taking `max(output_tokens)`
— the two documented transcript traps. See §4.3: this is delegated, not
reimplemented.
4. **Reconstruct invocations**: contiguous runs of attributed turns within one
`sessionId`, split on a gap in `attributionSkill` or a configurable idle gap.
5. **Classify each turn** as mechanical / judgment / retry per spec §5.2.
6. **Normalize tool signatures** and mine recurring n-grams (spec §5.2.2).
7. **Aggregate** per skill and per agent; compute MTR, `offload_waste`,
read amplification, retry density, judgment density, fan-out multiplier,
`offload_value`.
8. **Compute attribution coverage** (FR-3.4): attributed turns vs. turns in
sessions where the target was active at all.
9. **Emit JSON.** Masked signatures and counts only — never raw content (FR-3.5).
### 4.2 Output shape
```jsonc
{
"window": {"since": "...", "until": "...", "sessions": 34},
"coverage": {"attributed_turns": 812, "candidate_turns": 1104, "ratio": 0.735,
"caveat": "hook-driven and unattributed agent turns excluded"},
"totals": {"weighted_tokens": 41200000, "cost_estimate_usd": 118.4},
"skills": [{
"skill": "anxious:kanban-flow",
"invocations": 22, "turns": 190,
"weighted_tokens": 9100000, "share_of_plugin": 0.22,
"mtr": 0.68, "judgment_density": 0.11,
"offload_waste": 6900000, "read_amplification": 14.2, "retry_density": 0.09,
"fanout_multiplier": 1.0,
"ngrams": [{"sig": ["list_issues(<str>)", "issue_read(<n>)", "label_read(<n>)"],
"recurrences": 17, "waste": 3100000}],
"offload_value": 8400000
}],
"agents": [ ... same shape ... ]
}
```
### 4.3 The `cc-tokens` dependency (open question Q1)
The pricing table, tier weighting, and the two de-duplication traps must have
exactly one home. Reimplementing them here is how the numbers silently diverge —
the traps are specifically the kind that produce plausible wrong answers.
Resolution order at runtime:
1. `cc-tokens` on `PATH` → shell out with `--json` for money numbers.
2. Otherwise glob `~/.claude/plugins/cache/*/token-budget/*/bin/cc-tokens`,
newest semver.
3. Otherwise **fail loudly** with an install hint. Never vendor, never
approximate.
`offload-scan` performs its own pass for the things `cc-tokens` does not model —
tool-sequence signatures, turn classification, n-gram mining — but takes token
and cost arithmetic from `cc-tokens`.
**Preferred long-term fix, to be proposed to the user:** contribute
`cc-tokens attribute --by skill|plugin --json` upstream to `token-budget`. The
attribution fields are already in the transcripts; `cc-tokens` is already the
right place for that arithmetic; and it would make `offload-scan` a thin
consumer. Filed as a task, not assumed.
---
## 5. `bin/audit-snapshot` — persistence and diffing
```
audit-snapshot write --target <name> --from <inventory.json> <scan.json> --classified <c.json>
audit-snapshot list --target <name>
audit-snapshot diff --target <name> [--against <run_id>]
```
**Storage.** The wiki repo of `<org>/claude-plugin-inference-arbitrage`, cloned to
a local cache and pushed via `cluster:gitea-agent` (this plugin never writes
Gitea directly).
```
Data/<target>/snapshots.jsonl # append-only, one JSON object per run
Audits/<target>/<YYYY-MM-DD>.md # human-readable run summary
Audits/<target>/Latest.md # rollup + open candidate table
Methodology/Rubric.md # versioned copy; snapshots pin rubric_version
Methodology/Calibration.md # threshold changes, with the evidence for each
```
Append-only JSONL in a git-backed wiki gives durable history, a free audit trail
via `git log`, and diffability, with no database. `Latest.md` is regenerated, not
appended.
**Candidate identity (FR-5.2)** resolves in this order — issue number first,
because it is the only identifier a human also uses:
1. Gitea issue number, once filed.
2. Normalized tool-sequence signature hash.
3. `(skill, slug)` assigned at first sighting.
`diff` emits per-candidate `new | persisting | grown | shrunk | resolved |
signature-drifted | stale`, with deltas expressed as **cost-per-invocation** and
**share-of-plugin-spend** (FR-5.4) — absolute tokens are shown as context only,
so a quiet week does not read as progress.
`resolved` requires issue-closed **and** measured cost below threshold with ≥N
invocations in a later window; issue-closed alone reports
`claimed-fixed-unconfirmed` (FR-5.5).
---
## 6. Snapshot schema
```jsonc
{
"run_id": "2026-07-29T14:02:11Z",
"rubric_version": "1.0.0",
"tool_versions": {"inference-arbitrage": "0.1.0", "cc-tokens": "0.1.0"},
"target": {"name": "anxious", "version": "0.31.0", "source": "oleks-local",
"repo": "oleks/claude-plugin-anxious"},
"window": {"since": "2026-07-22", "until": "2026-07-29", "days": 7},
"coverage": {"ratio": 0.735, "attributed_turns": 812},
"totals": {"weighted_tokens": 41200000, "share_mechanical": 0.61},
"candidates": [{
"candidate_id": "anxious/steward/label-derivation",
"signature_hash": "b41f…",
"skill": "anxious:steward-agent",
"position": "llm-over-script-digest",
"boundary_confidence": "high",
"determinism_tests": {"T1": true, "T2": true, "T3": true, "T4": true, "T5": true},
"falsifiability": {"signature": "...", "examples": 3, "overrule_case": "..."},
"measurement": {"invocations": 22, "offload_waste": 6900000,
"offload_value": 8400000, "share_of_plugin": 0.22,
"cost_per_invocation": 381818},
"digest_schema": "…what crosses the boundary…",
"escalation_path": "…what happens on surprise…",
"issue": "oleks/claude-plugin-anxious#41",
"status": "persisting"
}],
"boundary_questions": [ { medium/low candidates, same shape, no issue } ],
"notes": []
}
```
---
## 7. Runtime constraints
Hard constraints from the environment, not preferences:
- **Single streaming stdlib process.** No fan-out of one subagent per skill —
that violates both the emmett heavy-workload rule and the recorded
subagent-fan-out-cost finding, and would be an especially embarrassing way for
a token-optimization plugin to burn tokens.
- **No compile, test, or lint toolchain locally.** CI verification goes through
Woodpecker (`.woodpecker/test.yaml`), per the global rule.
- **Bounded memory.** Line-by-line JSONL streaming; aggregate state is O(skills ×
distinct signatures), not O(turns).
- **Offline.** No network in either pass. Only the two output paths touch the
network, and both go through delegation.
---
## 8. Skills
### `offload-audit` — the main procedure
*Triggers:* "audit this plugin for token optimization", "can any of this be a
script", "offload analysis", "where is this plugin wasting inference", "check
<plugin> for scripting opportunities".
Procedure: resolve target → `plugin-inventory``offload-scan` → apply the
rubric to each candidate (delegating to `boundary-rubric` for the per-candidate
classification) → threshold → file path A → write path B → report.
Carries the filing threshold, the "nothing to offload is a valid answer" rule,
and the coverage-caveat reporting contract.
### `boundary-rubric` — classify one step
*Triggers:* "should this be a script or LLM", "where does the boundary go", "is
this judgment or algorithm", "can I automate this step".
Usable standalone, outside an audit, on a step someone is about to build. This is
deliberate: the rubric is more valuable applied *before* the code exists than
after, and a standalone skill lets it be reached during ordinary design work.
Body carries the five tests, the four positions, the positioning questions, and
the falsifiability triple. `references/boundary-rubric.md` carries the full text
with worked examples.
### `offload-trend` — history and follow-through
*Triggers:* "has this plugin got worse", "offload trend", "are my
recommendations being acted on", "diff the last audit", "what did we say last
time".
Runs `audit-snapshot diff`, reports the volume-normalized trend, lists
recommendations still unaddressed with their issue refs in full `owner/repo#num`
form, and promotes any `low`/`medium` boundary question that has recurred with
growing cost.
---
## 9. Agent
**`inference-arbitrage:offload-analyst-agent`** — `model: sonnet` (the work is
structured classification over pre-computed JSON, not open synthesis; if
calibration proves weak on the boundary questions, revisit).
*Tools:* `Bash`, `Read`, `Grep`, `Glob`, `Skill`, `Agent` (for the `anxious:issuer`
delegation only).
**Read-only with respect to the target.** It never edits a plugin it audits and
never implements its own recommendations — that separation is what preserves its
ability to conclude "no offload here."
Behavioral rules carried in the agent body:
- Lead with attribution coverage and the mechanical-share headline. A candidate
list without coverage context is misleading.
- Name skills, never UUIDs — borrowed directly from `usage-agent`'s reporting
discipline, which exists for the same reason.
- Rank by measured value and say so; a user who acts only on item one should have
captured most of the benefit.
- **"Nothing to offload here" is a complete and valuable answer.** Do not
manufacture candidates to look thorough.
- Never file without the overrule case (rubric §5).
---
## 10. Output path A — filing, in detail
Per FR-6, delegation only:
```
offload-analyst → Agent(anxious:issuer-agent) → cluster:gitea-agent → Gitea
```
The analyst formulates; `issuer` decides repo, title, labels, milestone;
`gitea-agent` writes. This plugin holds no Gitea credentials and calls no Gitea
MCP tool.
**Issue body contract** (`references/issue-template.md`):
```markdown
<!-- ia-candidate: anxious/steward/label-derivation -->
**Measured cost.** 8.4M weighted tokens over 22 invocations (2026-07-22 → 07-29),
22% of this plugin's audited spend. MTR 0.68, judgment density 0.11.
Recurring sequence observed 17×: list_issues → issue_read → label_read.
**Boundary position.** llm-over-script-digest — confidence high.
Determinism tests: T1 ✓ T2 ✓ T3 ✓ T4 ✓ T5 ✓
**Proposed signature.**
`derive_labels(repo: str, issue: Issue) -> tuple[Labels, list[Ambiguity]]`
**Examples.** (3, including one edge case) …
**When a human would overrule this.**
**Escalation path.**
**What crosses the boundary.**
Audit: `Audits/anxious/2026-07-29` on <org>/claude-plugin-inference-arbitrage wiki.
Rubric v1.0.0.
```
Labels: `token-offload` (new, cross-cutting), plus the `anxious` four-axis
taxonomy as `issuer` determines — typically `kind/chore` or
`kind/capability-gap`, `area/agent-behavior`, `domain/agents`,
`activity/automate`.
Before filing, search the target repo for the `ia-candidate` marker; if present,
comment the updated measurement instead of opening a duplicate (FR-6.3).
---
## 11. Publishing
Per the `plugin-publishing` skill, and per §1.1 above:
1. Source tree at `~/projects/claude-plugins/<name>/` (rename from
`token-offload-audit` once the name is settled). Note the workspace root is
deliberately **not** a git repo — commit inside the plugin subdir, never at the
top level.
2. `git init -b main`, the workspace `.gitignore`, commit.
3. Create `<org>/claude-plugin-<name>` on Gitea with no auto-init, push.
4. Optionally register in the local `oleks-local` marketplace for `@oleks-local`
testing before going public — recommended for this plugin, since Phase 7
acceptance wants real runs before anyone can install it.
5. Add the entry to `~/projects/claude-plugin-index/plugins.json` (separate repo)
and push; the Gitea webhook refreshes the registry, which also acts as the
git smart-HTTP proxy that makes `claude plugin add` work.
6. Verify at `claude-plugins.oleks.space/api/plugins`, then `claude plugin add`.
**Every content change requires a version bump** — the cache at
`~/.claude/plugins/cache/<source>/<name>/<version>/` is keyed on version and a
same-version push is silently ignored.
Also create the wiki on the plugin's repo (path B depends on it) and the
`token-offload` label on each target repo as it is first audited (via `issuer`,
implicitly).
+327
View File
@@ -0,0 +1,327 @@
# rubric.md — where the boundary between script and intelligence belongs
> *"the problem here is to correctly understand a points where we need
> intelligence between script entities correctly positioned"*
This is the crux of the whole plugin, so it gets its own document. It ships
verbatim as `references/boundary-rubric.md`, and every snapshot records the
`rubric_version` it was classified under.
The rubric has five parts, applied in order:
1. **Five determinism tests** — is this step a program at all?
2. **Four positions** — the answer is not binary, and the interesting cases are
the two in the middle.
3. **The positioning principle** — given a pipeline, where exactly do you cut?
4. **Two failure modes** — named, with detection questions for each.
5. **The falsifiability triple** — the gate a candidate must pass to be filed.
---
## 1. The five determinism tests
A step may be offloaded to a script only if it passes **all five**. These are
necessary conditions, not a score to be averaged — a step that fails one has
failed, however well it does on the others.
### T1 — Closed input domain
*Can the inputs be enumerated or typed?*
Paths, exit codes, JSON documents, git refs, file contents, timestamps, integers,
a fixed set of enum values — closed. Free-form natural language, a user's stated
intent, an arbitrary error message from an unknown tool, the contents of a web
page — open.
The tell: if you can write the function's parameter list without using the type
"whatever the human said," it is closed.
### T2 — Stable output oracle
*Is there one right answer, and can you assert it in a test?*
If two competent engineers looking at the same input would produce the same
output, and you can write `assert f(x) == y`, there is an oracle. If the answer
is "it depends what we're trying to do here," there is not.
**This is the strongest single test**, because it is constructive. If you cannot
write the assertion, you do not have a specification, and code without a
specification is just inference with worse error messages. The inability to write
the test *is* the finding.
### T3 — Invariance
*Same input, same output, every time — across runs, models, and days?*
Then a script is strictly better: it is free, instant, and cannot drift.
The diagnostic is subtle and worth stating carefully. Ask: **if this step gave a
different answer on two identical inputs, would that be a bug or a feature?** If
it would be a bug, the step is a script and running it on inference is
introducing nondeterminism into something that should not have any. If run-to-run
variation is *acceptable or desirable* — a differently-worded explanation, a
different but equally valid prioritization — that variance is the signature of
judgment, and freezing it into a script destroys something real.
### T4 — No semantic gap
*Does the answer depend on anything not present in the inputs?*
World knowledge, the user's goals, project history, what "good" means here,
whether a name is apt, whether a description is honest, whether this counts as a
bug — all live outside the inputs. A step that needs them cannot be a total
function over its arguments, and any script pretending otherwise is encoding a
guess as a fact.
T4 is the test that most often saves a step from being wrongly scripted.
### T5 — Bounded branch factor
*Is the decision tree enumerable?*
A dozen cases, or a hundred, or a table — enumerable. "Any of the open-ended ways
this could go wrong" — not. A step with an unbounded tail of special cases can
still be *partly* scripted (handle the enumerable cases, escalate the tail), but
that is position 3 below, not position 1.
### Reading the results
| Outcome | Position |
|---|---|
| All five pass | `pure-script` |
| T2 or T4 fails | `pure-inference` — stop, do not propose an offload |
| T1/T3/T5 fail but T2 and T4 hold | Decomposable — go to §3 and find the cut |
| Fails only because the *inputs* are unstructured | Often a **digest** problem, not a judgment problem — a script can structure the input even if it cannot make the decision |
That last row matters more than it looks. A great many steps look like judgment
only because nobody normalized the input.
---
## 2. The four positions
The binary "script or LLM" is the wrong frame, and forcing it is how both failure
modes in §4 happen. There are four positions, and the two in the middle are where
almost all the value is.
### Position 1 — `pure-script`
Deterministic end to end. Offload wholly; the model does not participate.
*Parse, glob, diff, count, aggregate, sort by a formula, validate against a
schema, git plumbing, format a report from structured data.*
Real instance: `cc-tokens`' entire JSONL pipeline. Nothing in the
de-duplication of `(message.id, requestId)` or the application of a pricing table
benefits from a model.
### Position 2 — `script-with-compiled-judgment`
The judgment is real, but it is made **once, at authoring time**, by a human or a
model, and then frozen into a rule, table, regex, threshold, or config. Runtime
is a pure script.
Real instance: `price-hunter`'s `require`/`exclude`/`broken`/`confirm` rules per
hunt. Deciding that a hunt for a *product* must match the product noun and not the
flavour word is judgment; applying that rule to ten thousand listings is not.
**The characteristic risk is rule rot.** The compiled judgment was correct for the
world as it was at authoring time. When the world moves, the script keeps applying
the old judgment, confidently and silently — and this is precisely why the
`tune-hunt` skill exists. Any position-2 recommendation must therefore carry a
staleness plan: how the rule is re-validated, and what signal says it has rotted.
A position-2 offload without a re-validation story is not a saving, it is a
deferred bug.
### Position 3 — `llm-over-script-digest` — **the dominant win**
The script gathers, filters, normalizes, joins, and ranks. The model judges only
the shortlist. This is the position the user is describing when they say
"intelligence between script entities correctly positioned," and it is where most
real candidates land.
Real instance: `worktree-audit`. The script enumerates worktrees and computes the
provably-safe classification (including the remote-upstream subtlety that
inference kept getting wrong); the model looks only at the must-reconcile pile
and decides whether a given dirty tree is someone's in-progress work worth
rescuing.
Prospective instance: `anxious:steward`. The deterministic axes of the label
taxonomy are computed by script across all repos; the model is handed only the
disagreements, the no-signal cases, and the weak-prior axes.
The whole design question in position 3 is **where** the cut goes — §3.
### Position 4 — `pure-inference`
Irreducible. Do not touch, and say so explicitly in the report, because "this is
correctly done by inference" is a finding.
*Intent, naming, prose, trade-off explanation for this particular user, novel
synthesis, disambiguation, adversarial or unstructured input, and — always — the
authorization call on an irreversible action.*
---
## 3. The positioning principle — where exactly to cut
Given a pipeline of steps, the boundary belongs at its **narrowest waist**: the
point where, upstream, volume is high and every transformation is mechanical and
total; *at* the point, the representation is small and typed; and downstream, the
decisions are few and semantic.
Formally, you are choosing the cut that minimizes
```
tokens_crossing_the_boundary × judgment_calls_downstream
```
subject to one hard constraint that dominates the optimization:
> **Semantic sufficiency.** The digest crossing the boundary must contain
> everything the judgment rests on.
A cut that saves tokens but strips the evidence the decision depends on is a
**false offload**, and it is *worse than no offload*: it doesn't slow the model
down, it makes it wrong, and it hides the reason. The model downstream cannot
know what the script filtered away, so it will answer confidently from a
mutilated view.
### The four positioning questions
Ask these in order about any proposed cut:
**P1 — What crosses?** Write the schema of the digest. If you cannot write it,
you have not found the boundary — you have only found a place where you would
like one to be.
**P2 — What is lost?** For each field the script drops, name the judgment that
would have used it. If any judgment downstream would have used a dropped field,
the cut is in the wrong place. Move it upstream.
**P3 — What happens on surprise?** When the script meets an input it was not
designed for, does it escalate to the model, or does it silently return a wrong
answer? **A correct boundary always has an escalation path.** A script whose only
options are "handle it" and "handle it wrong" is at the wrong altitude. This is
the same shape as the taxonomy's own axis-1 rule ending in *"no signal — defer to
human"*.
**P4 — Who owns the irreversible action?** The script may compute, propose, rank,
and prepare. It must not *authorize*. Anything that deletes, deploys, pushes,
sends, or spends stays behind a judgment gate — and this holds even when all five
determinism tests pass, because determinism is not the same as safety. Blast
radius is an independent axis and it overrides the rubric.
### The compression heuristic
Good cuts show a large information-compression ratio with near-zero semantic loss.
A script that turns 400 KB of `git worktree list` + `git log` + `git status`
output into a 12-row typed table of `(path, branch, dirty, has_remote_upstream,
classification)` is a good cut: three orders of magnitude of compression, and
nothing a human would have used got dropped.
A script that turns the same input into `"7 worktrees, 3 stale"` is a bad cut at
the same location — same compression, catastrophic semantic loss. The location was
right; the schema was wrong. **Most bad cuts are bad schemas at good locations**,
which is why P1 and P2 come before anything else.
---
## 4. The two failure modes
### 4a — Over-scripting (premature crystallization)
Freezing a judgment that legitimately varies. The system becomes fast, confident,
and wrong, and — the real damage — *nobody is watching anymore*, because the step
no longer produces output a human reads.
Detection questions:
- Does the correct answer depend on anything not in the inputs? (T4)
- Would the rule need to change when the world changes, and would anyone notice?
- Are you scripting the decision, or scripting the *gathering* for the decision?
The second is almost always available and almost always safe.
- Is there a case where a competent human would look at the script's output and
say "no, not this time"? If yes, that case must escalate, not be encoded.
### 4b — Under-scripting (inference as interpreter)
Using a model as a slow, expensive, nondeterministic interpreter for something a
program does correctly. This is the status quo the plugin exists to find, and it
is invisible precisely because it *works* — the answers are right, just paid for
at ~one full context re-read per step.
Detection questions:
- Does the same tool sequence recur across invocations? (§5.2.2 of `spec.md`
measures exactly this)
- Has the model ever got this step *wrong* in a way a script could not? A
recorded error here is strong evidence — the `@{u}` bug in worktree
classification is the canonical example.
- Is the model re-deriving, every invocation, a procedure that is already written
down in the skill body as numbered steps?
- Would a competent engineer, shown this step, reach for a shell one-liner?
### The asymmetry
These two errors are not equally bad, and the rubric is deliberately biased.
**Under-scripting costs money. Over-scripting costs correctness, silently, and
compounds.** A missed candidate means a plugin stays somewhat expensive and the
next audit will find it again. A wrong candidate means a script that gives a
confident wrong answer for a year with no one watching. Hence the conservative
filing threshold, and hence §5.
---
## 5. The falsifiability triple — the gate
Before any candidate may be filed as an issue (FR-4.2), the agent must be able to
write, **in the issue body itself**, all three of:
**(a) The function signature.** Name, typed parameters, typed return. If it
cannot be typed, T1 or T2 failed and this is not a script.
**(b) Three input→expected-output pairs, one of which is an edge case.** These are
the test assertions. If they cannot be written, T2 failed — there is no oracle,
and what looked like an algorithm is a judgment wearing an algorithm's clothes.
**(c) The case where a human would overrule the script.** The input where the
script returns its defined answer and a competent human says "no, not here" —
plus what the escalation path does about it.
**(c) is the load-bearing element**, and it is deliberately the hardest.
An agent that cannot name the overrule case has not understood the boundary; it
has only noticed that a step looks repetitive. That is exactly the state in which
over-scripting happens. So the rule is strict and mechanical:
> **No overrule case → downgrade to boundary question.** Report it to the user as
> an open question about where the boundary belongs. Never file it.
The inverse is also informative. If the overrule case turns out to be *common*
if the human would overrule often — then the step is not a position-1 script at
all. It is position 3, and the cut belongs earlier in the pipeline, at the
gather/normalize stage rather than at the decision.
Boundary questions are a first-class output, not a failure. They are the plugin
asking the user for the judgment the user is uniquely able to supply, which is
precisely the behavior the rubric is trying to teach.
---
## 6. Confidence scoring
| `boundary_confidence` | Requires |
|---|---|
| `high` | All five determinism tests pass, position 1 or 3, falsifiability triple complete including the overrule case, escalation path defined |
| `medium` | Tests pass but the triple is partial, or position 2 without a staleness plan, or the digest schema (P1) is not yet writable |
| `low` | A repetition or cost signal exists, but T2 or T4 is in doubt |
Only `high` is filed as an issue. `medium` and `low` are reported as boundary
questions in the run summary and carried forward in snapshots, so that a later
audit with more measured evidence can promote them.
Carrying them forward matters: a `low` candidate that recurs across four audits
with growing cost is itself a signal worth surfacing, even if no single run could
confidently classify it.
+549
View File
@@ -0,0 +1,549 @@
# spec.md — `inference-arbitrage`
**One-line description.** Audits a Claude Code plugin — its definition *and* its
measured transcript usage — for work that is currently paid for as LLM inference
but is really a deterministic algorithm, and proposes where to cut the boundary
between script and judgment.
**Status.** Spec only. No code exists.
**Naming note.** This spec proposes the plugin be named **`inference-arbitrage`**
(rationale in `plan.md` §1). The checkout currently lives at
`~/projects/claude-plugins/token-offload-audit/` under a placeholder directory
name. Per this workspace's convention the directory name must equal the manifest
`name`, so if the proposed name is accepted the directory is renamed at scaffold
time (TASKS 1.1) — the spec files simply move with it. If the user prefers
`token-offload-audit`, everything here holds with that substituted; nothing in
the design depends on the name.
**Target repo** (to be created): `claude-plugin-<name>`, in **either** the
`oleks` or the `kotkan` org. The org is not derivable from the name and must be
chosen deliberately — `kotkan` ("Kotkan Workload Development") already holds the
two plugins this one is most closely related to (`kotkan/claude-plugin-token-budget`,
`kotkan/claude-plugin-worktree-discipline`) as well as `anti-patterns`, which
argues for `kotkan`. Published to `claude-plugins.oleks.space` per the
`plugin-publishing` skill.
---
## 1. Why
### 1.1 The economics that make this worth building
`token-budget` established the governing fact of Claude Code cost, and this spec
inherits it rather than re-deriving it:
> ~95%+ of raw tokens are `cache_read`; well under 1% is generated output. Cost
> scales with **(session length × context size)**, largely independent of work
> produced.
The consequence for *this* plugin is sharp and non-obvious:
**Every turn the model spends emitting a mechanical tool call costs approximately
one full context re-read.** A step like "run `git worktree list`, parse it, and
decide which entries are stale" does not cost "a bit of reasoning" — it costs
~200k tokens of re-read, at 0.1× input rate, *per turn*, and it costs the same
whether the step was hard or trivial. A 6-turn deterministic loop inside a skill
is 6 full context re-reads to produce an answer a 40-line script produces for
zero tokens.
So the offload lever is not "the model writes too much." It is: **how many turns
does the model have to stay in the loop for, on steps that produce no judgment.**
That is a measurable quantity, and it is what this plugin measures.
### 1.2 The precedent — this already worked, twice, by hand
Two plugins in this environment already made the cut correctly, and both are
usable as calibration fixtures:
- **`token-budget`.** The entire transcript-parsing algorithm — JSONL streaming,
`(message.id, requestId)` de-duplication, streaming-snapshot `output_tokens`
maxing, the pricing table, per-session aggregation — lives in `bin/cc-tokens`
(606 lines of stdlib Python). The skills carry only what is irreducibly
judgment: *which* subcommand this question needs, how to frame the dollar
caveat, which of two remediation skills a finding routes to, and whether the
answer is "nothing is structurally wrong here." The boundary is drawn at
exactly the right place, and nobody has to say so out loud because the split
is structural.
- **`worktree-discipline`.** The rule that a worktree is only safe to remove if
its commits have a *remote* upstream was being got wrong by inference: a bare
`@{u}` check resolves even in a repo with no remote, because `git worktree
add -b` sets up local tracking. The global `CLAUDE.md` records this explicitly
— *"Use `worktree-audit`, which makes that call correctly; a bare `@{u}` check
does not."* The classification moved into `bin/worktree-audit`; the skill kept
the judgment (is this branch someone's in-progress work? is this worth
reconciling or rescuing?).
Both were discovered by an incident. The point of this plugin is to find the
third one **before** the incident, and to find it with a number attached.
### 1.3 The user's actual framing
> "searching for the possibility points to offload algorithms into scripts from
> bare inference … not only as a one-time optimization, but also as ongoing
> efficiency monitoring with usage analysis … the problem here is to correctly
> understand a points where we need intelligence between script entities
> correctly positioned"
Three requirements fall out, and they are the three functional pillars below:
find the candidates (static + dynamic), **place the boundary correctly** (the
crux — see `rubric.md`), and keep watching over time.
---
## 2. Scope
### 2.1 In scope
- Auditing **any** Claude Code plugin, addressed by directory path or by name
resolved against `~/.claude/plugins/cache/<source>/<name>/<version>/`.
- A **static pass** over the plugin definition (skills, agents, commands, hooks,
references, existing scripts).
- A **dynamic pass** over local session transcripts (`~/.claude/projects/**/*.jsonl`),
attributing measured token spend to the target plugin's skills and agents.
- A **boundary classification** of each candidate against a published rubric.
- **On-demand runs** that each persist a snapshot, so history accumulates and
successive runs can be diffed.
- Two output paths per run: tracked issues on the **target** plugin's repo, and a
durable summary on **this** plugin's own wiki.
### 2.2 Out of scope for v1
- **Implementing** the offload. The plugin proposes and measures; a human or a
separate work session writes the script. Rationale: an auditor that also
writes the code it recommends loses the ability to say "no offload here."
- Real-time hooks or a mandatory cron. Cadence is on-demand by explicit
instruction (user or another agent). A user may wrap it in the `loop` or
`schedule` skills if they want recurrence — that is their choice, not the
plugin's default.
- Auditing non-plugin agent code, MCP servers, or arbitrary repos.
- Any network-dependent token accounting. The dynamic pass is offline, like
`cc-tokens`.
### 2.3 Non-goals worth naming
- **Not a linter.** It does not enforce style on SKILL.md. It reasons about cost
and determinism only.
- **Not a benchmark.** It never re-runs a plugin to measure it. It reads what
already happened.
- **Not an optimizer of prose length.** Trimming a skill body is a `context-diet`
concern and belongs to `token-budget`. This plugin is about *turn count on
mechanical work*, which is a different and larger lever.
---
## 3. Functional requirements
### FR-1 — Target resolution
**FR-1.1** Accept a target as a filesystem path, or as a plugin name resolved
through `~/.claude/plugins/cache/*/<name>/` (picking the highest semver
directory, excluding any marked `.orphaned_at`).
**FR-1.2** Record, in every snapshot, the resolved absolute path, the plugin
`name` and `version` from `.claude-plugin/plugin.json`, and the source
marketplace. Audits of different versions must be comparable but distinguishable.
**FR-1.3** Refuse to audit a directory with no `.claude-plugin/plugin.json`, with
a clear error naming what was expected. Do not guess at a plugin-shaped layout.
### FR-2 — Static pass
**FR-2.1** Enumerate the target's surface deterministically, via a script, not by
having a model read files: `skills/*/SKILL.md` (name, description, trigger
phrases, `allowed-tools`, body word count, heading structure), `agents/*.md`
(name, `model`, `tools` allowlist, body word count), `commands/*`, `hooks/*` and
`hooks.json`, `bin/*`, `references/*`, `tests/*`.
**FR-2.2** Emit per-skill and per-agent **static signals** (defined in §5.1) as
structured JSON — never as prose. The model consumes the JSON.
**FR-2.3** Compute a **script-coverage ratio** per plugin: prose-instruction
volume vs. shipped executable volume, plus the count of *literal command blocks
appearing in more than one skill* (a duplicated command block is a missing
shared script).
**FR-2.4** The static pass must run to completion on a plugin with zero
transcript history. A brand-new plugin is auditable on definition alone; the
report must then state that no measured evidence exists and mark every candidate
`unmeasured`.
### FR-3 — Dynamic (usage) pass
**FR-3.1** Attribute measured token spend to the target plugin's skills using the
`attributionPlugin` / `attributionSkill` fields present on assistant lines in the
transcripts, and to its agents via `agentName` / `isSidechain`.
*Verified during spec research:* these fields exist and are populated
(`attributionPlugin: "memory"`, `attributionSkill: "memory:save"`, etc.), and
**every** attributed assistant line sampled carried a full `message.usage` block
(1129/1129 in a 60-transcript sample). Per-skill cost attribution is therefore
directly computable, not an estimate.
**FR-3.2** Reuse `token-budget`'s accounting rather than reimplementing it. The
two documented transcript traps — one message spanning many JSONL lines all
repeating the same `usage`, and `output_tokens` being streaming snapshots that
must be maxed per `(message.id, requestId)` — are exactly the kind of thing that
is silently got wrong on a reimplementation. See PLAN §4.3 for the mechanism and
its risk.
**FR-3.3** Compute the **offload signals** defined in §5.2 — Mechanical Turn
Ratio, repetition signature, read amplification, retry density, judgment density,
fan-out multiplier — per skill and per agent.
**FR-3.4** Report **attribution coverage** for every audit: what fraction of the
window's turns plausibly belonging to this plugin were actually attributable.
Work driven by a plugin's hooks, or by an agent invoked without a Skill call, may
carry no attribution. An audit that silently understates coverage is worse than
no audit, so coverage is a required headline field, not a footnote.
**FR-3.5** Emit only **shapes and counts** — normalized tool-call signatures,
counts, token sums. Never copy raw transcript content, file contents, command
arguments, or user prose into a snapshot, wiki page, or issue. Transcripts
contain secrets and private work. Argument literals are masked (see §5.2.2).
### FR-4 — Boundary classification
**FR-4.1** Every candidate is classified into exactly one of four positions —
`pure-script`, `script-with-compiled-judgment`, `llm-over-script-digest`,
`pure-inference` — by the rubric in `rubric.md`.
**FR-4.2** No candidate may be filed as an issue unless the agent can produce the
**falsifiability triple** (rubric §5): the proposed function signature, three
input→expected-output pairs including one edge case, and *the case where a human
would overrule the script*. Failure to produce the third element downgrades the
candidate to a **boundary question** reported to the user, never filed.
**FR-4.3** Each candidate carries a `boundary_confidence` (`high` / `medium` /
`low`) derived from which of the five determinism tests it passes, and an
`offload_value` in weighted tokens and as a share of the plugin's audited spend.
**FR-4.4** The audit must be able to conclude **"nothing to offload here"** and
say so plainly. `token-budget` is expected to produce exactly this result; if a
run against `token-budget` produces a list of confident offload candidates, the
rubric is miscalibrated and that is a bug in this plugin.
### FR-5 — Snapshot accumulation and diffing
**FR-5.1** Every run appends one machine-readable snapshot record to an
append-only store on this plugin's own Gitea wiki (schema in PLAN §6).
**FR-5.2** Candidates carry a **stable identity** across runs so that "3 of 5
prior recommendations are still unaddressed" is computable. Identity resolution
order: (1) an open/closed Gitea issue number, once filed — the strongest and
preferred identity; (2) the normalized tool-sequence signature; (3) the
`(skill, slug)` pair assigned at first sighting.
**FR-5.3** A run compares against the previous snapshot for the same target and
reports per-candidate status: `new`, `persisting`, `grown`, `shrunk`, `resolved`,
`signature-drifted`, `stale` (target version changed such that the skill no
longer exists).
**FR-5.4** Trend comparison must normalize for usage volume. Absolute token
deltas across windows of different activity are meaningless — a quiet week reads
as an improvement. Primary trend metrics are **cost per invocation** and **share
of the plugin's audited spend**; absolute tokens are reported as context only.
**FR-5.5** A candidate is `resolved` only when its issue is closed **and** the
signature's measured cost has dropped below threshold in a later window with at
least N invocations. Issue-closed-without-measurement is reported as
`claimed-fixed-unconfirmed`. Committed ≠ verified, in the same spirit as this
environment's `nixos-deploy-pending` discipline.
### FR-6 — Output path A: issues on the target repo
**FR-6.1** Each `high`-confidence candidate above the value threshold is filed as
a tracked issue **on the target plugin's own repo**, by delegating to
`anxious:issuer` (which shapes and dispatches; `cluster:gitea-agent` performs the
write). This plugin never writes to Gitea directly.
**FR-6.2** Follow the existing environment conventions rather than inventing a
scheme: full `owner/repo#num` references everywhere; labels from the `anxious`
four-axis taxonomy (`kind/chore` or `kind/capability-gap`,
`area/agent-behavior`, `domain/agents`, `activity/automate`), plus one new
cross-cutting label `token-offload`. Filing is implicit — the audit does not ask
permission to file, consistent with the global bug-handling rule.
**FR-6.3** **Idempotency is mandatory.** Each issue body carries a marker line
`<!-- ia-candidate: <candidate_id> -->`. Before filing, search the target repo
for that marker; if found, comment the updated measurement on the existing issue
instead of opening a duplicate. An on-demand audit that is re-run must not spam
the tracker — this is the single most likely way for the plugin to become
hated.
**FR-6.4** If the target plugin has no writable repo (a third-party or official
marketplace plugin), skip path A entirely, record `filing: unavailable` with the
reason, and report the candidate in the summary only.
### FR-7 — Output path B: the run summary
**FR-7.1** Each run writes a human-readable summary page to **this plugin's own**
Gitea wiki at `Audits/<target-plugin>/<YYYY-MM-DD>`, and updates
`Audits/<target-plugin>/Latest`.
**FR-7.2** Rationale for wiki-over-artifact, since the brief asks for one: the
summary must be durable, addressable across sessions, diffable by git, and
readable by a future agent that was not present for the run. A Gitea wiki page is
all four — the wiki is a git repo, so the snapshot JSONL and the prose report
version together and a `git log` shows the audit history for free. An Artifact is
a private, session-produced web page with no cross-run addressing and no
relationship to the repo; it cannot be the system of record. An Artifact **is**
offered as an optional rendered view — a trend chart across snapshots — when a
human explicitly asks to look at the history, generated from the same snapshot
data. Canonical store: wiki. Optional presentation: artifact.
**FR-7.3** The summary states, in this order: attribution coverage and window;
the composition headline (how much of the plugin's spend was mechanical); the
ranked candidate table with boundary classification and value; the diff against
the previous snapshot; the boundary questions (low-confidence candidates) posed
to the user; and the issues filed or updated, in full `owner/repo#num` form.
### FR-8 — Self-application
**FR-8.1** The plugin must pass its own audit. Its static pass is a script, its
transcript scan is a script, its snapshot diff is a script; the agent's job is
classification, judgment, and reporting. If any of these migrate into skill prose
executed by inference, the plugin is violating its own thesis.
**FR-8.2** Auditing `inference-arbitrage` with `inference-arbitrage` is an
acceptance test (TASKS §6).
---
## 4. Success criteria
| # | Criterion | How it is verified |
|---|---|---|
| S1 | Audits any plugin by path or name, with no target-specific code | Run against `token-budget`, `worktree-discipline`, `anxious`, `memory` |
| S2 | Correctly returns "nothing to offload" for a well-cut plugin | Run against `token-budget` → no `high`-confidence candidates |
| S3 | Rediscovers a known-correct historical cut | Run against `worktree-discipline` with `bin/worktree-audit` removed from the inventory → must flag the classification step as `high` confidence |
| S4 | Finds at least one real, defensible candidate in an unaudited plugin | Run against `anxious`; user agrees the candidate is genuine |
| S5 | Re-running does not duplicate issues | Run twice; second run comments, does not file |
| S6 | Trend across ≥2 snapshots is computable and volume-normalized | Two runs on different windows; diff reports per-invocation deltas |
| S7 | Passes its own audit | FR-8.2 |
| S8 | Runs within emmett's constraints | Single stdlib process, streaming, no subagent fan-out, no toolchain build |
S3 is the important one. It is the only criterion that tests the rubric against a
cut *known* to be correct, made by a human, for a reason that is written down.
---
## 5. Signal catalog
### 5.1 Static signals — what the definition says
These are computed by `bin/plugin-inventory` and consumed as JSON.
| Signal | Definition | Reads as |
|---|---|---|
| **Procedural density** | Fraction of a skill body that is numbered/imperative steps containing literal commands, vs. discursive prose | A recipe the model re-derives every invocation |
| **Script-verb count** | Occurrences of `count`, `sum`, `rank`, `sort`, `diff`, `parse`, `validate`, `enumerate`, `dedupe`, `format`, `check exists`, `compare` | Mechanical intent stated in prose |
| **Judgment-verb count** | Occurrences of `decide`, `judge`, `explain`, `prioritize`, `name`, `write`, `weigh`, `interpret`, `for this user` | Irreducible intent stated in prose |
| **Verb ratio** | script-verbs / (script-verbs + judgment-verbs) per skill | >0.6 with no `bin/` script is the loudest static smell |
| **Duplicated command blocks** | Identical or near-identical fenced command blocks appearing in ≥2 skills | A shared script that was never written |
| **Rule tables** | Markdown tables that are pure lookup (input → output, finding → action) | A dispatch table being narrated at inference time |
| **Script coverage** | executable LOC in `bin/` + `hooks/` vs. total skill+agent body words | Structural balance; low ratio + high verb ratio is the target profile |
| **Hook/prose drift** | Logic present both in a hook script and restated in skill prose | Two sources of truth; the prose will rot |
| **Tool allowlist shape** | An agent whose `tools` are dominated by read-only mechanical tools | The agent was built to fetch, not to judge |
**Worked example — `token-budget` (expected: clean).** `skills/token-audit` has
high procedural density (it prescribes three exact commands in order) and a rule
table (`finding → remediation skill`). But script coverage is high (606 LOC of
`cc-tokens` against ~3 skill bodies), the verb ratio in the *body* is
judgment-dominated ("report", "interpret", "state the caveat", "rank the fixes by
effect"), and — decisively — the prescribed commands are invocations of a script
that already exists. The static pass should score this as **already cut**, with
the residual candidate being at most "the fixed three-command opening sequence
could be one `cc-tokens audit` subcommand," classified `low` value.
**Worked example — `anxious` (expected: real candidates).** `agents/steward.md`
sweeps *every* `oleks/*` repo and reconciles labels, milestones, and board
placement. The taxonomy it reconciles against is a **pure lookup table** with a
documented deterministic ruleset (`references/taxonomy.md` axis 3 is explicitly
described as *"steward-owned, repo-derived (deterministic)"*, and axis 1's
derivation is a four-step priority list ending in "defer to human"). A
deterministic ruleset, written down, applied per repo by an LLM turn, over N
repos, is the canonical shape: the *derivation* is a script, the *"defer to
human"* branch and the axis-2/axis-4 weak priors are judgment. Expected
classification: `llm-over-script-digest` — a script computes proposed labels and
emits only the disagreements and the no-signal cases for the model to rule on.
### 5.2 Dynamic signals — what the transcripts show
Computed by `bin/offload-scan` over `~/.claude/projects/**/*.jsonl`, filtered to
turns attributed to the target.
#### 5.2.1 The primary metric — Mechanical Turn Ratio
A turn is **mechanical** if all of:
- it contains ≥1 `tool_use` block, and every one is drawn from the mechanical
tool set (read-only `Bash`, `Read`, `Grep`, `Glob`, `list_*` / `get_*` /
`*_read` MCP tools, `ToolSearch`), and
- its `output_tokens` are below a threshold (default 400) — the model emitted a
tool call and little else, and
- it is not immediately preceded by a tool error (that is a retry, counted
separately under 5.2.4).
```
MTR(skill) = mechanical_turns / attributed_turns
mechanical_tokens = Σ over mechanical turns of (cache_read + cache_write + input + output)
offload_waste(skill) = mechanical_tokens # weighted, via cc-tokens' tier normalization
```
`offload_waste` is the honest number: **tokens spent purely to keep the model in
the loop on steps that produced no judgment.** Because `cache_read` dominates, it
is very nearly `mechanical_turns × mean_context_size` — which is exactly why turn
count, not verbosity, is the lever.
An MTR above ~0.5 on a skill carrying meaningful spend is the primary flag. A
high MTR on a skill with trivial spend is noise and is not reported.
#### 5.2.2 Repetition signature — finding the algorithm in the wild
Normalize every tool call to a signature: tool name plus argument *shape*, with
volatile literals masked — paths → `<path>`, integers → `<n>`, hex/SHA → `<sha>`,
URLs → `<url>`, quoted free text → `<str>`. (This masking is also what satisfies
the privacy requirement FR-3.5.)
Within each invocation of a skill, extract the ordered signature sequence. Across
invocations, find n-grams of length ≥3 that recur in ≥3 distinct invocations.
**A frequently-recurring identical tool sequence is an algorithm, empirically
discovered.** This is the most direct possible answer to the user's ask, because
it does not rely on reading intent out of prose — it observes the procedure being
executed the same way repeatedly and prices it. Each recurring n-gram is reported
with its occurrence count and its measured `offload_waste`, and becomes a
candidate with a ready-made proposed script boundary.
#### 5.2.3 Read amplification
```
read_amplification = tokens_pulled_into_context_by_read_tools
/ tokens_of_that_material_referenced_in_subsequent_output
```
Approximated by comparing tool-result sizes to the model's next-turn output
length and its literal overlap with the result. High amplification means the
model is reading a haystack to find a needle — a filter/digest script belongs
upstream. This is the signature of the `llm-over-script-digest` position.
#### 5.2.4 Retry / self-correction density
Fraction of attributed turns that follow a tool error, or that repeat a
near-identical signature with adjusted arguments. Every retry is a full context
re-read. High density means the invocation is fiddly — the correct offload is
often a thin wrapper script that gets the invocation right once, rather than a
wholesale algorithm move. Cheap to build, immediately effective.
#### 5.2.5 Judgment density — the brake
Fraction of attributed turns with substantive `output_tokens` (above threshold)
and *no* `tool_use`, or containing `AskUserQuestion`. This is where the model was
actually thinking.
**Judgment density is a brake, not an accelerator.** High judgment density with
high spend means the plugin is doing what it should and must be reported as
healthy. The offload target is specifically **low judgment density × high spend**.
#### 5.2.6 Fan-out multiplier
Sidechain turns attributed to the target. Each subagent carries its own full
context, so a mechanical step performed inside a subagent costs a multiple of the
same step inline. A mechanical n-gram executing inside a fan-out is the
highest-value candidate class there is — and this environment already has a
recorded finding against one-subagent-per-tiny-step fan-out.
#### 5.2.7 Composite ranking
```
offload_value = offload_waste × (1 judgment_density) × repetition_factor
```
where `repetition_factor` = 1 + log₂(recurrences of the dominant n-gram).
Reported in weighted tokens and as a share of the plugin's audited spend.
**Filing threshold (all three required):** `offload_value` ≥ 2% of the audited
window's total spend **and** `boundary_confidence == high` **and** the
falsifiability triple was produced. Everything else is a boundary question or a
note in the summary. The threshold is deliberately conservative: the cost of a
missed candidate is a slightly expensive plugin; the cost of a wrong candidate is
a plugin that confidently does the wrong thing, forever, silently.
---
## 6. Risks and open questions
These want the user's input before or during the build.
**Q1 — Cross-plugin dependency on `cc-tokens`.** The dynamic pass should not
reimplement `token-budget`'s accounting (FR-3.2), but a plugin cannot rely on
another plugin's `${CLAUDE_PLUGIN_ROOT}`. Options: (a) require `cc-tokens` on
`PATH` and fail loudly with an install hint; (b) glob
`~/.claude/plugins/cache/*/token-budget/*/bin/cc-tokens` and pick the newest;
(c) vendor a copy. Recommendation: (a) with (b) as fallback, and **never** (c) —
a vendored copy silently diverges on exactly the two traps that are hardest to
notice. Better still: contribute `cc-tokens attribute --by skill|plugin --json`
upstream to `token-budget`, so the money math has one home. *Decision needed.*
**Q2 — Attribution coverage may be thin.** `attributionSkill` appears only when a
Skill was formally invoked. Hook-driven work, agent work started without a Skill
call, and plugin logic that runs as part of a larger session may carry no
attribution at all. A 60-transcript sample surfaced ~1100 attributed lines across
a dozen skills — real, but not obviously complete. If coverage for a given target
is low, the audit's dynamic pass is weak and must say so loudly (FR-3.4). *Open:
is there a better attribution path for agent-driven work than `agentName`?*
**Q3 — Read access to a target's repo.** The audit reads the local plugin cache,
which needs no repo access. But FR-6 files issues on the target's repo, which
does. For `oleks/*` plugins this is fine. For third-party or official-marketplace
plugins there is no writable tracker, and the audit degrades to summary-only
(FR-6.4). Confirm that degradation is acceptable rather than an error.
**Q4 — False positives are the real failure mode.** Recommending a script where
judgment is needed produces something worse than the status quo: a fast,
confident, wrong answer with no one watching. The rubric's five determinism
tests, the mandatory "when would a human overrule this" element, and the
conservative filing threshold are the mitigations, and the plugin never
implements its own recommendations. The residual risk is that a plausible-looking
candidate gets built by a later session that does not re-check the boundary —
which is why the issue body must carry the overrule case, not just the proposal.
**Q5 — Goodhart.** A plugin that scores plugins on token cost will, if followed
blindly, push toward brittle over-scripting. Proposed counter-metric tracked on
the trend page: the rate of `token-offload` issues later closed as `wontfix`, and
of offload scripts subsequently reverted. If that rate climbs, the rubric is too
loose. *Open: is this worth building in v1, or noted as a manual check?*
**Q6 — Privacy of the summary surface.** Snapshots and wiki pages are derived
from transcripts that contain secrets, private prose, and corp work. FR-3.5's
masking is the control. Worth confirming the masking list is sufficient, and
whether audits of `imagex:*` (corp) plugins should be excluded from the shared
wiki entirely.
**Q7 — Local compute budget.** Scanning a multi-gigabyte transcript history on
emmett must be a single streaming stdlib process, never a fan-out of one subagent
per skill. Stated as a hard constraint in PLAN §7, flagged here so it is a
conscious decision rather than an accident.
**Q8 — Rubric versioning.** Snapshots record `rubric_version`. When the rubric
changes, old candidate classifications are not directly comparable. Proposal:
diffs across a rubric-version boundary are annotated, not suppressed. Confirm.
---
## 7. Incidental finding
While surveying the plugin cache for this spec, two leaked git worktrees were
found inside the installed plugin cache itself:
```
~/.claude/plugins/cache/oleks-local/worktree-discipline/1.12.0/.claude/worktrees/wf_fdb5df0f-811-2/
~/.claude/plugins/cache/oleks-local/worktree-discipline/1.12.0/.claude/worktrees/wf_fdb5df0f-811-3/
```
Each carries a full copy of the plugin tree. These are unnamed-Workflow-subagent
worktrees (`wf_*`) that leaked into a *cache* directory — a location `sweep-worktrees`
is unlikely to be pointed at, and one that a plugin reinstall would silently
orphan. Filed as **kotkan/claude-plugin-worktree-discipline#9**; not in scope for
this spec.
+171
View File
@@ -0,0 +1,171 @@
# tasks.md — build order for `inference-arbitrage`
Ordered so that each phase is independently verifiable and the risky part (the
rubric) is calibrated against known-answer fixtures **before** anything is filed
to a real tracker.
Track as Gitea issues on `<org>/claude-plugin-inference-arbitrage` under a
`v0.1.0` milestone, via `anxious:issuer`. Reference in full `owner/repo#num` form.
---
## Phase 0 — Decisions to close first
These block design, not just code. Answers go into `Methodology/Calibration.md`.
- [ ] **0.1** Resolve Q1: the `cc-tokens` dependency. Shell-out + cache-glob
fallback, or propose `cc-tokens attribute --by skill|plugin --json`
upstream to `token-budget` first? Upstreaming is cleaner but adds a
dependency on landing a change in another plugin.
- [ ] **0.2** Resolve Q3: confirm summary-only degradation is acceptable for
plugins with no writable repo (official/third-party marketplace).
- [ ] **0.3** Resolve Q6: confirm the masking list, and decide whether corp
(`imagex:*`) plugins are excluded from the shared wiki entirely.
- [ ] **0.4** Resolve Q5: build the Goodhart counter-metric (wontfix/revert rate)
in v1, or note it as a manual check?
- [ ] **0.5** Confirm the plugin name `inference-arbitrage` (vs. keeping the
placeholder `token-offload-audit`). The directory name must equal the
manifest name, so this decision also fixes the checkout path.
- [ ] **0.6** Choose the Gitea **org**: `oleks` or `kotkan`. Not derivable from
the name. `kotkan` already holds `token-budget`, `worktree-discipline`, and
`anti-patterns` — the three closest relatives — which argues for `kotkan`.
## Phase 1 — Repo scaffold
- [ ] **1.1** Rename `~/projects/claude-plugins/token-offload-audit/` to match the
confirmed manifest name (workspace convention: directory == manifest
`name`), and build out the PLAN §2 layout: `plugin.json` v0.1.0, MIT
LICENSE, `.gitignore` with `.cache/`, `.claude/`, `*.log`. The four spec
files move with the directory and stay in-tree as the design record.
- [ ] **1.2** `git init -b main` **inside the plugin subdir** (the workspace root
is deliberately not a repo), create `<org>/claude-plugin-<name>` on Gitea
with no auto-init (via `cluster:gitea-agent`), push `main`.
- [ ] **1.2b** Verify `bin/` scripts carry the executable bit
(`git ls-files -s``100755`) — three shipped scripts make this a live
concern. Re-check `git status` after every push: the global pre-push hook
auto-formats markdown/JSON and may append its own commit.
- [ ] **1.2c** Register in the local `oleks-local` marketplace
(`~/projects/claude-plugins/.claude-plugin/marketplace.json`,
`"source": "./<name>"`) so Phases 27 can test via `@oleks-local` without
publishing anything publicly.
- [ ] **1.3** Initialize its wiki with `Methodology/Rubric.md` (rubric.md
verbatim, v1.0.0) and empty `Audits/` + `Data/` scaffolding.
- [ ] **1.4** `.woodpecker/test.yaml` running `tests/run-all.sh`. **CI verifies;
nothing heavy runs on emmett.**
## Phase 2 — Static pass
- [ ] **2.1** `bin/plugin-inventory`: target resolution (path or name, semver
pick, skip `.orphaned_at`), frontmatter parsing (~30 lines, no PyYAML).
- [ ] **2.2** Signal extraction: procedural density, verb ratio, command-block
signatures, rule tables, script coverage, hook/prose drift.
- [ ] **2.3** `references/signals-catalog.md` with the verb lexicons and
thresholds, loaded at runtime — tuning must be a reviewable diff.
- [ ] **2.4** `tests/inventory.test.sh` against a fixture plugin tree.
- [ ] **2.5** **Calibration run** against `token-budget`, `worktree-discipline`,
`anxious`, `memory`. Record the actual numbers in
`Methodology/Calibration.md` and tune thresholds until the ordering matches
the human read: `token-budget` cleanest, `anxious` richest in candidates.
## Phase 3 — Dynamic pass
- [ ] **3.1** `bin/offload-scan` skeleton: streaming JSONL reader, attribution
filter, `(message.id, requestId)` de-dup — delegating token/cost arithmetic
to `cc-tokens` per decision 0.1. **Do not reimplement the pricing table or
the streaming-snapshot max.**
- [ ] **3.2** Invocation reconstruction (contiguous attributed runs per session).
- [ ] **3.3** Turn classification: mechanical / judgment / retry, with the
mechanical tool set and the output-token threshold in the signals catalog.
- [ ] **3.4** Tool-signature normalization and masking (also satisfies FR-3.5).
- [ ] **3.5** N-gram mining: recurring sequences ≥3 long in ≥3 invocations.
- [ ] **3.6** Metrics: MTR, `offload_waste`, read amplification, retry density,
judgment density, fan-out multiplier, composite `offload_value`.
- [ ] **3.7** Attribution coverage computation (FR-3.4).
- [ ] **3.8** `tests/scan.test.sh` against **synthetic** fixture transcripts with
hand-computed expected metrics. Do not test against real transcripts —
they contain private content and they change.
- [ ] **3.9** Verify runtime and memory on the real history: one process,
streaming, bounded. Record wall time in the calibration page.
## Phase 4 — Rubric and classification
- [ ] **4.1** `references/boundary-rubric.md` = `rubric.md` verbatim.
- [ ] **4.2** `skills/boundary-rubric/SKILL.md` — usable standalone on a single
step, before any code exists.
- [ ] **4.3** `agents/offload-analyst.md` with the PLAN §9 behavioral rules,
including the hard gate: **no overrule case → boundary question, never
filed.**
- [ ] **4.4** **Calibration against known answers — the load-bearing test.**
- `token-budget`**must** yield zero `high`-confidence candidates (S2).
- `worktree-discipline` with `bin/worktree-audit` masked out of the
inventory → **must** flag worktree classification as `high` (S3).
If either fails, the rubric is wrong and Phase 5 does not start.
- [ ] **4.5** Record the calibration outcome and any threshold change, with its
evidence, in `Methodology/Calibration.md`.
## Phase 5 — Snapshots and trend
- [ ] **5.1** `bin/audit-snapshot write|list|diff`; schema per PLAN §6.
- [ ] **5.2** Candidate identity resolution (issue → signature → slug) and
signature-drift detection.
- [ ] **5.3** Volume-normalized diffing: cost-per-invocation and
share-of-spend primary, absolute tokens as context only (FR-5.4).
- [ ] **5.4** `resolved` vs `claimed-fixed-unconfirmed` logic (FR-5.5).
- [ ] **5.5** Wiki read/write through `cluster:gitea-agent`; `Latest.md`
regeneration.
- [ ] **5.6** `skills/offload-trend/SKILL.md`.
- [ ] **5.7** `tests/snapshot.test.sh` — two synthetic snapshots, assert the diff.
## Phase 6 — Output paths
- [ ] **6.1** `references/issue-template.md` per PLAN §10.
- [ ] **6.2** `anxious:issuer` delegation path; create the `token-offload` label
on a target repo on first audit (implicitly, via `issuer`).
- [ ] **6.3** **Idempotency (FR-6.3)**`ia-candidate` marker search before
filing; comment on the existing issue instead of duplicating. Verify by
running the same audit twice (S5). This is the highest-consequence
correctness detail in the whole plugin.
- [ ] **6.4** Summary page generation → `Audits/<target>/<date>` + `Latest`.
- [ ] **6.5** Graceful `filing: unavailable` for repo-less targets (FR-6.4).
- [ ] **6.6** `skills/offload-audit/SKILL.md` and `commands/offload-audit.md`
tying the whole procedure together.
## Phase 7 — Acceptance
- [ ] **7.1** S1 — audits four different plugins with no target-specific code.
- [ ] **7.2** S2 — `token-budget` → "nothing to offload."
- [ ] **7.3** S3 — masked `worktree-discipline` → rediscovers the known cut.
- [ ] **7.4** S4 — a real candidate in `anxious` that the user agrees is genuine.
**User sign-off required before the first issue is filed to a real repo.**
- [ ] **7.5** S5 — double-run files no duplicate.
- [ ] **7.6** S6 — two-snapshot trend is computable and volume-normalized.
- [ ] **7.7** S7 — **audit `inference-arbitrage` with `inference-arbitrage`**
(FR-8.2). Any parsing or aggregation found living in skill prose is a bug
to fix before release.
- [ ] **7.8** S8 — confirm the run stays inside emmett's constraints.
## Phase 8 — Publish
- [ ] **8.1** README with the thesis, the rubric summary, and worked examples.
- [ ] **8.2** Publish per `plugin-publishing`: push, index entry in
`~/projects/claude-plugin-index/plugins.json`, verify
`claude-plugins.oleks.space/api/plugins`, `claude plugin add`.
- [ ] **8.3** Optional, from decision 0.1: PR
`cc-tokens attribute --by skill|plugin` to `token-budget`.
- [ ] **8.4** Memory checkpoint — the rubric and the calibration findings into
`wing_claude_memory`, room `working-practice`.
---
## Dependency notes
- Phase 3 depends on decision **0.1**.
- Phase 4.4 gates Phase 5 and 6. **Do not file a single real issue until the
rubric passes calibration** — the plugin's credibility is spent on its first
wrong recommendation.
- Phase 2 and Phase 3 are otherwise independent and can proceed in parallel; both
feed Phase 4.
- The static pass alone (Phases 12 + 4) is already a shippable v0.1.0 if the
dynamic pass proves harder than expected. It answers "what could be a script"
without "and what is it costing you" — less compelling, but useful, and it
degrades honestly by marking every candidate `unmeasured` (FR-2.4).