Files
claude-plugin-inference-arb…/design/tasks.md
T
oleks dcec9cabfb FR-4.5: evidence-stability gate — refuse to file a candidate whose confidence moved with the window
Phase 7 filed anxious/agent-wip/release-policy-derivation as `measured` at
11.94% of audited spend on a 30-day window. The same candidate over 7 days came
back `thin` at 0.00%. Only the date range differed. A human reading either issue
body alone cannot know the other exists, which is rubric.md §4a's failure mode
reached through the plugin's own headline number.

That was caught by a human running a second skeptical review — luck, not a
control. It is now mechanical, in the same register as FR-4.2's overrule-case
gate:

- bin/stability-classify matches each `file` candidate against the most recent
  prior snapshot recording it (FR-5.2 identity) and downgrades it to a boundary
  question if `measurement_strength` changed or the share crossed the 2% filing
  threshold between windows. Movement within a label is a trend, not an
  instability; no prior window is `unchecked`, not unstable (FR-2.4).
- The downgrade is a finding against THIS repo, not the target's — the target
  did not change, the auditor described it two ways. Carries an `ia-stability`
  marker so a re-run comments rather than duplicating.
- bin/ia_store.py factors the store and identity resolution out of
  audit-snapshot so both scripts hash signatures identically. A divergence there
  would make the gate match nothing and fail open.

Documented as spec FR-4.5 and rubric §5b (shipped verbatim in references/).
Wired into skills/offload-audit as step 4b, before filing-plan; steps 5-6 now
consume its output rather than classified.json.

S4 is retired in its old form — a human triage catching a bad candidate is not a
repeatable test — and re-passed as: the gate catches the real
release-policy-derivation case with no human in the loop. Asserted in
tests/stability.test.sh case (f) against the verbatim two-window Phase 7 output.

Closes kotkan/claude-plugin-inference-arbitrage#11
Suite: 178 assertions, exit 0.
2026-07-29 20:25:04 +03:00

184 lines
10 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# tasks.md — build order for `inference-arbitrage`
Ordered so that each phase is independently verifiable and the risky part (the
rubric) is calibrated against known-answer fixtures **before** anything is filed
to a real tracker.
Track as Gitea issues on `<org>/claude-plugin-inference-arbitrage` under a
`v0.1.0` milestone, via `anxious:issuer`. Reference in full `owner/repo#num` form.
---
## Phase 0 — Decisions to close first
These block design, not just code. Answers go into `Methodology/Calibration.md`.
- [ ] **0.1** Resolve Q1: the `cc-tokens` dependency. Shell-out + cache-glob
fallback, or propose `cc-tokens attribute --by skill|plugin --json`
upstream to `token-budget` first? Upstreaming is cleaner but adds a
dependency on landing a change in another plugin.
- [ ] **0.2** Resolve Q3: confirm summary-only degradation is acceptable for
plugins with no writable repo (official/third-party marketplace).
- [ ] **0.3** Resolve Q6: confirm the masking list, and decide whether corp
(`imagex:*`) plugins are excluded from the shared wiki entirely.
- [ ] **0.4** Resolve Q5: build the Goodhart counter-metric (wontfix/revert rate)
in v1, or note it as a manual check?
- [ ] **0.5** Confirm the plugin name `inference-arbitrage` (vs. keeping the
placeholder `token-offload-audit`). The directory name must equal the
manifest name, so this decision also fixes the checkout path.
- [ ] **0.6** Choose the Gitea **org**: `oleks` or `kotkan`. Not derivable from
the name. `kotkan` already holds `token-budget`, `worktree-discipline`, and
`anti-patterns` — the three closest relatives — which argues for `kotkan`.
## Phase 1 — Repo scaffold
- [ ] **1.1** Rename `~/projects/claude-plugins/token-offload-audit/` to match the
confirmed manifest name (workspace convention: directory == manifest
`name`), and build out the PLAN §2 layout: `plugin.json` v0.1.0, MIT
LICENSE, `.gitignore` with `.cache/`, `.claude/`, `*.log`. The four spec
files move with the directory and stay in-tree as the design record.
- [ ] **1.2** `git init -b main` **inside the plugin subdir** (the workspace root
is deliberately not a repo), create `<org>/claude-plugin-<name>` on Gitea
with no auto-init (via `cluster:gitea-agent`), push `main`.
- [ ] **1.2b** Verify `bin/` scripts carry the executable bit
(`git ls-files -s``100755`) — three shipped scripts make this a live
concern. Re-check `git status` after every push: the global pre-push hook
auto-formats markdown/JSON and may append its own commit.
- [ ] **1.2c** Register in the local `oleks-local` marketplace
(`~/projects/claude-plugins/.claude-plugin/marketplace.json`,
`"source": "./<name>"`) so Phases 27 can test via `@oleks-local` without
publishing anything publicly.
- [ ] **1.3** Initialize its wiki with `Methodology/Rubric.md` (rubric.md
verbatim, v1.0.0) and empty `Audits/` + `Data/` scaffolding.
- [ ] **1.4** `.woodpecker/test.yaml` running `tests/run-all.sh`. **CI verifies;
nothing heavy runs on emmett.**
## Phase 2 — Static pass
- [ ] **2.1** `bin/plugin-inventory`: target resolution (path or name, semver
pick, skip `.orphaned_at`), frontmatter parsing (~30 lines, no PyYAML).
- [ ] **2.2** Signal extraction: procedural density, verb ratio, command-block
signatures, rule tables, script coverage, hook/prose drift.
- [ ] **2.3** `references/signals-catalog.md` with the verb lexicons and
thresholds, loaded at runtime — tuning must be a reviewable diff.
- [ ] **2.4** `tests/inventory.test.sh` against a fixture plugin tree.
- [ ] **2.5** **Calibration run** against `token-budget`, `worktree-discipline`,
`anxious`, `memory`. Record the actual numbers in
`Methodology/Calibration.md` and tune thresholds until the ordering matches
the human read: `token-budget` cleanest, `anxious` richest in candidates.
## Phase 3 — Dynamic pass
- [ ] **3.1** `bin/offload-scan` skeleton: streaming JSONL reader, attribution
filter, `(message.id, requestId)` de-dup — delegating token/cost arithmetic
to `cc-tokens` per decision 0.1. **Do not reimplement the pricing table or
the streaming-snapshot max.**
- [ ] **3.2** Invocation reconstruction (contiguous attributed runs per session).
- [ ] **3.3** Turn classification: mechanical / judgment / retry, with the
mechanical tool set and the output-token threshold in the signals catalog.
- [ ] **3.4** Tool-signature normalization and masking (also satisfies FR-3.5).
- [ ] **3.5** N-gram mining: recurring sequences ≥3 long in ≥3 invocations.
- [ ] **3.6** Metrics: MTR, `offload_waste`, read amplification, retry density,
judgment density, fan-out multiplier, composite `offload_value`.
- [ ] **3.7** Attribution coverage computation (FR-3.4).
- [ ] **3.8** `tests/scan.test.sh` against **synthetic** fixture transcripts with
hand-computed expected metrics. Do not test against real transcripts —
they contain private content and they change.
- [ ] **3.9** Verify runtime and memory on the real history: one process,
streaming, bounded. Record wall time in the calibration page.
## Phase 4 — Rubric and classification
- [x] **4.1** `references/boundary-rubric.md` = `rubric.md` verbatim.
- [x] **4.2** `skills/boundary-rubric/SKILL.md` — usable standalone on a single
step, before any code exists.
- [x] **4.3** `agents/offload-analyst.md` with the PLAN §9 behavioral rules,
including the hard gate: **no overrule case → boundary question, never
filed.**
- [x] **4.4** **Calibration against known answers — the load-bearing test.**
- `token-budget`**must** yield zero `high`-confidence candidates (S2).
- `worktree-discipline` with `bin/worktree-audit` masked out of the
inventory → **must** flag worktree classification as `high` (S3).
- the two-window `anxious` fixture → **must** downgrade
`release-policy-derivation` on the stability gate (S4).
If either fails, the rubric is wrong and Phase 5 does not start.
- [ ] **4.5** Record the calibration outcome and any threshold change, with its
evidence, in `Methodology/Calibration.md`.
## Phase 5 — Snapshots and trend
- [x] **5.1** `bin/audit-snapshot write|list|diff`; schema per PLAN §6.
- [x] **5.2** Candidate identity resolution (issue → signature → slug) and
signature-drift detection.
- [x] **5.3** Volume-normalized diffing: cost-per-invocation and
share-of-spend primary, absolute tokens as context only (FR-5.4).
- [x] **5.4** `resolved` vs `claimed-fixed-unconfirmed` logic (FR-5.5).
- [x] **5.5** Wiki read/write; `Latest.md` regeneration. Resolved as: the
script is offline and filesystem-only, mirroring the wiki tree into a
local store; `audit-snapshot pages` emits what to publish and the caller
(skill, or `cluster:gitea-agent`) performs the wiki writes, and
`audit-snapshot import` seeds a cold cache from a wiki-fetched
`snapshots.jsonl`. See the `bin/audit-snapshot` docstring.
- [x] **5.6** `skills/offload-trend/SKILL.md`.
- [x] **5.7** `tests/snapshot.test.sh` — two synthetic snapshots, assert the diff.
## Phase 6 — Output paths
- [x] **6.1** `references/issue-template.md` per PLAN §10.
- [x] **6.2** `anxious:issuer` delegation path; create the `token-offload` label
on a target repo on first audit (implicitly, via `issuer`).
- [x] **6.3** **Idempotency (FR-6.3)**`ia-candidate` marker search before
filing; comment on the existing issue instead of duplicating. Verify by
running the same audit twice (S5). This is the highest-consequence
correctness detail in the whole plugin.
- [x] **6.4** Summary page generation → `Audits/<target>/<date>` + `Latest`.
- [x] **6.5** Graceful `filing: unavailable` for repo-less targets (FR-6.4).
- [x] **6.6** `skills/offload-audit/SKILL.md` and `commands/offload-audit.md`
tying the whole procedure together.
## Phase 7 — Acceptance
- [x] **7.1** S1 — audits four different plugins with no target-specific code.
- [x] **7.2** S2 — `token-budget` → "nothing to offload."
- [x] **7.3** S3 — masked `worktree-discipline` → rediscovers the known cut.
- [x] **7.4** S4 — the FR-4.5 evidence-stability gate mechanically catches a real
unstable candidate in `anxious` (`agent-wip/release-policy-derivation`)
without human review. The original framing — "a candidate the user agrees
is genuine" — is retired: a human triage catching a bad candidate is not a
repeatable test, and the one that ran caught this candidate *because* its
evidence flipped across windows. That check is now
`bin/stability-classify`, asserted against the real two-window evidence in
`tests/stability.test.sh` case (f).
- [x] **7.5** S5 — double-run files no duplicate.
- [x] **7.6** S6 — two-snapshot trend is computable and volume-normalized.
- [x] **7.7** S7 — **audit `inference-arbitrage` with `inference-arbitrage`**
(FR-8.2). Any parsing or aggregation found living in skill prose is a bug
to fix before release.
- [x] **7.8** S8 — confirm the run stays inside emmett's constraints.
## Phase 8 — Publish
- [ ] **8.1** README with the thesis, the rubric summary, and worked examples.
- [ ] **8.2** Publish per `plugin-publishing`: push, index entry in
`~/projects/claude-plugin-index/plugins.json`, verify
`claude-plugins.oleks.space/api/plugins`, `claude plugin add`.
- [ ] **8.3** Optional, from decision 0.1: PR
`cc-tokens attribute --by skill|plugin` to `token-budget`.
- [ ] **8.4** Memory checkpoint — the rubric and the calibration findings into
`wing_claude_memory`, room `working-practice`.
---
## Dependency notes
- Phase 3 depends on decision **0.1**.
- Phase 4.4 gates Phase 5 and 6. **Do not file a single real issue until the
rubric passes calibration** — the plugin's credibility is spent on its first
wrong recommendation.
- Phase 2 and Phase 3 are otherwise independent and can proceed in parallel; both
feed Phase 4.
- The static pass alone (Phases 12 + 4) is already a shippable v0.1.0 if the
dynamic pass proves harder than expected. It answers "what could be a script"
without "and what is it costing you" — less compelling, but useful, and it
degrades honestly by marking every candidate `unmeasured` (FR-2.4).