The auto-improvement loop: what would be built¶
Status: specification of mostly unbuilt work. Written 2026-08-11; step 1 partly built in 5.4.0, step 3 in 5.5.0, and step 5 built narrow (verbatim runs only) in 5.7.0 -- see Build order.
python -m src.review agenda is not a command and no skill consumes
anything below. Of the review aids, verbatim scan emits JSON as of
5.4.0 (#127) and the other two do not yet. This document states what
would be built and what it must satisfy, in the order it would be
built.
It contains no argument. Every "why" -- why the aid sits in the review layer rather than the drafting one, why three of six item classes may not be acted on, why the loop stops instead of running overnight, and the one documented rule it cannot satisfy without the user's approval -- is in AUTO-IMPROVEMENT-RATIONALE.md. Read that first if you are deciding whether to build this; read this one if you are building it.
Written for someone implementing it. It assumes ARCHITECTURE.md for the four layers and DRAFT-ITERATION.md for the dossier and the drift sweep.
Not covered here: the prose and house-style half, which has its own
detectors, its own persistence and its own roadmap --
HOUSE-STYLE.md. Corpus growth is out of scope entirely:
nothing specified below fetches a paper, writes bibliography.bib, or
writes the ledger.
Table of contents¶
- The shape
- 1.
--jsonon all three review aids - 2. The
agendaaid - 3. The
agenda-reviserskill - 4. Acceptance and rollback
- The requirements
- Who calls it, and when
- How the loop is reached
- The cost ladder
- What accumulates across drafts
- Build order
- What this does not change
The shape¶
Three sentences.
- The deterministic half is a fourth review aid: it reads the other three aids' findings plus the dossier's drift report and emits one ranked, deduplicated worklist.
- The generative half is a skill: it consumes that worklist, repairs what may be repaired unattended, re-verifies each repair, and hands the rest to the human.
- The human closes the loop: they accept the diff, and no code path runs the skill automatically.
1. --json on all three review aids¶
Partly built (5.4.0). src/review/__init__.py owns one output
contract for the layer. Extend it with a JSON sibling beside the
Markdown, at content/review/<topic>/<stem>.<aid>.json.
- The JSON is an additional serialisation of the same findings list, never a second computation. The printed Markdown stays the default and stays authoritative.
- No timestamp, per the layer's existing rule: two runs over an unchanged draft and corpus produce byte-identical JSON.
- This is #127's change applied to the layer rather than to
verbatim_checkalone, so the report contract does not fork.
What 5.4.0 built, per #127's scope: the layer-level plumbing --
review.envelope() (the payload's provenance, and the not-a-verdict
notice, as data) and review.write_json() -- plus verbatim scan
--json, which prints the payload and files it under --write.
provenance and coverage reuse that plumbing when their own issues
land; until then the agenda below finds one aid's JSON and not the
other two, which is the case step 2 already accounts for.
2. The agenda aid¶
python -m src.review agenda <draft> -- a fourth key in review.AIDS.
Deterministic, stdlib-only, no LLM, tier 1, takes no lock, exits 0 whatever
it finds.
Reads:
- the three aids'
.jsonfor this draft -- each optional, and skipped with a note when absent; src.dossier.status(draft), for missing citekeys and candidates;rejected.md-- a candidate already turned down with a reason is never re-proposed;sections.md, so every item carries a section anchor.
Writes: <stem>.agenda.md and <stem>.agenda.json under
content/review/<topic>/, via the layer's existing write().
Merges. One finding may appear in two aids' output; the agenda emits one item. This cross-signal merge is the work no individual aid can do.
Every finding carries a stable identity -- (aid, class, section
anchor, citekey, hash of the matched span) or equivalent -- so that "this
finding is gone" and cross-aid dedup are both decidable across runs.
Order: class order as the table below lists it, then #128's severity bucket within a class, then position in the draft.
Item classes¶
| Class | Source | Kind | Unattended? |
|---|---|---|---|
missing-citekey |
drift | defect -- the gate will fail on it | yes |
verbatim-run |
verbatim scan | defect above a span threshold | yes, except the long runs #129 reserves for the human. Built: overlap-reviser |
prose |
style_check (#107), steering.md |
no evidence delta | only the mechanically re-checkable subset -- HOUSE-STYLE.md |
unsupported-claim |
provenance | judgement | no -- surfaced |
uncited-source |
coverage | judgement | no -- surfaced |
candidate |
drift | a decision, usually correct to decline | no -- surfaced |
The prose class has no producer until #103 and #107 land; until then it
is an empty list.
3. The agenda-reviser skill¶
A skill, not a src/ module. Named for its input, like the two revisers
it joins: draft-reviser works from the dossier, corpus-reviser from a
whole-corpus re-search, this one from the agenda.
Per item:
- Dispatch the existing
draft-reviserdiscipline -- readscope.mdandsteering.mdfirst, edit inside the named section withEdit, never a whole-fileWrite. - Re-run
python -m src.draft gateand the aid that raised the finding. Accept only if both come back clean. - Re-run every other aid. If the total count of objective-class findings rose, revert the edit and escalate the item.
- Log the attempt in
revisions.md-- outcome included, refusals included.
Termination: at most two attempts per item; a second failure escalates the item and the loop moves on. One agenda pass per invocation. The loop never adds a claim, and on the unattended classes only removes or rewords existing ones.
4. Acceptance and rollback¶
Two levels.
- Before the pass: one
python -m src.draft dossier export <slug>. - Within the pass: the skill holds each section's pre-edit text and re-applies it when an item fails. Reverting one item leaves every earlier accepted item intact.
A failed attempt is logged in revisions.md and never in
rejected.md.
The human is presented with a diff plus the revisions.md entries. The
loop proposes and repairs; the human accepts.
The requirements¶
Eleven obligations, each phrased so a reviewer can tell whether it has been met. AUTO-IMPROVEMENT-RATIONALE.md says where each comes from.
| Requirement | |
|---|---|
| R1 | The skill's write-set is exactly the draft and revisions.md. It may execute an aid, the gate and style_check; it may not edit them, nor #128's allowlist, rejected.md, scope.md, or anything under the corpus layer. |
| R2 | Every finding carries an identity stable across runs. |
| R3 | An unattended item's check is binary. No continuous score is ever the thing being optimised. |
| R4 | After each accepted edit, every aid re-runs and the total objective-class count must not rise, else the edit reverts. |
| R5 | Reverting one item leaves every earlier accepted item intact. |
| R6 | Every attempt is logged with its outcome, refusals included, and no machine outcome is ever written to rejected.md. |
| R7 | Two attempts per item, one pass per invocation, then hand back. |
| R8 | Where a deletion and a rewrite both pass, the smaller diff wins. |
| R9 | The agenda taken before the pass is the recorded baseline, and the closing report is stated against it. |
| R10 | The aid is registered in both review.AIDS and __main__.AIDS; the skill's description names its triggers; and both appear in AGENTS.md's layer bullets, CLI.md, the README tables and mkdocs.yml. |
| R11 | No hook, no scheduled job and no other skill invokes the agenda-reviser skill. Its only trigger is a person asking. |
Who calls it, and when¶
The two halves have different answers, and conflating them is how this design would go wrong.
The aid: anyone, at any time. python -m src.review agenda <draft> is
free, deterministic, read-only and exits 0. It has exactly the standing of
the other three aids -- you run it because you want to know. No occasion is
privileged and none is required.
The skill: only a person, and only on a draft they consider finished.
Its SKILL.md description is the whole trigger, so in practice it runs
when a user says something like "clean up this draft" or "what is left to
fix here". Three occasions to name in that description:
- before rendering or submitting;
- after a
syncmoved the corpus, whendossier status --allhas named this draft; - on picking a draft back up after weeks away.
Nothing else may call it (R11): not a PostToolUse hook, not a
scheduled job, not a genre skill at the end of its own run, and not
draft-reviser on its own initiative.
Why each.
A stale input is reported, not merged. Reports carry no timestamp -- deliberately, so they diff cleanly -- so the check is file mtime. An aid report older than the draft is named as stale in the agenda's header and its findings marked, rather than presented as current; the header says to re-run that aid.
How the loop is reached¶
Nothing here is discovered by scanning the filesystem. Each half is reached by a different mechanism, and a piece that is built but not registered is dead code.
| Piece | How it is found | Consequence of omitting it |
|---|---|---|
The agenda aid |
a fourth key in review.AIDS (src/review/__init__.py) and in __main__.AIDS |
src/review/__main__.py raises RuntimeError if the two dicts disagree, so a half-registered aid fails loudly at import rather than writing a report nothing can find |
The agenda-reviser skill |
its SKILL.md frontmatter name and description |
This is the only trigger mechanism. A skill whose description does not match how a user phrases the request is never invoked, however correct its body |
| Both, for an agent working on a draft | AGENTS.md's layer bullets, which enumerate the aids (Layer 4) and the skills (Layer 2) | An agent following AGENTS.md would not know either exists |
| Both, for a human | CLI.md for the command and its flags; GENRE.md for which reviser handles what; README's review-aid block | Undiscoverable outside the source |
| The docs themselves | mkdocs.yml nav and the README documentation tables |
Invisible in the site nav. Not a build failure: nav.omitted_files is INFO-level, and mkdocs build --strict still passes |
The first two rows are load-bearing rather than administrative, and SOUL.md is deliberately absent from the list -- why.
The cost ladder¶
Do the free thing first, and pay only for what it could not decide -- LADDERS.md's existing shape.
- Detection and rejection, at zero tokens. Every aid,
style_checkand the gate are stdlib, deterministic and modelless. This rung must run to exhaustion before rung 2 begins. - A single-shot edit where the fix is local and the re-check binary: a dialect slip, a defect marker, an acronym. No subagent, no retrieval, no dossier read -- the finding already names the span.
- A dispatched reviser, only for items needing the surrounding argument in context. The expensive rung, and the short list.
Model tiering is the fourth rung. #75 settled the policy and #76 the measurement behind it, both closed; what is left is applying that policy to whatever mechanical stages this loop adds, which is a question for the build rather than a blocker on it.
What accumulates across drafts¶
Instrumentation the loop would produce that outlives any one draft. None of it is built, and none of it is read today.
- Which retrieval queries paid.
retrieval.mdlogs every call;evidence.mdandrejected.mdrecord what was kept and turned down. Across drafts, that is which query shapes yield kept evidence -- the evidence #63's parked evaluation harness would otherwise have to synthesise. - Which item classes the human accepts. Accepted and reverted items per class, across drafts, is a labelled record of the loop's own reliability. #130 requires the gating threshold to be tuned against real reports rather than guessed; this is those reports.
- Where the tokens went.
dossier statusalready totals retrieval cost per revision. Across drafts, that is the measurement TOKENS.md currently estimates.
The house-style counterpart -- standing preferences, the glossary, the allowlist -- is in HOUSE-STYLE.md.
Build order¶
Issue #126 already fixes this order; the change is to its scope, not its sequence.
- Settle the amendment. Not a coding task -- AUTO-IMPROVEMENT-RATIONALE.md.
- #127, widened to all three aids. Hard prerequisite for everything
below. Done for
verbatim scanin 5.4.0, on layer-level plumbing the other two aids reuse; they are the remainder of this step. - #128 -- severity buckets and the boilerplate allowlist. Done in
5.5.0 -- the allowlist shipped as per-host, gitignored data (like
config.toml), not version-controlled as first framed in HOUSE-STYLE.md; the constraints above (read-only to the loop, etc.) hold either way. agenda, the fourth aid. New. Useful on its own the day it lands, whether or not step 5 follows.- #129, widened -- the
agenda-reviserskill, over all defect classes rather than verbatim runs alone. Built narrow first, in 5.7.0:overlap-reviseris #129 as filed, over theverbatim-runclass alone, consumingverbatim scan --jsondirectly rather than an agenda. It did not wait for steps 2 and 4 because it did not need to -- one aid's JSON already existed, and a loop that repairs one class is the report step 7 has to be tuned against. Widening it is now a matter of giving it the agenda as an input and the other classes as work; the write-set, the two-attempt limit, the binary re-check and the person-only trigger are already what R1-R11 ask for.
Two pieces of that step landed with it, both in the review layer
rather than the skill: the scan payload's id (R2's stable identity,
for the verbatim-run class) and verbatim recheck, which is R3's
binary check and R4's did-anything-else-break count made
deterministic. agenda should reuse both rather than restate them.
6. #103 and #107 -- the copy-edit branch and style_check.py, giving
the prose class a producer and a consumer.
7. #130 -- the gating decision, last, tuned against real reports from
step 5.
Steps 4 and 5 are the only new work; the rest are open issues.
What this does not change¶
- No new gate.
src.draft gateremains the only one. #130 remains the only place that decision is taken. - No corpus growth. The loop never fetches, never writes the ledger, and never proposes a paper that is not already in it.
- No new layer. Four layers, one new aid in the fourth, one new skill in the second.
- No new entry point.
python -m src.review agenda <draft>is one verb under an existing front door, at depth 1. - The review layer still never blocks.
agendaexits 0 with a full worklist, exactly as the other three aids do with findings.