📚 Write a survey, start to finish¶
Status: tutorial. Written 2026-09-15.
Written for an author who wants a literature survey, a related-work section or a "state of the art" chapter out of their own library, and who has not used this pipeline before. Assumed: nothing. This page repeats what other documents also say, deliberately -- you should be able to finish a draft without leaving it. Not covered here: how the retrieval ranking works (RETRIEVAL.md) and why each rule exists (WRITING-STANDARDS.md).
Sister tutorials: a thesis chapter, a textbook chapter, a tutorial, a deep-research report.
🧭 Table of contents¶
- What you will have at the end
- Before you start
- Step 1: settle the slug, reader and scope
- Step 2: open the dossier
- Step 3: write an outline (optional, recommended)
- Step 4: ask for the draft
- Step 5: what the skill does while you wait
- Step 6: gate, references, render
- Step 7: read the review aids
- Step 8: change something
- Step 9: back it up
- When something goes wrong
🎯 What you will have at the end¶
For a draft you decide to call dt/survey:
| Path | What it is |
|---|---|
content/drafts/dt/survey.md |
the draft itself, with [@citekey] markers -- the canonical copy |
content/dossiers/dt/survey/ |
why it says what it says: scope, kept evidence, rejected candidates, every search run |
content/rendered/dt/survey.pdf |
the typeset PDF, IEEE-numbered |
content/rendered/dt/survey.tex |
the same, as LaTeX |
content/rendered/dt/survey.md |
a numbered Markdown copy, for a reader who will not open a PDF |
content/rendered/dt/survey.evidence.pdf |
the evidence sidecar: each cited source with the verbatim spans that justified it |
content/review/dt/survey.*.md |
the review reports you chose to keep |
Every citekey in that draft appears in your own .bib export and was
picked up by a real parse of a real PDF. That is the one guarantee this
pipeline exists to make, and the gate in step 6 is what enforces it.
🔧 Before you start¶
Three things, once per machine.
1. Install it and get a project directory. Either
1 2 3 | |
or clone the repository and cp config.toml.example config.toml.
Nothing on this page differs by which you chose.
2. Export your library to BibTeX, at papers/bibliography.bib:
1 2 | |
Your reference manager must export the PDF file paths with the entries, or the sync below will record citekeys with no text behind them and retrieval will find nothing. Zotero users: see EXPORT-ZOTERO-GROUPS.md.
3. Build the corpus. This reads the bib, parses each PDF and writes
content/ledger.sqlite:
1 | |
It takes minutes on a small library and can be re-run any time. Check it
found text, not just entries -- with no flags, ledger prints a summary:
1 2 | |
If nothing is parsed, the bib has no usable PDF paths -- fix that
before going further, because a survey cannot be written from a corpus
with no text in it.
Every command on this page also works as
python -m chitragupta.<layer> ...--python -m chitragupta.corpus syncis the same thing. Use whichever your install gives you.
📐 Step 1: settle the slug, reader and scope¶
Decide three things before any searching happens. They are cheap now and expensive later.
The slug is the draft's path under content/drafts/, without the
suffix. It may contain directories, and the dossier and every render
mirror it:
| You are writing | Use |
|---|---|
| a standalone survey | survey -> content/drafts/survey.md |
| one of several documents on one topic | dt/survey -> content/drafts/dt/survey.md |
| a chapter of a book | books/digital-twins/survey |
Moving it later means moving the dossier and the renders too, so pick now.
The reader is one sentence: "a first-year PhD student who knows control theory but not digital twins", "a grant reviewer outside the field", "the related-work section of a paper for IEEE TSE". Everything downstream -- how much is explained, which sources earn space -- follows from this.
The scope is two lists: what the survey covers, and what it deliberately does not. Write the second one. A reader who can tell an omission from an oversight trusts the rest.
🗂 Step 2: open the dossier¶
The dossier is the draft's working memory. Create it before drafting:
1 | |
That writes content/dossiers/dt/survey/ with eight files. Exactly one
of them is yours to fill in: scope.md. The other seven are written
for you as the draft is produced -- evidence.md (what was kept and
why), rejected.md (what was turned down and why), sections.md (which
section cites which citekey), retrieval.md (every search that ran),
steering.md, revisions.md, and a README.md explaining the rest.
What goes in scope.md¶
init writes the file with every heading present and empty. Five things
are yours:
| Field | What goes in it | Why it is asked for |
|---|---|---|
- language: |
a BCP-47 tag: en-GB, en-US, en-IN |
ships unset; a draft whose dialect nobody chose silently gets the model's own |
## Reader |
one concrete sentence -- who this is for and what they already know | every later revision is judged against it |
## Covers |
the themes the survey will address | the positive half of scope |
## Does not cover |
what it deliberately will not, including any sub-theme the corpus turned out too thin to support | so a reader can tell an omission from an oversight |
## Glossary |
each recurring term with the one definition the whole survey uses | this is what stops terminology drifting between revisions |
The genre:, draft:, created: and corpus: lines are stamped by
init. draft digest: is filled later by dossier stamp.
Set the dialect with the command rather than editing the line, so the format is right:
1 | |
en-US for most IEEE and ACM venues, en-GB for most European funders,
en-IN where that is the house style.
A filled-in survey scope.md, which you can copy and edit -- the whole
file is at
examples/dossiers/survey/scope.md:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 | |
Note what the third exclusion does: it records a corpus finding, not a preference. Six months later that sentence is the difference between "we decided not to" and "we could not", and only one of those is worth revisiting.
Check it any time with:
1 | |
🗺 Step 3: write an outline (optional, recommended)¶
You can let the skill decide the sections. You will usually get a better survey if you declare them yourself. Add the file:
1 2 | |
Before filling it in, see what the corpus actually holds, so you do not declare a section it cannot support:
1 | |
Then edit content/dossiers/dt/survey/outline.md. It has exactly three
fields, all optional per section but at least one of the first two
required:
| Field | What goes in it | What the skill does with it |
|---|---|---|
brief: |
steering, in your own words -- what to emphasise, what to skip, how long | consumed once and never appears in the draft |
claim: |
your own prose: a sentence or short block you believe is true | rewritten into the draft, and grounded -- any sentence the corpus cannot support is reported back rather than shipped |
queries: |
a - list of search terms |
run verbatim, instead of the skill inventing its own sub-themes |
Three rules worth knowing before you write one:
- A section needs at least a
brief:or aclaim:.--checkexits 1 if one has neither. queries:is optional even then. A framing or gap-analysis section usually has nothing to retrieve, and leaving it out is correct rather than lazy.- A
#level-1 line is the file's title and is passed over; sections start at##.
Declared queries bind: the skill runs yours rather than inventing
sub-themes. If a section comes up thin it may add its own, logged
distinctly, so dossier status can later tell you whether the draft ran
what you declared -- "did this draft follow my outline?" becomes a
question with an answer.
A worked example for a survey of digital-twin literature -- the whole
file is at
examples/dossiers/survey/outline.md:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 | |
Validate the shape before drafting:
1 | |
It exits 1 if a section has neither a brief: nor a claim:.
🗣 Step 4: ask for the draft¶
You do not invoke the skill by name. Ask in ordinary words, in a session whose working directory is this project:
Write a literature survey on digital-twin fidelity and validation, into
content/drafts/dt/survey.md. The dossier and outline are already there.
That phrasing -- "write a survey" / "literature review" / "related-work
section" -- is what selects survey-writer. If you name the draft path
and say the dossier exists, it will use yours rather than making a second
one.
Two things worth saying in the same breath if they matter to you: the venue or length you are aiming at, and any source you already know must be in there.
⏳ Step 5: what the skill does while you wait¶
Not a black box. In order:
- Retrieves broadly, over-fetching on purpose, per sub-theme or per
declared
queries:line. - Scores every candidate itself before it counts as evidence, and
writes both the keeps (
evidence.md) and the rejects with reasons (rejected.md). A source you can see was considered and dropped is the point of that second file. - Re-searches any sub-theme that came up thin, with reformulated queries.
- Clusters by judgement into themes, and checks for disagreement between sources before writing a word.
- Drafts in Markdown with
[@citekey]markers, a comparison table, and a gap analysis -- the part of a survey that is actually worth reading. - Never writes a citekey it did not get from a retrieval result.
You will see it working. Where the corpus is thin it says so rather than filling the hole with a plausible sentence.
✅ Step 6: gate, references, render¶
The skill runs these itself. Run them yourself after any hand edit -- this is the sequence that turns a draft into a document.
The gate is the one hard check in this pipeline:
1 | |
OK means every [@citekey] in the draft is real. FAIL names the
offending line; the fix is to correct the key or drop the claim, never to
add the key to the bib by hand.
The references section, built from exactly the gated citekeys:
1 | |
Leave the body's [@citekey] markers alone -- do not hand-number them.
Pandoc assigns [1], [2] at render time.
The renders:
1 2 3 | |
The evidence sidecar, which a survey should always emit:
1 | |
It lists each cited source with the verbatim spans that justified it,
grouped by the section that leans on them. If it prints no quoted
evidence recorded, the run captured no quotations -- that is a real
answer about the draft, not a broken command.
🔍 Step 7: read the review aids¶
None of these can block anything. All of them exit 0. They are there to be read and disagreed with.
1 2 3 4 5 6 7 8 9 10 | |
What each is for, and what a finding from it actually means:
| Aid | Reads for | A finding means |
|---|---|---|
draft style |
defect markers, an acronym never expanded at first use, dialect against scope.md |
a place to look. The first run of this check over this project's own docs kept 59 of 73 marker hits on inspection |
review verbatim |
wording shared with any parsed source, cited or not | a run of words that also appears in a source. A quoted run that cites its source is a legitimate quotation; long and short are the buckets to read |
review coverage |
how much of what retrieval surfaced actually got cited | a query whose top results the draft ignored -- sometimes correct, sometimes a theme you dropped by accident |
review synthesis |
paragraphs that summarise sources in sequence instead of synthesising them | the classic survey failure: three sentences, three citekeys, no connection drawn |
review uncited |
claims that read like they need a source and have none | a sentence making a factual claim on nobody's authority |
review quotation |
whether a quoted span matches the source it cites | a quotation that has drifted from what the paper says |
review support |
whether the cited source actually entails the claim | a citation that is real but does not carry the sentence's weight |
Add --write to any of them to file the report under
content/review/dt/, in Markdown plus JSON.
The agenda: all of them as one worklist¶
AGENDA.md explains every section of an agenda file in full; what follows is the short version for this genre.
Rather than reading seven reports, merge them:
1 | |
The agenda reads the aids' filed JSON and never runs an aid, so run
the aids you want included first (with --write), then the agenda. It
names any aid whose report is absent rather than silently omitting it.
Every item carries three things: a class, a section anchor, and
whether it is unattended or surfaced.
| Meaning | Classes | |
|---|---|---|
[unattended] |
safe for an automated pass to repair without asking | prose, the short runs a verbatim scan finds, missing-citekey |
[surfaced] |
a judgement only you can make | unsupported-claim, claim-support, uncited-claim, recorded-but-uncited, misquoted |
What the report looks like for the survey we have been building -- this is the shape, with the counts and ids a real run produces:
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 | |
How to act on it. The [surfaced] items are the ones worth your
time -- in the excerpt above, the unsupported-claim in the Gaps section
is the survey's weakest sentence and the uncited-claim is a real claim
about practice that no source in the corpus makes. Fix those by asking
for a revision (step 8).
The [unattended] ones can be handed off: ask to "work the review
agenda" and agenda-reviser repairs them one at a time, re-running the
gate and a baseline recheck after each, logging every attempt --
including refusals and reverts -- in revisions.md.
To see whether a round of edits actually helped, re-run against the previous agenda:
1 2 | |
That is the one mode that re-runs the aids. It reports each finding as
resolved, persisting, new or accepted, with the objective count
before and after -- so a repair that fixes one thing and breaks another
shows up as a flat count with a non-empty new list, rather than as
success.
If you have considered a surfaced item and decided it stands, record that instead of re-reading it every round:
1 | |
Only claim-support, uncited-claim and unsupported-claim may be
accepted; anything else is refused with exit code 2.
📝 Step 8: change something¶
Never re-run the genre skill to make a change. It would re-search the whole corpus and write a different draft. Ask for a revision instead:
Shorten section 3 and drop the two 2019 sources in favour of the newer ones.
That selects draft-reviser, which reads the dossier, edits only the
affected sections, and logs what it changed in revisions.md. It is
cheap. Re-running the genre skill is the most expensive mistake available
here (TOKENS.md).
After a hand edit of your own, re-stamp so later drift reports know what they are comparing against:
1 2 | |
If you later re-sync the corpus and want to know whether any draft now cites something that left:
1 | |
💾 Step 9: back it up¶
content/drafts/ and content/dossiers/ are gitignored on purpose --
your writing is not this repository's to commit. Bundle them yourself:
1 | |
That writes a .tar.gz holding the draft and its dossier. Add
--with-rendered to include the PDFs.
🚑 When something goes wrong¶
| What you see | What it means | What to do |
|---|---|---|
FileNotFoundError: papers/bibliography.bib |
no bib export | export your library there and re-run corpus sync |
| The ledger has entries but 0 parsed | the bib has no PDF paths | re-export with file paths attached, then corpus sync |
| The skill says the ledger is empty and refuses | no corpus yet | chitragupta corpus sync |
gate says FAIL |
a citekey is not in the corpus | correct it, or drop the claim -- never add it by hand |
[missing-binary] from render |
no pandoc/pdflatex |
install them; the .md draft is unaffected |
| The survey reads thin in one theme | the corpus is thin there | add papers to the bib, corpus sync, then ask corpus-reviser for a whole-corpus pass |
| You want a change and are tempted to re-run the skill | -- | ask for a revision instead; see step 8 |