๐ Developer guide¶
Material for working on this repository itself, as opposed to using it to draft content -- test running, the full source layout, and known gaps. See README.md for the user-facing Quickstart/Configuration/ Architecture docs and DOCKER.md for running a container; DOCKER-DEVELOPER.md here is this repo's own record of how the two images are built and verified.
๐งญ Table of contents¶
- Running tests
- Benchmarking the parser
- Writing a script that drives the enrichment layer
- Repository layout
- Figures and copyright
- Citation provenance
- Open questions and unbuilt features
๐งช Running tests¶
1 2 3 4 5 6 7 8 9 10 | |
tests/ covers both the corpus layer and chitragupta/enrich/* -- the enrich group's
dependencies (docling, chromadb, bertopic,
sentence-transformers) are mocked via sys.modules for fast,
deterministic unit tests, so the
dev-deps group alone is not enough on its own: the enrich group
(python-deps, step 1 of Quickstart) must already be installed too, since
tests/test_bib_reader.py needs bibtexparser and the chitragupta/enrich/ test
modules need docling/chromadb/bertopic/sentence-transformers.
One module is the deliberate exception.
tests/test_enrich_real_libraries.py drives the real chromadb through
build_index()/search(), and asks the real
sentence_transformers/bertopic classes whether they still accept the
keywords the fakes accept (#514). The fakes are faithful enough that the
expensive failure is the day they quietly stop being; that module is what
notices. It does not download an embedding model -- its own docstring has
why -- so it stays as fast and as offline as the rest.
A handful of tests run real dependencies end to end rather than mocking
them, and skip automatically when the dependency is absent:
tests/test_feature_workflows.py and the TestRenderReal/
TestExtractTextReal classes elsewhere probe for the
pdftotext/pandoc/pdflatex binaries on PATH;
tests/test_enrich_real_libraries.py probes for an importable library
instead, which is the same idiom against a different kind of absence.
โก Benchmarking the parser¶
bench/ measures what a full docling parse of the bib corpus costs on
a given machine, and is deliberately kept out of tests/: it takes a couple of
hours, needs real PDFs and a GPU, and answers a "how long / what's the
bottleneck" question rather than a pass/fail one. It is excluded from the
release zip for the same reason tests/ is.
- docs/PERFORMANCE.md -- what each setting costs,
organised by setting. Ships in the release archive, unlike
bench/ - docs/PARALLELISM.md -- parallel parse design: architecture, components, and the roadmap
- bench/README.md -- how to run it, and what each switch measures
- bench/RESULTS.md
-- the dated measurement record,
newest last, with raw per-run data in
bench/results/. Read its "Which sections are current" table first: several early conclusions were overturned by later runs and are kept, marked, rather than deleted - bench/PARALLELISM-PLAN.md -- what is still unknown, and what to measure before changing it
The headline, in the order it was found:
- Parsing all 501 bib PDFs with
doclingtook ~1.6 hours (later measured at 1h 56m), with the A40 at ~7% utilization and three CPU cores of 48 busy. The GPU was worth only 1.79x over CPU-only -- the work was CPU-bound. - Turning OCR off (v0.12.0) was worth more than the GPU: 2.08x serially, 3.91x at 12 workers and 4.79x at 24, since OCR competes for the same CPU the parallelism needs. (An earlier 2.46x, from a 16-PDF serial sample, is still quoted in older text; it estimated the serial case only.)
- Parallelising
sync(v1.0.0) was worth 3.60x at four workers. - That moved the bottleneck onto a single GPU:
AcceleratorDevice.AUTOresolves tocuda:0in every worker, so GPU 0 ran at 100% while GPUs 1-3 idled. - Giving each worker its own card (v1.1.0) was worth a further 1.62x on the full corpus -- 528s to 326s at twelve workers. The whole 501-PDF corpus now parses in 5m 10s, against 1h 56m where this started.
- Per-worker startup (v2.1.0) turned out to be 3.2s of importing torch and docling plus ~5s of loading Docling's models, and only the first is shareable between processes. A forkserver pool with those modules preloaded, started before the bibliography is read, takes a fixed ~1.5-2s off pool startup -- 9.6% of an 8-document run, 2.5% of a 60-document one.
- Measuring the whole corpus instead of extrapolating from a 16-PDF
sample (2026-08-04) found the serial baseline was 55m 30s, not the
~39m every document had quoted -- 41% low. Correcting it showed
12-worker efficiency is 89%, not the 60% previously reported, and that
worker_ceiling()'scpus // 4clamp costs 1.41x: 32 workers beat the 12 it allows. - Asking whether a quotable passage survives a re-parse (2026-08-07,
bench/repro_check.py) found that ~1% of documents come back with a different passage text, and -- correcting what this project had asserted twice -- that two runs of the same configuration are not exempt either. The artifact-by-artifact contract that came out of it is in docs/ARCHITECTURE.md.
The lesson worth carrying: every one of those steps was measured, and
seven intermediate conclusions were wrong until the next measurement
corrected them -- including two that sat in the code as stated fact, and
one that had been written into three documents. bench/ exists so that
the next one is checked too.
๐ง Writing a script that drives the enrichment layer¶
chitragupta.enrich.docling_parse.parse_corpus and python -m
chitragupta.corpus sync both use
a worker pool when [parser].workers is above 1, and every start method
they can pick (forkserver or spawn -- see [parser].start_method)
re-imports the calling program's __main__ in each worker. Any script of
your own that calls them must guard its top level:
1 2 | |
Without it, every worker re-runs the script on startup and the pool dies
with BrokenProcessPool. chitragupta/enrich/__main__.py and chitragupta/sync.py
are both guarded already; this only bites ad-hoc scripts, and it bites
immediately rather than subtly.
๐ Repository layout¶
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 | |
๐ Figures and copyright¶
With [enrich].docling_images on (off by default), the Docling stage writes
each paper's figure bitmaps to content/docling/<doc>_artifacts/ and an
index of them to content/docling/<doc>.figures.json.
Those images are a reading aid, not draft content. Nothing in this
repo inserts them into content/drafts/, and nothing should start doing
so. A figure's copyright belongs to the publisher or the authors, and
citing a paper grants no right to reproduce its figures -- citation_gate
gates citekeys, and there is deliberately no equivalent gate for
images. The ledger also has no license column, so the pipeline genuinely
cannot tell a CC BY paper from an all-rights-reserved one; that judgment
stays with you, per figure.
The supported way to reference a figure is therefore textually, and
each record in <doc>.figures.json carries a ready-to-paste cite
string:
1 2 3 4 5 6 | |
Two details worth knowing about that cite string:
- The number comes from the caption's own text, never from the picture's position. Publisher logos and licence badges are pictures too -- on a real 17-page MDPI paper, 6 of the 13 extracted pictures were furniture rather than figures -- so the Nth picture is routinely not the paper's Figure N.
- The number is captured whole, including chapter-scoped forms
(
Fig. 1.1...Fig. 1.4, the convention in edited book chapters) and sub-figure letters (Figure 2a). Matching only the leading integer would collapse a chapter's four distinct figures onto oneFigure 1-- a citation pointing at the wrong picture. - A picture whose caption carries no number is cited by page instead
(
"the figure on p.1 of [@key]"), rather than being given a number this repo would have to invent. Two panels of one figure (captions beginning(a)/(b)) therefore share a page-based citation; that is the fallback behaving correctly, not a collision.
Every figure's cite string is a real [@citekey], because every
document the enrichment layer parses comes from the bib file (see
chitragupta/enrich/corpus.py).
A wholly original diagram -- not derived from any source paper's figure -- is a different case this section doesn't restrict; see docs/WRITING-STANDARDS.md ยง10 for the supported form and why it's plain ASCII rather than Unicode box-drawing.
๐ Citation provenance¶
python -m chitragupta.review provenance content/drafts/<slug>.md reports, for
every citation in a draft, what in the cited source supports it and where
-- ordered worst match first. It writes
content/review/<the draft's path minus its suffix>.provenance.md plus
.tex/.pdf renders beside it. The report mirrors the draft's own place
under content/drafts/, the same rule rendered/ and dossiers/ follow
(config.mirrored_dir), and chitragupta/review/__init__.py owns that contract
for all
six review-layer commands.
Run it directly rather than wrapping it in an enrichment stage: that
would have the enrichment layer importing the review layer, and would
make an advisory report wait on sync's write lock.
Advisory, not a gate, deliberately: matching is lexical, so it
cannot tell "the source doesn't say this" from "the source says it in
words I didn't recognise". citation_gate blocks because it checks
something exact (ledger membership); this reports because it doesn't.
Passage quality depends on what has been parsed. With the Docling stage
run, content/docling/<citekey>.passages.json supplies reading-ordered
paragraphs and the report quotes them. Without it, pdftotext output is
used and the report gives a page number without quoting -- on a
two-column paper that text splices two columns onto every line, so any
excerpt would be a collage of two arguments.
Full design rationale, including the measurements behind those choices: docs/CITATION-PROVENANCE.md.
โ Open questions and unbuilt features¶
Running this pipeline on a schedule was the long-standing goal here.
Most of it now exists: a rotating logs/pipeline.log, shared by the
corpus and enrichment layers (see chitragupta/logging_setup.py), a pages/s
throughput figure, exit codes an unattended caller can branch
on, and worked cron and systemd units in
docs/CLI.md -- including the
absolute-interpreter-path detail that cron's minimal environment
requires.
One blocker is left, and it is not a coding task. With no continuous
auto-export, bibliography.bib is a manual, point-in-time snapshot: a
scheduled sync re-reads whatever was last exported, so it keeps the
corpus consistent with the bib file but cannot keep the bib file
consistent with your reference manager. A schedule watching only its
mtime does nothing until a human re-exports. Closing that properly means
either a Zotero auto-export plugin (outside this repo) or accepting that
the export stays a deliberate human step.
๐ซ content/topics.json has no consumer¶
chitragupta/enrich/topic_model.py writes it and nothing reads it -- no module,
no
genre skill. survey-writer groups themes by judgement and says so
explicitly ("With a small corpus there's no BERTopic step"). That is
defensible today: clustering is whole-corpus, so assignments are not
stable between runs, and on a small corpus every document legitimately
lands in the outlier topic. If it is ever wired in, survey-writer's
"Cluster by judgment" step is the seam, gated on the file existing and on
there being non--1 assignments -- the same shape the existing skills use
to gate on content/chroma/.