🐳 Running with Docker¶
Two images, for two different jobs. Neither is a variant of the other and neither replaces the other:
| File | What it is | Size | Needs a checkout |
|---|---|---|---|
docker/Dockerfile |
The toolchain image. Builds the full TeX Live/Pandoc/Poetry/torch stack from this repository's own scripts/install_full_pipeline.sh, for hosts where you don't hold root. Ubuntu 26.04 LTS / Python 3.14, with the C headers and compiler triton JITs against (#668). Mount your checkout, get a shell with everything resolved. |
5.09GB (cpu) / 11.6GB (gpu, unmeasured since the base bump) | Yes -- it is the checkout you mount and work in |
docker/Dockerfile.claude |
The agent image. Node, Claude Code, and a /opt/venv install of the published chitragupta-cli package, driven by docker/docker-compose.yml as a long-lived claude remote-control host. The toolchain is not baked in; you add what you need at runtime with chitragupta install. |
~2.5GB | No -- it installs from PyPI |
docker/Dockerfile is documented first, then the agent image under
"The agent container".
There's nothing Docker-exclusive about any individual piece of the
toolchain image -- scripts/install_full_pipeline.sh is the single
install path for both the host and it.
Getting docker/ without a checkout. chitragupta init doesn't
scaffold docker/, and a pip install chitragupta-cli doesn't carry it
either. Every tagged GitHub
Release attaches
a standalone chitragupta-docker-<version>.zip -- docker/ plus this
page -- for exactly that case. That's enough on its own for the agent
image: Dockerfile.claude's build context is docker/ alone (see
"Build and run" below) and it installs from PyPI, so
the commands under "The agent container"
work with nothing else present. The toolchain image still needs a
checkout regardless of how you got docker/ -- the table above already
says so -- because its Dockerfile COPYs scripts/install_full_pipeline.sh
and its Run section mounts your own working copy for config.toml to
resolve from. DOCKER-DEVELOPER.md at the repository root is this
project's own record of how these two images are verified before a
release, not something a Docker user needs to read.
🔧 Build¶
1 | |
This runs scripts/install_full_pipeline.sh twice as separate,
independently cached layers -- os-deps, then python-deps (via
Poetry, with SKIP_VENV=1 so it installs into /opt/venv instead of
creating its own) -- so editing later Dockerfile lines or unrelated repo
files doesn't force earlier layers to rebuild.
Exception: the script itself is COPY'd once, before either of the
two stages runs, so editing scripts/install_full_pipeline.sh
invalidates both layers --
Docker's cache keys each layer on the exact command and any files that
command's COPY depends on, and this file feeds both of them. The
python-deps layer pulls torch and Docling's models; expect a long
first build.
⚖ CPU-only vs. GPU-capable: TORCH_VARIANT¶
torch isn't pinned in pyproject.toml (it's a transitive dependency of
sentence-transformers/docling/accelerate; see that file's own
comment) -- poetry.lock resolves whatever plain PyPI's default Linux
wheel currently is, and that wheel bundles the full CUDA runtime as
separate nvidia-* packages (cublas, cudnn, nccl, triton, ...)
regardless of whether the build host or the eventual container ever sees
a GPU. The docker/Dockerfile build arg TORCH_VARIANT controls which
one you get:
TORCH_VARIANT |
Build command | Measured image size | When to use it |
|---|---|---|---|
gpu (default) |
docker build -t chitragupta -f docker/Dockerfile . |
11.6GB | docker run --gpus deployments -- the bundled CUDA runtime is enough on its own, no host CUDA toolkit needed, only a matching driver |
cpu |
docker build -t chitragupta -f docker/Dockerfile --build-arg TORCH_VARIANT=cpu . |
5.09GB | Everything else -- embeddings/clustering/rendering all run fine on CPU, and this is what you want for build-verification or a host with no GPU at all |
The cpu variant reinstalls torch/torchvision from PyTorch's own
CPU-only wheel index, at the exact version poetry.lock resolved (read
back via pip show rather than pinned a second time in the Dockerfile,
so it can't drift from a poetry lock re-resolution), then removes the
now-orphaned nvidia-*/triton packages. Both the swap and the removal
happen inside the same RUN as the original poetry install, so the
CUDA wheels never end up committed to a layer in the first place --
doing this as a later, separate RUN only marks them deleted and leaves
the image larger, since the earlier layer's bytes are still stored.
Sizes above also depend on the Poetry/pip download cache being purged in
that same layer (rm -rf /root/.cache/pypoetry /root/.cache/pip, at the
end of the python-deps RUN) -- without it, every build leaves the
full set of downloaded wheel archives sitting in the final image on top
of the installed packages (measured: 24.9GB for the gpu variant with
the cache left in place, more than double its 11.6GB with it purged).
🚀 Run¶
Mount your repo and a volume for content/ so it survives container
restarts:
1 2 3 4 | |
The image deliberately doesn't bake the repo in -- it mounts it -- so the
config.toml the container reads is the one in your working copy. That
file is gitignored, so create it before the first run or chitragupta.config
will refuse to import:
1 | |
✅ Verify the toolchain¶
Inside the running container, check that the render and enrichment dependencies actually resolved:
1 2 | |
🤖 The agent container (docker/Dockerfile.claude)¶
A second, much thinner image whose job is to host one long-lived
claude remote-control session with this pipeline installed, rather
than to build the toolchain. It installs the published
chitragupta-cli package from PyPI into /opt/venv and copies no part of
this repository in, which is the whole difference: an image built from a
checkout has to be rebuilt to follow a release, and this one only has to
be restarted.
What that buys, and what it costs: the build is seconds rather than
tens of minutes and the image is ~2.5GB rather than 5.09GB, because
TeX Live, Pandoc and torch are not in it. You add whatever the work
actually needs on first run, with chitragupta install -- which is why
the image gives its user passwordless sudo (see below).
The shell you land in is zsh with Oh My Zsh, pinned to a commit
rather than installed by piping an unpinned script from the network, and
neovim is there for the config you inevitably want to fix by hand. The
venv reaches zsh through /etc/zsh/zshenv, not /etc/profile.d --
Debian's zsh does not read the latter, and getting that wrong brings
back the failure the next section is about.
🧩 One environment, and why it is not pipx¶
chitragupta-cli goes into a venv at /opt/venv which is first on
PATH -- the same arrangement docker/Dockerfile uses -- so python,
python3, pip, chitragupta and cg are all the same environment.
That is not a style preference. .claude/settings.json, the one
chitragupta init scaffolds, launches every hook as command:
"python". Install the package into a private venv, as pipx does,
and that python is the system interpreter, which cannot import it --
so the citation gate returns
1 2 | |
on every draft write. This image was built with pipx first and
that is exactly what happened: a container that came up looking healthy
and refused to let anything be drafted in it. The gate is right to
block — it will not pass a draft it could not check — which is why the
fix belongs here rather than in the hook.
A /etc/profile.d/ snippet carries the same PATH for login
shells, which source /etc/profile and rebuild PATH from scratch,
discarding what ENV PATH set. Worth knowing that the pipx version
appeared to survive a login shell only because Debian's ~/.profile
re-adds $HOME/.local/bin, where pipx's shims happen to live -- luck,
not design.
🔧 Build and run¶
Build with a temporary tag while you are changing anything here, so
an existing :latest that other containers are running survives a
failed attempt:
1 | |
Note the build context is docker/, not the repository root -- nothing
outside that directory is needed, and a root context would send the
whole tree (worktrees included) to the daemon.
Then run it through compose, from inside docker/ so its relative
defaults and its .env resolve:
1 2 3 4 5 | |
docker/.env.example is the tracked documentation of every variable
the compose file reads; docker/.env is gitignored. Nothing has to be
edited but CHITRAGUPTA_WORKSPACE.
That mkdir is not optional politeness. The Docker daemon creates a
missing bind-mount source itself, owned by root:root -- and for the
read-only papers mount the result is a directory nobody can put a PDF
in: read-only inside the container, root-owned outside it, silently.
Create both directories first and they stay yours.
🎛 Two profiles: cpu and gpu¶
Neither service starts without a profile, on purpose. docker
compose up with none selected prints no service selected and exits
having done nothing:
1 2 | |
or set COMPOSE_PROFILES in .env (the example ships cpu) and plain
docker compose up -d works.
The two services are identical apart from a deploy: reservation, and
they share the YAML anchor that says so, so they cannot drift. They also
share a container name, since only one runs at a time.
Why a profile rather than one service that adapts: a reservation naming
the nvidia driver makes up fail outright on a host without the
NVIDIA Container Toolkit -- could not select device driver -- rather
than starting without a GPU. Compose has no conditional for a volume or
a reservation, so two services under two profiles is the only shape that
lets one file serve both machines. Verified both ways on a host with the
toolkit: --profile gpu reaches all three A40s and nvidia-smi inside
the container, --profile cpu shows no /dev/nvidia* at all.
⚙ What compose reads from the environment¶
Every path and name is a variable. The first version of this file
carried one developer's home directory and a release number in its own
filename, which made it unusable to anyone else and stale on the next
release. Set these in docker/.env or in the environment:
| Variable | Default | What it does |
|---|---|---|
CHITRAGUPTA_WORKSPACE |
required | Host directory mounted at /workspace -- where chitragupta init scaffolds and where drafts land. Compose refuses to start without it rather than mounting something arbitrary |
CHITRAGUPTA_CONTAINER_NAME |
chitragupta-claude |
The container's name, so docker exec -it <name> bash works. Also becomes the container hostname and the agent's name in the remote-control UI |
CHITRAGUPTA_PROJECT_NAME |
chitragupta |
Compose's project name, which prefixes the network. Set it per workspace to run two of these side by side |
CHITRAGUPTA_IMAGE_TAG |
latest |
Which tag to run. This is how you point one workspace at a :tmp build without disturbing the :latest everything else uses |
CHITRAGUPTA_CLAUDE_HOME |
./claude |
Host directory for /home/prasad/.claude. This is what persists the login, so a restart is not a re-authentication |
CHITRAGUPTA_PAPERS |
$CHITRAGUPTA_WORKSPACE/papers |
Host directory of PDFs, mounted read-only at /workspace/papers. Point it elsewhere to share one corpus across workspaces |
COMPOSE_PROFILES |
none | cpu or gpu -- see below. Compose's own variable, not one of ours |
CHITRAGUPTA_USER |
required | The unprivileged account inside the image -- your own login name (id -un). A build arg as well as a runtime setting, so changing it needs docker compose build, not just a restart |
CHITRAGUPTA_UID / CHITRAGUPTA_GID |
required | That account's numeric ids -- your own (id -u, id -g). They decide who owns a file the container writes into the workspace. Avoid 1000: the Node base image has already taken it for its own node account |
Nothing in the image hardcodes an account name, and the three account
settings have no defaults anywhere -- not in docker-compose.yml
and not in Dockerfile.claude. docker compose refuses to start
without them, naming the variable and the file; a bare docker build
that forgets the matching --build-arg fails on an explicit guard
rather than creating an account called "".
That is deliberate. The account belongs to your machine, not to this
project, and a default that is wrong for almost everyone is worse than
no default at all: it works just well enough to write files owned by the
wrong uid into your workspace, which you then cannot edit from outside
the container. Copy .env.example to .env and fill in the three
values it shows you how to read off your own account.
One variable does the whole job: CHITRAGUPTA_USER reaches the
useradd, the /etc/sudoers.d/<user> file, $HOME, the PATH entry
the venv installs into, and the mount point for
CHITRAGUPTA_CLAUDE_HOME -- five places that cannot disagree because
they read one value. The entrypoint lives at
/usr/local/bin/entrypoint.sh rather than in the home directory for the
same reason.
🚀 First run¶
The container starts before it can do anything useful, deliberately, and in this order:
1 2 3 4 5 6 | |
chitragupta install os-deps is what the passwordless sudo in this
image exists for: it apt-get installs the toolchain that the image
does not ship. (chitragupta install still refuses all, dev-deps
and python-deps by name, naming what reaches them instead -- those are
checkout-shaped stages with no analogue here.)
The enrichment stack is its own step rather than part of the image, because it is torch and would take this from ~2.5GB past 6GB:
1 | |
That is the CLI doing it, not a pip incantation you have to know: it
resolves 'chitragupta-cli[enrich]' pinned to the version already
running, so adding the extra cannot quietly upgrade the tool.
chitragupta install enrich needs 6.76.0 or newer, and this image
installs whatever PyPI serves -- so an image built from an older release
will not have the verb yet and chitragupta install --help will not
list it. pip install 'chitragupta-cli[enrich]' is the equivalent
there, and still works everywhere.
✅ Verify the container¶
1 2 3 4 5 6 7 | |
🔍 When the agent is not there¶
The container can report Up with no agent running in it, and the
common cause is the first-run one: claude remote-control exits
immediately if the mounted .claude carries no login. entrypoint.sh
reports that in docker logs rather than leaving it silent:
1 2 3 4 5 | |
It reports rather than exiting non-zero on purpose: the fix needs a
running container to docker exec into, and a crash loop would take
that away. docker logs <name> is therefore the first thing to read
when the remote-control UI does not list your agent -- not docker ps,
which will say Up either way.
Attach to the live session with:
1 | |
⌨ Running pipeline commands inside the container¶
The same commands as the main README's Quickstart work directly with no
venv prefix, since /opt/venv is already on PATH (and exported as
VIRTUAL_ENV, so Poetry installs into it rather than creating its own)
inside the container:
1 2 3 | |