ADR 001 — The nightly build runs off-box; only artifacts ship¶
Status: Accepted · Date: 2026-08-08 · Deviates from: spec §3, §5.7, §11.1, §11.2
Context¶
Should the data build run on the machine that serves the application?
The original design said yes: ingest → dbt → embeddings as a cron-triggered container alongside the web app (spec §11.1), with the atomic swap happening locally (§5.7). It is the simpler arrangement, and for many projects it is the right one.
It is the wrong one here, for reasons that generalise beyond this project:
- A build and a server want opposite things. A web server wants steady, low, predictable resource use. A data build wants to consume everything available for a few minutes and then stop. Putting them on one host means the build's peak sets the server's floor.
- The heavy step is rarely the one you plan for. Here it is not the database — that is under 100 MB. It is bulk ingestion churning tens of megabytes of compressed archives per source, and an embeddings toolchain that pulls roughly 2 GB of dependencies to produce one column.
- Blast radius. Anything sharing a host with a periodic heavy job inherits its worst minute. Where that host also serves unrelated things, a build failure stops being a build problem.
Decision¶
The build runs on the developer's machine, not the server. make all produces a dated
database file, the rendered semantic artifacts, the documentation site and the favicon
cache. A publish script ships them, verifies a checksum at the far end, swaps a symlink
atomically, and healthchecks. The server runs one application container and a backup
job — nothing that ingests, transforms or embeds.
A note on what this document does and does not contain. These ADRs exist to share the reasoning — the trade-offs, the things that turned out to be wrong, the constraints that actually bit. They deliberately do not describe the hosting estate: no hostnames, addresses, providers, neighbouring services or utilisation figures. None of that would make the argument clearer, and publishing it hands anyone reading a direct route to the origin, around the CDN and the protections in front of it. The architecture is the public artifact; the infrastructure is not.
Rationale, and a correction¶
The original argument had two legs — disk pressure and CPU/RAM contention. The disk leg turned out to be wrong and is withdrawn. The S0 feasibility spike measured the full World Bank WDI pull (411 Nepal-covered series × 14 geos × 1960–2025, 198,969 observations) at 1.2 MB parquet / 2.6 MB DuckDB. Extrapolated to global scope across all Phase-1 sources, the database is roughly 200 MB — not the multi-GB artifact the spec's MotherDuck 10 GB threshold implies. Shipping it nightly is trivially cheap.
The decision stands on the remaining leg, which the spike did not undermine:
- Ingestion is the heavy step, not storage. FAOSTAT's REST API is down (HTTP 521), so the connector must use bulk ZIPs — 69 datasets, several 50–70 MB each, unzipped and parsed. That is sustained disk and CPU churn on a box with 26 GB free.
- Embeddings pull a large toolchain.
sentence-transformersplus torch is ~2 GB of dependencies and is CPU-bound at build time. Installing that alongside 19 running containers buys nothing — the output is aFLOAT[384]column that ships inside the.db. - Blast radius. A runaway build on a shared box takes down six other people's projects. Off-box,
the worst case is a stale database and a failed publish, and the previous
.dbkeeps serving.
Consequences¶
make allmust be run somewhere deliberate. This is a real cost: the pipeline is no longer self-driving, and a forgotten build shows up as a stale freshness badge./healthzreportsdb_age_hoursand returns 503 beyond 48h, so staleness is detected rather than silently served.- Rollback is
ln -sfnto the previous.db— cheaper than the spec's in-place design. - CI can take over the build later without re-architecting; the seam is
ops/publish.sh.