Skip to content

Architecture

This page describes how Groundfact turns a dozen disagreeing public statistical sources into one cited answer, and — as much as the shape of the pipeline — what it deliberately does not do. It does not describe where any of it runs: no hostnames, addresses, providers, regions, neighbouring services or utilisation figures appear below. ADR 001 explains why in more detail; the short version is that the architecture is the public artifact, the infrastructure is not, and naming the latter would only help someone route around the protections in front of it.

The pipeline

public APIs → dated snapshots → conformed warehouse → semantic layer → cited answers

Public APIs → dated snapshots. Every ingestion run writes a timestamped, immutable Parquet file per source and never modifies it afterwards. Nothing downstream reads a live API — every later stage reads a snapshot that was true at a specific moment and stays true forever. That is also the revision mechanism: a number cited from a given date is reproducible from the file that produced it, not from whatever the upstream API happens to return today.

Dated snapshots → conformed warehouse. Each source's raw rows are staged and then conformed against two canonical keys: a geography id and, where the indicator is curated, an indicator id. The same country arrives under half a dozen different codes across nine sources — a crosswalk resolves all of them to one geo_id, and a row whose geography cannot be resolved is dropped rather than kept with a null key. An unresolved join is a wrong answer waiting to happen, not a row worth keeping "just in case."

Conformed warehouse → semantic layer. The warehouse on its own is just a fact table with dimensions. The semantic layer is a curated subset of it — one YAML record per indicator, each naming its unit, its plausible range, its preferred source among the ones that carry it, its aliases, its taxonomy placement — checked into the same repository as the code. That YAML is the single upstream source of truth for everything a reader or a model is allowed to know exists: the documentation you're reading, the compact grounding document a language model is given, and the warehouse's own indicator dimension are all rendered from the same commit by the same build step. A CI check re-renders all of it and fails if the output differs from what's committed, so the docs and the model's view of the world cannot silently drift apart from the data or from each other.

Semantic layer → cited answers. A question in plain English is resolved into a small, validated query specification — which indicators, which geographies, which time range, which transform — never into SQL. Deterministic code turns that specification into a parameterised query against the warehouse. A second pass writes commentary grounded in the actual rows returned and proposes a chart shape; the server re-binds the real values before rendering anything, so what reaches the page is never something the model produced unchecked. Every number on screen carries a citation — source, dataset, snapshot date, licence — back to where it came from.

Reconciliation is surfaced, not hidden. Where two sources report the same indicator for the same place and period and disagree, the difference is computed and shown rather than resolved by picking a winner silently. A modelled estimate and a survey measurement are both real information; showing the gap is more honest than averaging it away.

Why the model never writes SQL

The tempting design is to let a language model write SQL directly against the warehouse and trust it to write a correct query. That trust is exactly what Groundfact's design refuses to extend, for a specific reason: a wrong query that still parses is indistinguishable, from the model's own output, from a right one. Nothing downstream would catch it.

The query specification pattern removes that trust requirement instead of trying to harden it. Every field the model can set is closed: an indicator id that must already exist in the semantic layer, a geography id from a fixed set, a time range, a small enum of transforms (compare, index, growth rate, and so on). There is no field for "the rest of the query" — no free-text filter, no arbitrary join, nothing that could smuggle intent past validation. Deterministic code, written and tested like any other code path, turns that specification into the actual parameterised SQL. The model is therefore not merely discouraged from inventing a number or a join condition — it is never in a position to supply one. What it chooses is which indicators, geographies and time window answer the question, and how to describe and chart the numbers the query layer independently returned. It cannot express a query the semantic layer doesn't already know how to answer safely, which means there is no version of "the model got creative" that reaches the database.

What I deliberately didn't build

No Postgres. The whole warehouse is comfortably under the size where an embedded analytical engine reading its own file stops making sense; there is no server process to run, patch, or lose sleep over, and no write concurrency to arbitrate because there is exactly one writer (the nightly build) and many read-only readers. A client-server database earns its operational cost at a scale this project isn't at.

No Kubernetes. One application process reads one immutable file and serves read-only traffic; that is what a genuinely single-node deployment looks like when it isn't apologising for being one. Orchestrating a fleet to serve a workload that fits comfortably on a single machine would be solving a scaling problem nobody has yet, at the cost of a much larger surface to operate.

No vector database. Indicator and topic search runs over a few hundred rows, not millions. A brute-force cosine scan over an embedding column computed once at build time answers in milliseconds at this size — a dedicated vector store would be infrastructure standing in for a loop. The seam to swap it in later, if the catalogue ever reaches a size where that stops being true, is deliberately left open rather than built ahead of the need.

No orchestrator. There is exactly one scheduled job: the nightly build, triggered by cron and shipped by a script that verifies a checksum, swaps a symlink atomically, and healthchecks before and after. A workflow engine coordinates many interdependent tasks with retries and branching; there is one task, and it either succeeds or it doesn't, in which case the previous build keeps serving. Reaching for a scheduler that manages dependency graphs, for a graph with one node, would add a system to operate in exchange for nothing a cron entry doesn't already do.

No SPA framework. Every page is rendered server-side from data the server already computed; there is no client-side application state to synchronise because there is no client-side application. A small amount of interactivity — a streaming chat response, a chip tap that fills a question in — is real, and is handled with a lightweight hypermedia approach layered onto plain HTML rather than by shipping a virtual DOM and a build toolchain to manage state the server was always going to compute anyway.

No free-form text-to-SQL. This is the one on this list with a genuine safety argument behind it rather than just an operational-cost one, and it's covered above: letting a model write arbitrary SQL means trusting it not to write a wrong query that still executes. The validated query-specification layer exists specifically so that trust is never required.

Further reading

  • Sources — the public licence register: every source, what Groundfact takes from it, its licence, and whether the data is re-hosted or only ever linked to at origin.
  • Indicator reference — the curated indicators, generated from the same semantic layer.
  • Architecture decisions — the specific deviations from the original design, with the reasoning and, where it applies, a correction of an earlier assumption that turned out to be wrong.