Skip to content

ADR 004 — The indicator crosswalk has no dimension key

Status: Accepted (documents a known limitation) · Date: 2026-08-09 Affects: spec §5.3, §6.1 · Discovered while: crosswalking ILOSTAT, IMF, OWID, UN SDG and HDX

Context

conformed.fact_observation resolves a raw source code to a canonical indicator by joining the generated crosswalk on (source_id, code) — and nothing else.

That is sufficient for a source that publishes one series per code. The World Bank does: SP.DYN.TFRT.IN is the total fertility rate and nothing else.

Several sources do not. ILOSTAT, UN SDG and UNESCO UIS publish many series under one code, distinguished only by dimensions — sex, age band, geographic location, education level — that arrive in obs_attrs rather than in the code itself. Under a single UN SDG code you may find the total, the female figure, the male figure, the urban figure and the rural figure.

Because the join ignores obs_attrs, wiring such a code maps every one of those series onto the same canonical indicator. The result is not an error. It is a chart with several overlapping lines that all claim to be the same measure, and an aggregate that silently mixes a national total with a female sub-population.

Decision

Crosswalk only codes that carry a single series for a given geography and period. Where a source multiplexes dimensions under one code, do not map it until the crosswalk can express the dimension.

This is the direct reason source coverage is narrower than ingestion:

Source Rows ingested Rows crosswalked Why
ILOSTAT 1,271,592 limited most codes carry sex/age/status dimensions
UNESCO UIS 64,351 partial education-level and sex dimensions
UN SDG 14,071 49 most series carry Sex/Age/Location
HDX 15,054 0 subnational detail collapsed onto one national geo_id

Those rows remain ingested and queryable; they are simply not governed. That is the honest position, and the sources table on the homepage states it rather than implying nine governed sources when there are fewer.

Consequences

  • Coverage is capped by the join, not by curation effort. Adding more indicator YAML does not unlock ILOSTAT's 1.27 million rows; changing the key does.
  • The fix, when it is worth doing: add an optional dimension filter to the crosswalk — (source_id, code, dimension_predicate) — and have fact_observation apply it, so a mapping can say "the total series only". That is a schema change to a generated seed plus a join change, and it should be measured against the reconciliation gain before being built.
  • Until then, every new mapping must be checked for multiplexing. Comparing the row count for one geography and year against the expected 1 is enough: anything higher means the code carries dimensions.

Why this is recorded rather than fixed

The blend is silent. It produces plausible numbers, passes every not-null and uniqueness test, and only shows up as a chart that looks subtly wrong — which is the failure mode this project has hit repeatedly and now actively hunts for. Writing the limitation down, with the reason coverage stops where it does, is more useful than a partial fix that makes the boundary harder to see.