ADR 004 — The indicator crosswalk has no dimension key¶
Status: Accepted (documents a known limitation) · Date: 2026-08-09 Affects: spec §5.3, §6.1 · Discovered while: crosswalking ILOSTAT, IMF, OWID, UN SDG and HDX
Context¶
conformed.fact_observation resolves a raw source code to a canonical indicator by
joining the generated crosswalk on (source_id, code) — and nothing else.
That is sufficient for a source that publishes one series per code. The World Bank does:
SP.DYN.TFRT.IN is the total fertility rate and nothing else.
Several sources do not. ILOSTAT, UN SDG and UNESCO UIS publish many series under one
code, distinguished only by dimensions — sex, age band, geographic location, education
level — that arrive in obs_attrs rather than in the code itself. Under a single UN SDG
code you may find the total, the female figure, the male figure, the urban figure and the
rural figure.
Because the join ignores obs_attrs, wiring such a code maps every one of those series
onto the same canonical indicator. The result is not an error. It is a chart with several
overlapping lines that all claim to be the same measure, and an aggregate that silently
mixes a national total with a female sub-population.
Decision¶
Crosswalk only codes that carry a single series for a given geography and period. Where a source multiplexes dimensions under one code, do not map it until the crosswalk can express the dimension.
This is the direct reason source coverage is narrower than ingestion:
| Source | Rows ingested | Rows crosswalked | Why |
|---|---|---|---|
| ILOSTAT | 1,271,592 | limited | most codes carry sex/age/status dimensions |
| UNESCO UIS | 64,351 | partial | education-level and sex dimensions |
| UN SDG | 14,071 | 49 | most series carry Sex/Age/Location |
| HDX | 15,054 | 0 | subnational detail collapsed onto one national geo_id |
Those rows remain ingested and queryable; they are simply not governed. That is the honest position, and the sources table on the homepage states it rather than implying nine governed sources when there are fewer.
Consequences¶
- Coverage is capped by the join, not by curation effort. Adding more indicator YAML does not unlock ILOSTAT's 1.27 million rows; changing the key does.
- The fix, when it is worth doing: add an optional dimension filter to the crosswalk —
(source_id, code, dimension_predicate)— and havefact_observationapply it, so a mapping can say "the total series only". That is a schema change to a generated seed plus a join change, and it should be measured against the reconciliation gain before being built. - Until then, every new mapping must be checked for multiplexing. Comparing the row count for one geography and year against the expected 1 is enough: anything higher means the code carries dimensions.
Why this is recorded rather than fixed¶
The blend is silent. It produces plausible numbers, passes every not-null and uniqueness test, and only shows up as a chart that looks subtly wrong — which is the failure mode this project has hit repeatedly and now actively hunts for. Writing the limitation down, with the reason coverage stops where it does, is more useful than a partial fix that makes the boundary harder to see.