The Codex normalization standard.
One contract for every catalog dataset. The Codex normalization standard shapes each record so that any Codex dataset feels like part of a single corpus rather than a loosely-related dump.
Seven principles.
Each one is a property buyers can rely on across every dataset.
Schema-locked
Every dataset publishes a versioned schema. Breaking changes bump the major version. Buyers pin to a version and trust forward compatibility within a major.
Source-attributed
Every record points back to the upstream document, system, and the moment it was observed. Citations and audits are derivable from the row, not reconstructed from logs.
Bitemporally honest
Real-world time and system time are kept separate. A buyer asking ‘what did the world look like on date X?’ and ‘what did Codex know on date Y?’ get different answers — both correct.
Spatially consistent
Each package documents its geographic grain and join keys, including H3 resolution 8 where supported. Align geography and coverage before joining records.
Documented source fields
Source categories and derived fields follow each package’s schema. They are not automatically adjudicated labels or authorized training data; governed training use requires separate approval.
Versioned snapshots
Versioned, dated releases; new releases as each family is refreshed. Buyers can pin a release for reproducible analysis, subject to its source terms and approvals.
Joinable by construction
Schemas document the keys available for cross-dataset joins. Use keys both packages supply and align grain, geographic coverage, and time periods; not every dataset joins directly to every other.
What every record carries.
The minimum metadata envelope. Each dataset extends it with its own payload.
- Identity
Stable record and chunk identifiers, plus a fetchable URL pointing to the original source document.
- Schema & lineage
Per-dataset schema version and the version of the pipeline that produced the row, both semver.
- Time
Real-world event time, source publication time, system ingest time, system modification time, and effective-from / effective-to dates where the record carries legal force.
- Confidence & provenance
A confidence score and an ordered chain of the transformations that produced the row.
- Access tier
The APRS “acl” object carries the release’s access tier and a stable license URL for retrieval-time access checks; the applicable source terms still govern permitted uses.
Aligned to the standards your auditors already accept.
Every dataset publishes a per-field crosswalk to the binding external standard for its domain — DCAT-US v3.0 (federally mandated for U.S. open data starting 2026), W3C PROV-DM, schema.org, IHO S-100 + IMO A.600 for maritime, OSCRE IDM for real estate, GS1 GLN for points of interest, NAICS for industry classification.
Browse the per-dataset crosswalks →What this page leaves out.
This page is a summary. Field-by-field schemas and the APRS minimum ingestion contract live in the versioned Codex standard, and each published package documents its own columns, sources, and upstream license notices in its dataset card.