Summary
Before Washington, D.C. could connect a person’s path from early learning through the workforce, it had to solve a narrower problem first: reliably identifying when records from five different agencies belonged to the same person.
[Estimated read time: 2 minutes]
Following a person’s path across systems
A student enrolls in an early learning program. Years later, they show up in a wage record. Somewhere in between, they may have completed a credential, entered workforce training, or moved through several public systems that never talked to each other.
Understanding what happened along that path, and which education and workforce experiences contributed to later outcomes, means connecting records that live in separate agency systems.
Before any of that connecting can happen, there’s a narrower problem to solve first: how to determine when the record in one system and the record in another belong to the same individual.
One of the most consequential problems in a longitudinal system
Names and addresses change over time, fields may be incomplete, and agencies can record and format the same information differently. Matching only identical records can miss connections that should be made. Looser matching can create connections that shouldn’t.
Without a reliable, privacy-preserving way to link individuals across agency systems, each agency’s data can only tell its own part of the story.
How Washington, D.C. built the first P20W+ longitudinal system on Databricks with Resultant
The District of Columbia’s Office of Education Through Employment Pathways (ETEP) encountered this problem while developing its Education Through Employment Data (ETE) System. In its first two years, the system has incorporated 30 datasets from six partner agencies, covering more than a decade of education and workforce history.
Resultant’s proprietary identity-resolution methodology, Privacy Preserving Probabilistic Record Linkage (P3RL), provides a reliable way to match records across datasets while keeping raw personally identifiable information secure throughout the process.
But reliable identity resolution was only the beginning of what it takes to build a longitudinal system agencies can trust, researchers can use, and the District can sustain as its data and research priorities evolve.
Where the whitepaper picks up
The full whitepaper goes inside the technical and governance decisions behind D.C.’s ETE Data System, including the architecture, researcher access, inter-agency data sharing, and knowledge transfer that supported long-term District ownership. It also shares lessons from the implementation for other jurisdictions considering similar work.