Group of young people in technical vocational training with teacher

The Problem Every P20W+ Data System Has to Solve

Summary

Before Washington, D.C. could connect a person’s path from early learning through the workforce, it had to solve a narrower problem first: reliably identifying when records from five different agencies belonged to the same person.

[Estimated read time: 2 minutes]

 

Following a person’s path across systems

A student enrolls in an early learning program. Years later, they show up in a wage record. Somewhere in between, they may have completed a credential, entered workforce training, or moved through several public systems that never talked to each other.

Understanding what happened along that path, and which education and workforce experiences contributed to later outcomes, means connecting records that live in separate agency systems.

Before any of that connecting can happen, there’s a narrower problem to solve first: how to determine when the record in one system and the record in another belong to the same individual.

 

One of the most consequential problems in a longitudinal system

Names and addresses change over time, fields may be incomplete, and agencies can record and format the same information differently. Matching only identical records can miss connections that should be made. Looser matching can create connections that shouldn’t.

Without a reliable, privacy-preserving way to link individuals across agency systems, each agency’s data can only tell its own part of the story.

 

How Washington, D.C. built the first P20W+ longitudinal system on Databricks with Resultant

The District of Columbia’s Office of Education Through Employment Pathways (ETEP) encountered this problem while developing its Education Through Employment Data (ETE) System. In its first two years, the system has incorporated 30 datasets from six partner agencies, covering more than a decade of education and workforce history.

Resultant’s proprietary identity-resolution methodology, Privacy Preserving Probabilistic Record Linkage (P3RL), provides a reliable way to match records across datasets while keeping raw personally identifiable information secure throughout the process.

But reliable identity resolution was only the beginning of what it takes to build a longitudinal system agencies can trust, researchers can use, and the District can sustain as its data and research priorities evolve.

Where the whitepaper picks up

The full whitepaper goes inside the technical and governance decisions behind D.C.’s ETE Data System, including the architecture, researcher access, inter-agency data sharing, and knowledge transfer that supported long-term District ownership. It also shares lessons from the implementation for other jurisdictions considering similar work.

Read the whitepaper here.

Connect

Find out how our team can help you achieve great outcomes.

Insights delivered to your inbox