Introduction
DC’s Vision for a Connected Data System
Unifying Data through Identity Resolution Using P3RL Methodology
What P3RL Is
P3RL (Privacy Preserving Probabilistic Record Linkage) is Resultant’s proprietary identity resolution solution. It assigns a unique identifier to each individual across all contributing datasets, making longitudinal analysis possible while keeping raw PII secure throughout the process.
What makes P3RL distinctive is its ability to resolve identity across datasets that were never designed to share a common identifier, and to do so reliably at scale. Real-world government data carries the kinds of quality issues that rule-based matching can’t handle without building an ever-growing and increasingly fragile ruleset.
How It Works
P3RL draws on multiple layers of matching logic, including:
Domain knowledge rules are encoded as version-controlled configurations, keeping matching logic transparent, auditable, and adaptable as the system grows to include new data sources. Each new source added to the system extends coverage and strengthens the accuracy of matches already made.
The result is a unique identifier assigned to each individual across all data-contributing agencies, creating the common thread that makes longitudinal analysis possible.
Platform Architecture
Azure Databricks for Security and Scale
Workspace Structure
Medallion Architecture
Governance via Unity Catalog
Data Integration at Scale
Initial Research Outputs
Knowledge Transfer and District Ownership
Lessons for Other Jurisdictions
Inter-agency data sharing requires dedicated project management, not just technical coordination.
ETEP led the data-sharing agreements across partner agencies. Keeping that workstream moving in parallel with technical development was essential to keeping the project on track.
Identity resolution is one of the most consequential problems to solve in a P20W+ system.
Without a reliable, privacy-preserving way to link individuals across agency systems, each agency's data can only tell its own part of the story. Getting this right early is what makes cross-agency longitudinal analysis possible at all.
Platform selection should follow client needs, not the other way around.
For DC, that process led to building on Azure Databricks, a platform already in use across several agencies. Starting from a clear understanding of what ETEP needed, who would use the system, and how it would grow shaped every technical decision that followed.
Governance needs to be configured into the system from the beginning.
On this project, that meant building Unity Catalog's access controls, audit logging, and data lineage tracking into the architecture as it was designed, not after the fact. A research environment that can't demonstrate how data moves and who can see it won't earn the trust of partner agencies or researchers.
On this project, embedding knowledge transfer throughout delivery rather than scheduling it at the end made a measurable difference.
Design decisions were documented and explained as they were made. The goal from the start was for ETEP to understand the system fully, not just operate it.
Resultant works with state and municipal agencies to design and implement P20W+ data systems built for long-term use. Connect with our Education and Workforce team at resultant.com.
We’re proud to help organizations thrive, and we’d love to tell you more.