Tutorials
Clean up duplicates on sample data
Find, merge, and prove identity on messy sample records.
Scenario
Two fictional systems — a CRM export and a billing export — describe the same 80 customers with different spellings, missing fields, and conflicting tiers. Your job: one customer list with no duplicates, every merge justified by evidence. The sample hides five hard pairs that naive matching gets wrong.
Image: duplicate candidates side by side with differing spellings.
Definition of done
- One merged customer list, zero duplicates, verified by count.
- Every merge linked to the identity evidence that justified it.
- The five hard pairs resolved deliberately — including any you chose not to merge.
Steps
- Connect both sample sources. Two systems, two scopes — the whole exercise is what happens when the same world arrives twice.
- Define the Customer entity with identifiers from both systems: CRM id and billing id as alternate keys, names and tiers as properties.
- Write identity rules before mapping data. Which fields prove sameness? Name similarity alone is the naive trap the five hard pairs punish — require corroborating keys (matching id, matching email, or matching address plus tier).
- Map and materialize, then count. Eighty plus eighty rows in, and the merged count should read eighty out. Any other number names your next task: over eighty means missed merges, under means false ones.
Image: count check — 160 rows in, merged total out, variance labeled as missed or false merges.
- Work the five hard pairs one by one. Open each candidate, read the evidence for and against, merge with reasons or split with reasons. "Looks the same" is not evidence — cite the keys.
- Rerun and prove stability. Same inputs, same eighty. Then change one source row and rerun again: the merge should absorb it, not fork. Record the absorption as evidence the rules generalize — stability across a changed input is worth more than stability across identical ones.
Image: second rerun diff showing the changed row absorbed with no new duplicates.
Image: merge review showing evidence for and against, with the recorded decision.
What good looks like
- Count correct and stable across reruns.
- Every merge cites keys, not hunches.
- Deliberate non-merges recorded with reasons — restraint is also a result.
Next
- Identity rules reference: Ontology mappings.
- Same flow, your data: Shape your first ontology.