Semogram Docs
Tutorials

Clean up duplicates on sample data

Find, merge, and prove identity on messy sample records.

Scenario

Two fictional systems — a CRM export and a billing export — describe the same 80 customers with different spellings, missing fields, and conflicting tiers. Your job: one customer list with no duplicates, every merge justified by evidence. The sample hides five hard pairs that naive matching gets wrong.

Image: duplicate candidates side by side with differing spellings.

Definition of done

  • One merged customer list, zero duplicates, verified by count.
  • Every merge linked to the identity evidence that justified it.
  • The five hard pairs resolved deliberately — including any you chose not to merge.

Steps

  1. Connect both sample sources. Two systems, two scopes — the whole exercise is what happens when the same world arrives twice.
  2. Define the Customer entity with identifiers from both systems: CRM id and billing id as alternate keys, names and tiers as properties.
  3. Write identity rules before mapping data. Which fields prove sameness? Name similarity alone is the naive trap the five hard pairs punish — require corroborating keys (matching id, matching email, or matching address plus tier).
  4. Map and materialize, then count. Eighty plus eighty rows in, and the merged count should read eighty out. Any other number names your next task: over eighty means missed merges, under means false ones.

Image: count check — 160 rows in, merged total out, variance labeled as missed or false merges.

  1. Work the five hard pairs one by one. Open each candidate, read the evidence for and against, merge with reasons or split with reasons. "Looks the same" is not evidence — cite the keys.
  2. Rerun and prove stability. Same inputs, same eighty. Then change one source row and rerun again: the merge should absorb it, not fork. Record the absorption as evidence the rules generalize — stability across a changed input is worth more than stability across identical ones.

Image: second rerun diff showing the changed row absorbed with no new duplicates.

Image: merge review showing evidence for and against, with the recorded decision.

What good looks like

  • Count correct and stable across reruns.
  • Every merge cites keys, not hunches.
  • Deliberate non-merges recorded with reasons — restraint is also a result.

Next