HeadlinesBriefing favicon HeadlinesBriefing.com

Deterministic Stages Beat Similarity Scores for Deduplication

Towards Data Science •
×

Deduplicating a 10,000-row supplier list in Python reveals that string matching is the easy part — deciding what to do with similarity scores is the real job. The author's agency buys editorial placements from independent publishers, and duplicate supplier records cost money when invoices get paid twice or price history splits across spellings nobody searches for.

The same site shows up as solartravelmag.com, https://www.solartravelmag.com/, blog.solartravelmag.com, and Solar Travel Mag. COM with tracking parameters. Since January, an external marketplace catalog has been synced through its API on top of a decade of manual entry. One marketplace admits its listed prices match reality about 90% of the time, making cross-checking mandatory.

A synthetic rebuild of the list contained 11,531 rows describing 7,180 actual vendors, with 2,429 surface variants, 888 subdomain rows, 529 country-domain siblings, 505 typos, and 180 near-name traps. Two deterministic stages — normalization and subdomain collapse using the Public Suffix List — removed 76% of duplicates for free. The fuzzy matcher scored 33.7 million pairs in two seconds, but no threshold could safely merge what remained. The matcher's real output is a ranked review queue, not a set of merges.