Sen CutlerData Engineer

Data Matching & Deduplication Across 10.8M Rows

The project called for me to pull, clean, and normalize the datasets. I used a PostgreSQL database to handle the 10.8M+ rows efficiently, and built and refined matching logic to reduce false positives/negatives, by incorporating names, companies, and locations. I fixed data formatting issues, reprocessed missing records, and debugged inconsistencies in source files. I delivered structured reports with all requested details.

  • Python
  • PostgreSQL
  • Data Cleaning
  • Data Engineering

Data Processing · Data Quality