DA0-002 Data Acquisition and Preparation Practice Question
A data analyst needs to identify duplicate customer records. Which TWO methods are commonly used? (Select two.)
⚠ Common exam trap
Watch out — candidates often choose 'Exact match on all fields' (Option E) thinking it is a reliable deduplication method, but in practice it fails to catch real-world duplicates that have any minor variation, and the exam expects you to recognize that fuzzy matching and sorted adjacency comparisons are the standard techniques for duplicate detection.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Fuzzy matching using Levenshtein distance
Option A (Fuzzy matching using Levenshtein distance) is correct because it measures the minimum number of single-character edits (insertions, deletions, substitutions) needed to transform one string into another, making it ideal for catching near-duplicate customer records that differ slightly due to typos or formatting variations. Option B (Sorting and comparing adjacent rows) is correct because once records are sorted by a key field such as name or email, duplicate or near-duplicate entries naturally cluster together, allowing efficient pairwise comparison of neighboring rows to flag matches. Option C is not a reliable, scalable method since random sampling cannot guarantee detection of all duplicates and is subjective. Option D does not help because hashing a primary key, which is unique by definition, will never reveal duplicates. Option E is too strict, as exact matching on all fields will miss duplicates that differ in even one attribute, such as a middle initial or apartment number.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Fuzzy matching using Levenshtein distance
Why this is correct
Levenshtein distance measures the minimum single-character edits between strings, so records differing by typos or transpositions still match. This satisfies the need to catch near-duplicates that exact key comparison would miss across customer names and addresses.
- ✓
Sorting and comparing adjacent rows
Why this is correct
Sorting on a matching key clusters similar records adjacently, so comparing neighbouring rows reveals exact duplicates cheaply. This satisfies the duplicate-detection requirement for large datasets where pairwise comparison of every record would be computationally prohibitive.
- ✗
Visual inspection of random sample
Why it's wrong here
Inspecting a random sample cannot reliably surface duplicates across the full dataset, since matching pairs may fall outside the sample. It tempts as a quick sanity check, but systematic methods such as sorting, grouping, fuzzy matching or SQL self-joins are needed to detect duplicates deterministically.
- ✗
Using a hash function on primary key
Why it's wrong here
Hashing a primary key yields a unique value per record, so identical customers with different keys produce different hashes and no duplicates surface. Hashing suits detecting changed rows or building surrogate keys, not matching records that already carry distinct identifiers.
- ✗
Exact match on all fields
Why it's wrong here
Requiring every field to match misses duplicates where any attribute differs, such as a changed surname or postcode, so near-identical customers stay undetected. Exact full-row matching suits confirming identical rows in a clean dataset, not fuzzy customer deduplication.
Go deeper
Related to this question
About these practice questions
One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.