DA0-002 Data Acquisition and Preparation Practice Question
A data analyst needs to merge two customer tables from different sources. One table uses 'CUST_ID' as the primary key, the other uses 'CustomerID'. To ensure accurate merging, the analyst should first:
⚠ Common exam trap
Test-takers frequently assume deduplication (Option C) is the most critical first step, but without first standardizing the join keys, any deduplication logic would operate on mismatched or incomplete data, leading to incorrect results.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Normalize the key column names to a common format
Normalizing key column names to a common format (Option B) is the correct first step because the merge operation requires a consistent join key. Without aligning 'CUST_ID' and 'CustomerID' to a single name and data type, the database or ETL tool will treat them as different columns, resulting in a cross join or an error. This step ensures referential integrity and enables an accurate inner or outer join based on the primary key.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Perform a fuzzy match on names
Why it's wrong here
Fuzzy matching on names compares string similarity and cannot equate 'CUST_ID' with 'CustomerID', so no reliable join key emerges. Fuzzy matching is correct when identifiers are absent and records must be linked on approximate textual likeness, such as inconsistent customer name spellings.
- ✓
Normalize the key column names to a common format
Why this is correct
Mismatched key names such as 'CUST_ID' and 'CustomerID' cause the join to fail or produce a cartesian result. Standardising both columns to one common name and format lets the merge match rows correctly, satisfying the requirement for accurate joining across sources.
- ✗
Remove duplicate rows from both tables
Why it's wrong here
Deduplicating rows does not reconcile the differing key column names, so the join still cannot match records across the two tables. Duplicate removal is correct when a single table contains repeated records that would inflate counts, not when two sources use incompatible identifier labels.
- ✗
Aggregate data by region
Why it's wrong here
Aggregating by region collapses individual customer records into summary groups, destroying the row-level granularity a merge requires. Aggregation is correct when producing regional totals or reporting metrics, not when joining two tables whose key columns are named differently and must be reconciled.
About these practice questions
This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.