Courseiva

DA0-002 Data Acquisition and Preparation Practice Question

A data analyst needs to merge two customer tables from different sources. One table uses 'CUST_ID' as the primary key, the other uses 'CustomerID'. To ensure accurate merging, the analyst should first:

⚠ Common exam trap

Test-takers frequently assume deduplication (Option C) is the most critical first step, but without first standardizing the join keys, any deduplication logic would operate on mismatched or incomplete data, leading to incorrect results.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Normalize the key column names to a common format

Normalizing key column names to a common format (Option B) is the correct first step because the merge operation requires a consistent join key. Without aligning 'CUST_ID' and 'CustomerID' to a single name and data type, the database or ETL tool will treat them as different columns, resulting in a cross join or an error. This step ensures referential integrity and enables an accurate inner or outer join based on the primary key.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Perform a fuzzy match on names

    Why it's wrong here

    Fuzzy matching on names compares string similarity and cannot equate 'CUST_ID' with 'CustomerID', so no reliable join key emerges. Fuzzy matching is correct when identifiers are absent and records must be linked on approximate textual likeness, such as inconsistent customer name spellings.

  • ✓

    Normalize the key column names to a common format

    Why this is correct

    Mismatched key names such as 'CUST_ID' and 'CustomerID' cause the join to fail or produce a cartesian result. Standardising both columns to one common name and format lets the merge match rows correctly, satisfying the requirement for accurate joining across sources.

  • ✗

    Remove duplicate rows from both tables

    Why it's wrong here

    Deduplicating rows does not reconcile the differing key column names, so the join still cannot match records across the two tables. Duplicate removal is correct when a single table contains repeated records that would inflate counts, not when two sources use incompatible identifier labels.

  • ✗

    Aggregate data by region

    Why it's wrong here

    Aggregating by region collapses individual customer records into summary groups, destroying the row-level granularity a merge requires. Aggregation is correct when producing regional totals or reporting metrics, not when joining two tables whose key columns are named differently and must be reconciled.

About these practice questions

This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.