Courseiva

DA0-002 Data Acquisition and Preparation Practice Question

A marketing company is building a customer segmentation model. The data team has access to two sources: a CRM database with customer demographics and purchase history, and a third-party data provider that offers social media activity scores. The CRM data is updated daily, while the third-party data is refreshed weekly on Sundays. The analyst needs to create a unified dataset for the model training scheduled for Wednesday morning. The analyst runs a SQL query to join the two tables on CustomerID, but the resulting dataset has far fewer rows than expected. Upon investigation, the analyst finds that many customers in the CRM do not have matching records in the third-party data. Additionally, some customers in the third-party data have multiple entries due to unresolved duplicates. The analyst must produce the most complete dataset possible while maintaining data quality. Which course of action should the analyst take?

⚠ Common exam trap

Many candidates choose an INNER JOIN (Option B) thinking it ensures data quality by only including matched records, but they overlook the requirement for completeness, which necessitates preserving all CRM customers even without third-party matches.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

First deduplicate the third-party data by keeping the most recent record per CustomerID, then perform a LEFT JOIN from CRM to the deduplicated third-party data.

It first resolves the duplicate issue in the third-party data by keeping the most recent record per CustomerID, ensuring each customer has a single, current social media score. Then, a LEFT JOIN from CRM to the deduplicated third-party data preserves all CRM customers, maximizing completeness while maintaining data quality. This approach aligns with the goal of producing the most complete dataset for model training, as the CRM is the primary source with daily updates.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    First deduplicate the third-party data by keeping the most recent record per CustomerID, then perform a LEFT JOIN from CRM to the deduplicated third-party data.

    Why this is correct

    Deduplicating the third-party table to one row per CustomerID prevents the join from multiplying CRM records, and a LEFT JOIN preserves every CRM customer, including those with no social media match, maximising completeness while maintaining quality.

  • ✗

    Perform an INNER JOIN on CustomerID and then remove duplicates from the result.

    Why it's wrong here

    An INNER JOIN retains only CustomerIDs present in both sources, which is exactly what already produced the shrunken result; removing duplicates afterwards does not restore the dropped CRM customers. INNER JOIN suits scenarios where only mutually matched records are valid.

  • ✗

    Use only the third-party data because it provides the social media scores needed for segmentation.

    Why it's wrong here

    Discarding CRM entirely abandons demographics and purchase history, the core segmentation features, and inherits unresolved duplicate rows. Third-party-only data suits enrichment or social-listening analysis, not a model requiring complete customer coverage joined to transactional attributes.

  • ✗

    Perform a LEFT JOIN from the third-party data to CRM, then aggregate duplicates by averaging scores.

    Why it's wrong here

    Reversing the join direction keeps only third-party rows, discarding CRM customers lacking social scores and so shrinking the dataset rather than maximising completeness. A LEFT JOIN from CRM to third-party, with duplicate scores aggregated, is the pattern that preserves every CRM customer.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.