Courseiva

DA0-002 Data Acquisition and Preparation Practice Question

A data analyst is preparing a dataset for analysis and discovers that the 'customer_id' column has missing values. The analyst decides to remove all rows with missing 'customer_id' because it is a primary key. Which data preparation technique is being applied?

⚠ Common exam trap

A common mix-up: candidates confuse missing value handling with data transformation techniques like normalization or discretization, which do not address missing data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Listwise deletion

Listwise deletion is the correct technique because it involves removing records with missing values. Since customer_id is a primary key, missing values cannot be reliably filled, so deleting those rows is a standard data cleaning step. This ensures that each record has a valid unique identifier, which is essential for relational integrity and accurate analysis.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Discretization

    Why it's wrong here

    Discretization converts continuous data into discrete bins, such as grouping ages into ranges. It does not handle missing values. Applying discretization to customer_id would be inappropriate because it is an identifier, not a continuous variable, and the missing values would remain unresolved.

  • ✓

    Listwise deletion

    Why this is correct

    Listwise deletion removes entire records that have missing values in any column. Here, the analyst removes rows where customer_id is missing, which is exactly listwise deletion applied to a specific column. Since customer_id is a primary key, missing values cannot be imputed, so removal is a valid approach.

  • ✗

    Normalization

    Why it's wrong here

    Normalization is a scaling technique that adjusts numerical values to a common scale, such as 0-1. It does not address missing values. Applying normalization to customer_id would be meaningless because it is an identifier, not a numeric measure, and the missing values would still need handling separately.

  • ✗

    Imputation

    Why it's wrong here

    Imputation involves filling missing values with estimated ones, such as the mean or median. Since customer_id is a primary key, imputing values would create artificial identifiers that do not correspond to real customers, leading to data integrity issues. Therefore, imputation is not appropriate here.

About these practice questions

One of 1,004 original DA0-002 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.