Courseiva
Data Analysis →hardMultiple Select

DA0-002 Data Analysis Practice Question

A data analyst is performing data cleaning. Which THREE steps are part of this process? (Choose three.)

⚠ Common exam trap

Many candidates confuse data cleaning with data transformation or feature engineering, leading them to select normalization or feature engineering as cleaning steps, when in fact cleaning strictly addresses data quality issues like consistency, completeness, and uniqueness.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Correcting inconsistent data

Data cleaning is the process of detecting and correcting (or removing) corrupt or inaccurate records from a dataset, so option A (Correcting inconsistent data) is correct because fixing mismatched formats, units, or values (e.g., 'NY' vs 'New York') is a core cleaning task. Option C (Handling missing values) is correct because cleaning must address nulls or blanks through imputation, deletion, or flagging so downstream analysis is not skewed. Option E (Removing duplicate records) is correct because duplicate rows inflate counts and distort aggregates, and deduplication is a standard cleaning step. Option B (Normalization) is not part of cleaning; it is a data transformation/scaling technique (e.g., min-max or z-score) typically applied during preprocessing/modeling. Option D (Feature engineering) is also not cleaning; it is the creation of new derived variables for modeling, which occurs after cleaning.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Correcting inconsistent data

    Why this is correct

    Correcting inconsistent data resolves conflicting formats, units and values so records agree across sources. This is a core data cleaning activity, satisfying the stem's requirement to identify steps that standardise and reconcile raw data before analysis.

  • ✗

    Normalization

    Why it's wrong here

    Normalization scales data, part of transformation, not cleaning.

  • ✓

    Handling missing values

    Why this is correct

    Handling missing values addresses nulls through imputation, removal or flagging, preventing biased or failed analysis. This is a core data cleaning step, satisfying the stem's requirement to identify activities that remediate incomplete records before analysis.

  • ✗

    Feature engineering

    Why it's wrong here

    Feature engineering creates new variables from existing data, which occurs after cleaning rather than as part of it. It is tempting because both stages sit in the preparation pipeline, and it would be correct when deriving predictive attributes for a model.

  • ✓

    Removing duplicate records

    Why this is correct

    Removing duplicate records directly satisfies the data cleaning requirement by eliminating redundant rows that would otherwise skew counts, sums and averages. Deduplication ensures each entity appears once, preserving analytical accuracy before aggregation or modelling. This is a core cleaning task, distinct from transformation or validation steps performed later in the pipeline.

About these practice questions

Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.