DA0-002 Data Analysis Practice Question
A data analyst at a logistics company receives a shipment-tracking table in which the carrier identifier column contains values such as 'FEDX', 'FedEx', 'fed ex', and 'FX'. The analyst needs to consolidate these into a single carrier value before building a delivery-performance dashboard. Which data preparation step should the analyst perform?
⚠ Common exam trap
The trap here is reaching for a transformation like one-hot encoding or normalization because the column looks messy, when the actual defect is inconsistent labels that only cleansing and standardization can merge.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Standardize and normalize the categorical values through data cleansing.
The carrier column suffers from inconsistent representation of the same entity, a classic data cleansing issue. Trimming whitespace, normalizing case, and mapping aliases to a canonical label collapses the variants into one category. Numeric rescaling, missing-value imputation, and one-hot encoding all address different problems and would leave the duplicated carrier categories intact in the delivery dashboard.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
One-hot encode the carrier column into binary indicator fields.
Why it's wrong here
One-hot encoding converts each distinct category into its own binary column. Performed on the raw data, it would create four separate carrier indicators for what is really one carrier, actively worsening the fragmentation. Encoding is a modeling-prep step applied after categories are clean and consolidated, not a remedy for inconsistent label spellings.
- ✓
Standardize and normalize the categorical values through data cleansing.
Why this is correct
These entries are the same carrier recorded inconsistently, so the analyst should cleanse the column by trimming whitespace, unifying case, and mapping known aliases to one canonical value. Consolidating variants into a single label prevents the dashboard from splitting one carrier's shipments across four bogus categories, which would distort on-time rates and volume comparisons.
- ✗
Impute the missing carrier values using the column mean.
Why it's wrong here
Nothing is missing here; every row contains a carrier string. Imputing a mean is a numeric operation reserved for absent values in quantitative columns, and a mean carrier identifier is conceptually meaningless. Applying it would neither detect nor resolve the alias variants, so the analyst would still see four separate carrier categories in the dashboard.
- ✗
Apply min-max normalization to rescale the carrier column.
Why it's wrong here
Min-max normalization rescales numeric values into a bounded range such as 0 to 1. It has no meaning for text labels and would fail outright on this column. The problem is inconsistent spelling and casing, not numeric magnitude, so a scaling transformation cannot merge 'FEDX' with 'FedEx' and leaves the dashboard fragmentation untouched.
Go deeper
Related to this question
About these practice questions
Courseiva writes every DA0-002 question from scratch — 1,004 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.