DA0-002 Data Acquisition and Preparation Practice Question
While profiling a customer dataset, an analyst finds that the 'country' column contains values including 'USA', 'United States', 'U.S.A.', and 'US' for the same nation, plus 'usa ' with trailing whitespace. Reports grouped by country show fragmented counts. Which preparation step resolves this issue?
⚠ Common exam trap
The trap here is mistaking a representation inconsistency for missing or invalid data, which leads to deletion or imputation instead of standardization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply standardization by trimming whitespace and mapping all variants to a single canonical country code.
The column is not missing data; it holds the same nation under several spellings and spacing. Standardizing with trimming plus a canonical mapping such as an ISO code unifies those variants so grouped counts consolidate correctly. Deleting rows, widening the column, or imputing a single value either destroys data or leaves the fragmentation untouched.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Impute the mode of the column for every row to make all values identical.
Why it's wrong here
Overwriting every row with the most frequent value erases genuine differences between customers in different countries, producing a column that is uniform but false. Imputation is meant for missing values, not for replacing observed valid ones. It also destroys the analytical value of the country dimension entirely, making any geographic reporting meaningless.
- ✓
Apply standardization by trimming whitespace and mapping all variants to a single canonical country code.
Why this is correct
Standardization collapses the synonymous spellings and whitespace variants into one canonical representation, which is exactly what eliminates the fragmented grouping. Mapping to a controlled code such as ISO 3166 also makes future joins and comparisons reliable. Trimming handles the formatting defect while the mapping handles the semantic equivalence, so both parts of the problem are addressed together.
- ✗
Remove all rows containing any country value that is not 'USA' to enforce consistency.
Why it's wrong here
Deleting valid records to force uniformity destroys real data and biases every downstream metric by silently dropping customers from other nations or from variant spellings of the same nation. It confuses consistency with correctness. The fragmentation problem is one of representation, not of having too many countries, so removal is both destructive and ineffective.
- ✗
Increase the column length to accommodate the longest variant string.
Why it's wrong here
Column width has no bearing on grouping behavior when the values themselves differ textually; 'USA' and 'United States' would still be treated as distinct even in a very wide column. This change addresses storage sizing, not semantic equivalence. It leaves the fragmented counts completely unchanged while adding unnecessary schema churn.
Visual reference
Go deeper
Related to this question
About these practice questions
This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.