DA0-002 Data Acquisition and Preparation Practice Question
A marketing analyst is reviewing a dataset of campaign responses. The 'response' column contains values 'Yes', 'No', 'Y', 'N', 'yes', and 'no'. Before analysis, the analyst needs to count how many customers responded positively. Which data preparation step is most appropriate?
⚠ Common exam trap
The trap here is assuming that converting to uppercase alone solves the problem, but it ignores abbreviations like 'Y' that still need mapping.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Standardize the values to 'Yes' and 'No' using case normalization and mapping.
The analyst should standardize the values to 'Yes' and 'No' by normalizing case and mapping abbreviations. This ensures all positive responses are consistently represented, enabling accurate counting and analysis. Without standardization, variations like 'Y' and 'yes' would be treated as separate categories, leading to incorrect insights.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert the column to uppercase to make all values consistent.
Why it's wrong here
Converting to uppercase only addresses case differences, not the abbreviation issue. 'Y' would become 'Y' and 'Yes' would become 'YES', still leaving two distinct representations. While case normalization is part of the solution, it alone does not fully standardize the values. The analyst must also map abbreviations to a single form.
- ✗
Filter out rows with 'Y' and 'N' because they are abbreviations.
Why it's wrong here
Filtering out abbreviations would discard valid data and bias the analysis. 'Y' and 'N' are common shorthand for 'Yes' and 'No', and removing them would reduce the sample size and potentially skew results. The analyst should retain and standardize them instead. This approach is not a preparation step but data loss.
- ✗
Create a new binary column where 'Yes' and 'Y' are 1, and all else 0.
Why it's wrong here
Creating a binary column based on the original values would still miss 'yes' and 'Y' if case is not handled. It also assumes that any value not exactly 'Yes' or 'Y' is negative, which could misclassify 'yes'. This approach is incomplete without prior standardization. It could lead to incorrect counts and is not a robust preparation step.
- ✓
Standardize the values to 'Yes' and 'No' using case normalization and mapping.
Why this is correct
Standardizing the values ensures consistency, so that 'Y', 'yes', and 'Yes' are all recognized as positive responses. This simplifies counting and prevents undercounting due to case or abbreviation variations. It is a fundamental data cleaning step that improves data quality and analysis accuracy. Without it, aggregations would be misleading.
Go deeper
Related to this question
About these practice questions
This DA0-002 question is part of Courseiva's 1,004-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This DA0-002 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the DA0-002 exam.