AI0-001 AI Models and Data Engineering Practice Question
Exhibit
Data Validation Report: Table: customers - column "age": null values: 0, unique values: 87, min:18, max:99 - column "income": null values: 12, unique values: 1500, min:0, max:500000 - column "region": null values: 0, unique values: 4, values: ["North", "South", "East", "West"] - column "gender": null values: 0, unique values: 2, values: ["M", "F"]
Refer to the exhibit. A data engineer runs a validation report on the customers table. The "income" column has 12 null values. Which imputation strategy is most appropriate for this column?
⚠ Common exam trap
CompTIA often tests the misconception that a global mean or median is always the safest imputation, when in fact ignoring subgroup structure (like region) can introduce significant bias and violate the assumption of missing-at-random conditioned on observed features.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Replace nulls with the median income per region
Imputing missing income values with the median per region preserves the central tendency of each regional subgroup, which is robust to outliers and maintains the distributional characteristics of the data. This strategy is particularly appropriate for income data, which often exhibits skewness and regional variation, ensuring that the imputed values are contextually relevant and do not distort downstream analytics or machine learning models.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Remove rows with null income
Why it's wrong here
Deleting rows discards the remaining valid attributes of those 12 customers, biasing any downstream analysis. Row removal suits columns with very few nulls where the lost records are negligible and no meaningful distribution shift results; here imputation preserves the data.
- ✓
Replace nulls with the median income per region
Why this is correct
Income is typically right-skewed and contains outliers, so the mean would be distorted; the median is robust. Computing it per region preserves local wage differences, satisfying the requirement to impute the 12 nulls without flattening genuine geographic variation in the column.
- ✗
Replace nulls with 0
Why it's wrong here
Zero income is unrealistic and would introduce bias.
- ✗
Replace nulls with the mean income of the entire dataset
Why it's wrong here
Mean imputation assumes a symmetric distribution; the exhibit shows income is skewed, so the mean is pulled by high earners and distorts the 12 replacements. It is tempting because mean imputation is the standard default for roughly normal numeric columns, where it would indeed be the correct choice.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.