AI0-001 AI Concepts and Foundations Practice Question
Which TWO techniques are commonly used to handle missing data in a dataset?
⚠ Common exam trap
CompTIA often tests the distinction between data preprocessing techniques that handle missing values versus those that transform or reduce features, so candidates may confuse feature scaling or PCA with missing data handling because they are all part of data preparation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Remove rows with missing values
Removing rows with missing values is a straightforward technique to handle missing data, especially when the missingness is random and the dataset is large enough that dropping a few rows does not significantly reduce the sample size or introduce bias. Option D is correct because imputing missing values with the mean or median is a common statistical method that preserves the dataset size and is simple to implement, though it can reduce variance and may distort relationships if the data is not missing completely at random.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Feature scaling
Why it's wrong here
Feature scaling rescales existing numeric ranges to comparable magnitudes; it neither detects nor fills absent values, so nulls remain. It is tempting because scaling is a routine preprocessing step, and it would be correct when features differ wildly in units or magnitude and the algorithm is distance- or gradient-based.
- ✗
One-hot encoding
Why it's wrong here
One-hot encoding converts categorical levels into binary indicator columns; a missing value simply yields all-zero or an error, not an imputed entry. It is tempting because encoding is standard preprocessing, and it would be correct when a nominal feature must be represented numerically for a model that cannot consume categories directly.
- ✓
Remove rows with missing values
Why this is correct
Dropping rows containing missing values is a straightforward deletion technique that keeps remaining data untouched. It works when missingness is random and sparse, but discards information and can bias results if the missing records differ systematically from the retained ones.
- ✓
Impute with mean or median
Why this is correct
Replacing missing values with the column mean or median preserves the row count and keeps the dataset usable for algorithms that reject nulls. It suits numerical data and reduces bias from deletion, though it shrinks variance and can distort correlations.
- ✗
Principal component analysis (PCA)
Why it's wrong here
PCA reduces dimensionality by projecting correlated features onto orthogonal components; it cannot impute absent values, so records with gaps still break the analysis. It is tempting because PCA is a standard preprocessing step, and it would be the right choice when the goal is compressing many correlated numeric features rather than filling missing entries.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.