AI0-001 AI Concepts and Foundations Practice Question
A data scientist is preparing a dataset for a classification task. The dataset contains 10,000 rows and 50 features, but many features have missing values. Which approach should the scientist take first to address the missing data?
⚠ Common exam trap
CompTIA often tests the misconception that immediate imputation (e.g., mean/median) or row deletion is the safest first step, when in reality, a diagnostic analysis of missingness patterns is required before any data modification.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Analyze the pattern and proportion of missing values to choose an appropriate imputation strategy.
The first step in handling missing data is to understand the pattern and proportion of missingness (e.g., MCAR, MAR, MNAR) to select an appropriate imputation method. Blindly applying imputation or deletion without analysis can introduce bias or reduce model performance. This diagnostic step ensures the chosen strategy aligns with the data's underlying structure and the classification task's requirements.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a deep learning model to predict missing values without preprocessing.
Why it's wrong here
Imputation requires training data, so a model cannot predict missing values while those values remain absent from its inputs. It is tempting because learned imputation can outperform simple statistics, and would be correct after an initial baseline imputation step.
- ✓
Analyze the pattern and proportion of missing values to choose an appropriate imputation strategy.
Why this is correct
Before imputing, the scientist must understand whether missingness is random or systematic and how much data is affected, since this determines whether mean, median, or model-based imputation is valid. This diagnostic step precedes any imputation choice.
- ✗
Remove all rows with any missing values to ensure a clean dataset.
Why it's wrong here
Dropping every incomplete row discards most of the 10,000 records when missingness spans many of the 50 features, biasing the classifier. It is tempting because listwise deletion is quick and clean, and would suit a dataset with only a handful of affected rows.
- ✗
Replace missing values with the mean of each feature immediately.
Why it's wrong here
Mean substitution distorts each feature's distribution and understates variance before any missingness pattern is examined. It is tempting because it is fast and keeps all rows, and would be reasonable for a roughly symmetric feature with very few gaps.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.