Courseiva
AI Concepts and Foundations →mediumMultiple Select

AI0-001 AI Concepts and Foundations Practice Question

Which TWO techniques are commonly used to handle missing data in a dataset?

⚠ Common exam trap

CompTIA often tests the distinction between data preprocessing techniques that handle missing values versus those that transform or reduce features, so candidates may confuse feature scaling or PCA with missing data handling because they are all part of data preparation.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Remove rows with missing values

Removing rows with missing values is a straightforward technique to handle missing data, especially when the missingness is random and the dataset is large enough that dropping a few rows does not significantly reduce the sample size or introduce bias. Option D is correct because imputing missing values with the mean or median is a common statistical method that preserves the dataset size and is simple to implement, though it can reduce variance and may distort relationships if the data is not missing completely at random.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Feature scaling

    Why it's wrong here

    Feature scaling rescales existing numeric ranges to comparable magnitudes; it neither detects nor fills absent values, so nulls remain. It is tempting because scaling is a routine preprocessing step, and it would be correct when features differ wildly in units or magnitude and the algorithm is distance- or gradient-based.

  • ✗

    One-hot encoding

    Why it's wrong here

    One-hot encoding converts categorical levels into binary indicator columns; a missing value simply yields all-zero or an error, not an imputed entry. It is tempting because encoding is standard preprocessing, and it would be correct when a nominal feature must be represented numerically for a model that cannot consume categories directly.

  • ✓

    Remove rows with missing values

    Why this is correct

    Dropping rows containing missing values is a straightforward deletion technique that keeps remaining data untouched. It works when missingness is random and sparse, but discards information and can bias results if the missing records differ systematically from the retained ones.

  • ✓

    Impute with mean or median

    Why this is correct

    Replacing missing values with the column mean or median preserves the row count and keeps the dataset usable for algorithms that reject nulls. It suits numerical data and reduces bias from deletion, though it shrinks variance and can distort correlations.

  • ✗

    Principal component analysis (PCA)

    Why it's wrong here

    PCA reduces dimensionality by projecting correlated features onto orthogonal components; it cannot impute absent values, so records with gaps still break the analysis. It is tempting because PCA is a standard preprocessing step, and it would be the right choice when the goal is compressing many correlated numeric features rather than filling missing entries.

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.