Courseiva
mediumMultiple Select

AIF-C01 Practice Question: A data scientist is preparing data for a…

A data scientist is preparing data for a classification model. The dataset contains missing values in several features. Which TWO approaches are appropriate for handling missing data? (Select TWO.)

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Remove rows with any missing values

Option A (Remove rows with any missing values) is correct because listwise deletion is a standard, valid preprocessing approach when missingness is sparse or completely at random, producing a clean dataset that most classifiers can consume directly. Option E (Impute missing values with the median of the feature) is correct because median imputation is a widely accepted technique that preserves all records and is robust to outliers and skewed distributions, making it suitable for numeric features feeding a classification model. Option B (Set missing values to zero) is not appropriate in general because zero is a real value that can distort the feature's distribution and mislead the model unless zero is a meaningful sentinel. Option C (Ignore missing values during training) is not valid because most classification algorithms cannot natively process NaN values and will error or produce biased results. Option D (Replace missing values with -1) is also not generally appropriate because -1 is an arbitrary constant that can introduce artificial patterns and skew numeric features unless it is a documented missing-value code.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Remove rows with any missing values

    Why this is correct

    Dropping rows containing any missing value is valid when missingness is sparse and random, since the remaining complete records still represent the population. It preserves data integrity without fabricating values, though it reduces sample size and can bias results if missingness is systematic.

  • ✗

    Set missing values to zero

    Why it's wrong here

    Setting missing values to zero fabricates a real data point, distorting distributions and correlations for features where zero is not a valid measurement, so the classifier learns spurious patterns. It is tempting because zero-filling is trivial and defensible for count features, where zero genuinely represents "no occurrences" and imputation is unnecessary.

  • ✗

    Ignore missing values during training

    Why it's wrong here

    Ignoring missing values during training discards incomplete rows, shrinking the dataset and biasing the model toward records with no gaps; scikit-learn estimators reject NaN outright, so training fails. It is tempting because dropping nulls is a legitimate preprocessing step when missingness is random and sparse, and the remaining volume still represents the population adequately.

  • ✗

    Replace missing values with -1

    Why it's wrong here

    Replacing missing values with −1 injects a fabricated numeric value that the model treats as a genuine observation, distorting distributions and any distance or gradient computation. It is tempting because sentinel encoding suits tree-based models when −1 lies outside the valid feature range, but for general classification preprocessing it manufactures false signal rather than addressing the missingness.

  • ✓

    Impute missing values with the median of the feature

    Why this is correct

    Median imputation replaces missing entries with the feature's median, which is robust to outliers and skew compared with the mean. It retains all rows, preserving sample size, and is appropriate for the numerical features described, though it reduces variance and can distort correlations.

About these practice questions

One of 862 original AIF-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.