Courseiva
mediumMultiple Select

AIF-C01 Practice Question: A data scientist is building a regression model…

A data scientist is building a regression model to predict house prices. During feature engineering, they have categorical variables (e.g., neighborhood) and numerical variables (e.g., square footage) with missing values. Which TWO actions should the data scientist take? (Choose two.)

⚠ Common exam trap

In the AWS AI Practitioner exam, a common trap is that candidates may incorrectly choose label encoding (Option D) because it seems simpler, without recognizing that it introduces an arbitrary ordinal relationship that degrades model performance.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Impute missing square footage values with the median

Option C is correct because imputing missing square footage values with the median is a robust strategy for numerical features, especially when outliers may be present, and it preserves the row count for model training. Option E is correct because one-hot encoding is the appropriate technique for nominal categorical variables like neighborhood, as it creates binary columns without imposing an artificial ordinal relationship. Option A is not ideal because dropping all rows with any missing values can discard substantial useful data and introduce bias. Option B is not required for regression models and may not be necessary unless the algorithm is sensitive to feature scale. Option D is incorrect because label encoding would impose an arbitrary order on neighborhood categories, which is inappropriate for a nominal variable.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Drop all rows with any missing values

    Why it's wrong here

    Dropping rows reduces data size; imputation is generally preferred unless missingness is excessive.

  • ✗

    Normalize all numerical features to [0,1] range

    Why it's wrong here

    Normalization is not always required; tree-based models are scale-invariant.

  • ✓

    Impute missing square footage values with the median

    Why this is correct

    Square footage is numerical and skewed by outliers, so median imputation preserves the central tendency without being distorted by extreme values, unlike the mean. This keeps the regression model's feature distribution stable while retaining rows that would otherwise be dropped.

  • ✗

    Apply label encoding to the neighborhood variable

    Why it's wrong here

    Label encoding implies ordinality; neighborhood has no inherent order.

  • ✓

    Apply one-hot encoding to the neighborhood variable

    Why this is correct

    Neighbourhood is categorical with no ordinal relationship, so one-hot encoding creates a binary column per category, letting the regression model treat each neighbourhood independently rather than implying a false numeric ordering that label encoding would introduce.

About these practice questions

This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.