mediumMultiple Select
AIF-C01 Practice Question: A data scientist is building a regression model…
A data scientist is building a regression model to predict house prices. During feature engineering, they have categorical variables (e.g., neighborhood) and numerical variables (e.g., square footage) with missing values. Which TWO actions should the data scientist take? (Choose two.)
⚠ Common exam trap
In the AWS AI Practitioner exam, a common trap is that candidates may incorrectly choose label encoding (Option D) because it seems simpler, without recognizing that it introduces an arbitrary ordinal relationship that degrades model performance.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Impute missing square footage values with the median
Option C is correct because imputing missing square footage values with the median is a robust strategy for numerical features, especially when outliers may be present, and it preserves the row count for model training. Option E is correct because one-hot encoding is the appropriate technique for nominal categorical variables like neighborhood, as it creates binary columns without imposing an artificial ordinal relationship. Option A is not ideal because dropping all rows with any missing values can discard substantial useful data and introduce bias. Option B is not required for regression models and may not be necessary unless the algorithm is sensitive to feature scale. Option D is incorrect because label encoding would impose an arbitrary order on neighborhood categories, which is inappropriate for a nominal variable.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Drop all rows with any missing values
Why it's wrong here
Dropping rows reduces data size; imputation is generally preferred unless missingness is excessive.
- ✗
Normalize all numerical features to [0,1] range
Why it's wrong here
Normalization is not always required; tree-based models are scale-invariant.
- ✓
Impute missing square footage values with the median
Why this is correct
Square footage is numerical and skewed by outliers, so median imputation preserves the central tendency without being distorted by extreme values, unlike the mean. This keeps the regression model's feature distribution stable while retaining rows that would otherwise be dropped.
- ✗
Apply label encoding to the neighborhood variable
Why it's wrong here
Label encoding implies ordinality; neighborhood has no inherent order.
- ✓
Apply one-hot encoding to the neighborhood variable
Why this is correct
Neighbourhood is categorical with no ordinal relationship, so one-hot encoding creates a binary column per category, letting the regression model treat each neighbourhood independently rather than implying a false numeric ordering that label encoding would introduce.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.