mediumMultiple Choice
AIF-C01 Practice Question: A data scientist is building a model to predict…
A data scientist is building a model to predict housing prices. The dataset contains features such as square footage, number of bedrooms, and location. After training a linear regression model, the RMSE on the test set is significantly higher than on the training set. What is the MOST likely cause?
⚠ Common exam trap
In the AWS AI Practitioner exam, this scenario tests the distinction between overfitting and underfitting. A large gap between test and training RMSE indicates overfitting, not data leakage or insufficient data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Overfitting
Overfitting. When RMSE on the test set is significantly higher than on the training set, it indicates the model has learned noise and specific patterns from the training data that do not generalize to unseen data. This is the classic symptom of overfitting in linear regression, where the model may have too many features or excessive complexity relative to the data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Data leakage from test to training set
Why it's wrong here
Leakage from test into training typically inflates test performance, making test error artificially low rather than high. The observed gap points to the model memorising training data. Leakage is the correct diagnosis when validation scores look implausibly good compared with deployment.
- ✓
Overfitting
Why this is correct
A large gap between low training error and high test RMSE indicates the model has memorised training noise rather than the underlying relationship. Overfitting is therefore the most likely cause, since the model fits training data well but fails to generalise to unseen housing records.
- ✗
Underfitting
Why it's wrong here
Underfitting produces high error on both training and test sets, since the model fails to capture the underlying pattern. Here training error is low while test error is high, indicating variance. Underfitting would be the diagnosis when both errors remain large after training.
- ✗
Insufficient training data
Why it's wrong here
Insufficient training data usually raises training error too, or affects both sets similarly; it does not by itself create a wide train-test gap. The pattern shown reflects variance from an over-complex model. Insufficient data is the answer when the model cannot learn the pattern at all.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.