mediumMultiple Choice
AIF-C01 Practice Question: A machine learning practitioner is training a…
A machine learning practitioner is training a model to forecast product demand and observes that the model performs well on training data but poorly on unseen data. Which of the following is the MOST likely cause?
⚠ Common exam trap
The AWS AI Practitioner exam often tests the distinction between overfitting and underfitting by describing performance on training vs. unseen data; the trap here is that candidates may confuse 'high bias' (underfitting) with 'high variance' (overfitting) when the model performs well on training data but poorly on test data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Overfitting due to excessive model complexity
The model performs well on training data but poorly on unseen data, which is the classic symptom of overfitting. Overfitting occurs when the model is excessively complex (e.g., too many parameters, deep neural networks with high capacity) and learns noise and random fluctuations in the training data rather than the underlying pattern. This leads to high variance, causing excellent training performance but poor generalization to new data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
High bias-variance tradeoff favoring bias
Why it's wrong here
High bias yields consistent underperformance on training and test sets alike, not the strong-training, weak-test split described. Bias-dominated models are chosen when the algorithm is too constrained for the data's complexity; the stem's gap instead indicates variance, where the model fits noise.
- ✗
Underfitting due to model being too simple
Why it's wrong here
Underfitting produces poor results on both training and unseen data, since a too-simple model cannot capture the underlying pattern. It is the correct diagnosis when training error itself is high; here training performance is strong, indicating the model has memorised rather than generalised.
- ✗
Data leakage causing artificially high training scores
Why it's wrong here
Leakage inflates training scores, but it typically also lifts validation performance when the leaked feature persists at inference; the described pattern is classic overfitting. Leakage is the answer when a feature encodes the target, such as including future sales figures in demand forecasting inputs.
- ✓
Overfitting due to excessive model complexity
Why this is correct
Excessive model complexity lets the network memorise training examples, including noise, so training performance stays high while generalisation to unseen demand data degrades. The gap between strong training results and poor unseen-data results directly indicates overfitting rather than underfitting or data leakage.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.