AIF-C01 Fundamentals of AI and ML Practice Question
A data scientist trains a model to predict whether a loan applicant will default. After deployment, the model performs well on applicants similar to the training data but poorly on applicants from a newly added geographic region that was underrepresented in training. Which statement best describes the underlying problem?
⚠ Common exam trap
The trap here is labeling any train-versus-deployment gap as overfitting, when the failure is isolated to a subpopulation missing from the training distribution.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The training data does not represent the new region, so the model cannot generalize to that subpopulation.
The model fails specifically on a subpopulation that was underrepresented in training, which is a data coverage and distribution shift problem. Good performance on familiar applicants confirms the algorithm works, but without representative training examples from the new region the model cannot generalize there. The remedy is additional representative data or techniques that address distribution shift.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The training data does not represent the new region, so the model cannot generalize to that subpopulation.
Why this is correct
When a subpopulation is underrepresented or absent in training, the model learns patterns that do not transfer to it, producing poor predictions for that group. This is a data coverage and distribution shift problem, and it matches the observed failure on the newly added geographic region exactly.
- ✗
The model has overfit the training set because it memorized noise in the original regions.
Why it's wrong here
Overfitting means strong training performance but weak test performance on data drawn from the same distribution. Here the model performs well on data resembling training but poorly on a new region, which points to a distribution difference rather than memorization of noise. Overfitting is not the primary issue described.
- ✗
The model requires more training epochs to converge on the new region's data.
Why it's wrong here
Additional epochs on the same data would not teach the model about a region it has barely seen; it would only reinforce existing patterns or worsen overfitting. The gap stems from missing representation of the new region, not insufficient training time, so more epochs do not address the root cause.
- ✗
The evaluation metric used is inappropriate and should be replaced with accuracy.
Why it's wrong here
Changing the metric does not change the model's predictions; it only changes how they are summarized. The model genuinely predicts poorly for the new region, so switching metrics would mask rather than fix the issue. The problem is data representation, not measurement.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Amazon Web Services exam blueprint
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.