AIF-C01 Fundamentals of AI and ML Practice Question
Which TWO techniques are commonly used to prevent overfitting in machine learning models? (Select TWO.)
⚠ Common exam trap
AWS often tests the misconception that adding more data or features always helps model performance, when in fact irrelevant features or reducing training data can worsen overfitting, and candidates may incorrectly associate 'more complexity' with better generalization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use cross-validation
Option B (Use cross-validation) is correct because techniques like k-fold cross-validation estimate how well a model generalizes to unseen data by training and validating on different data splits, which helps detect and mitigate overfitting during model selection and hyperparameter tuning. Option E (Use regularization) is correct because methods such as L1 (Lasso) and L2 (Ridge) regularization add a penalty term to the loss function that constrains the magnitude of model weights, reducing variance and preventing the model from fitting noise in the training data. The unmarked options do not belong: adding more irrelevant features (A) increases dimensionality and noise, making overfitting worse; increasing model complexity (C) raises variance and typically worsens overfitting; and reducing the amount of training data (D) gives the model less information to learn generalizable patterns, also promoting overfitting.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Add more irrelevant features
Why it's wrong here
Adding irrelevant features increases dimensionality without signal, giving the model more spurious correlations to memorise and worsening overfitting. Feature selection or dimensionality reduction removes such noise; this option does the opposite of the required effect.
- ✓
Use cross-validation
Why this is correct
Cross-validation partitions the data into folds, training and validating on different subsets, so performance estimates reflect generalisation rather than memorisation of one split. It detects overfitting by exposing the gap between training and held-out fold error.
- ✗
Increase model complexity
Why it's wrong here
Increasing model complexity, such as adding more layers or parameters, directly increases variance and exacerbates overfitting by allowing the model to memorise training noise rather than generalising patterns. This option is tempting because, in scenarios where the model is underfitting (high bias), raising complexity can improve performance on training data, but it fails here because the stem specifically asks for techniques that *prevent* overfitting, not those that cause it.
- ✗
Reduce the amount of training data
Why it's wrong here
Reducing training data removes the examples a model needs to generalise, so it fits the remaining sample more tightly and overfits more. Regularisation, cross-validation or early stopping are used instead; more representative data is the remedy, not less.
- ✓
Use regularization
Why this is correct
Regularization adds a penalty to the loss function to limit model complexity.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AIF-C01 question from scratch — 862 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.