hardMultiple Select
AIF-C01 Practice Question: A machine learning team is building a binary…
A machine learning team is building a binary classifier using Amazon SageMaker. The dataset has 10,000 features and 1,000 samples. The model overfits severely. Which TWO approaches are MOST likely to reduce overfitting? (Choose two.)
⚠ Common exam trap
AWS often tests the misconception that increasing batch size or training longer always improves generalization, when in fact these techniques can worsen overfitting in high-dimensional, low-sample scenarios.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Perform feature selection to reduce the number of features
Option C is correct because with 10,000 features but only 1,000 samples, the model has far more parameters than data points, so performing feature selection to reduce the number of features directly lowers model complexity and the variance that causes overfitting. Option D is correct because adding L2 regularization to the loss function penalizes large weights, shrinking them toward zero and constraining the model's effective capacity, which is a standard, effective remedy for overfitting. Options A, B, and E do not belong: increasing batch size to the full dataset mainly affects gradient noise and optimization dynamics rather than reducing model capacity, adding more layers increases model complexity and would worsen overfitting, and training for more epochs lets the model fit the training data even more closely, also worsening overfitting.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the batch size to the full dataset
Why it's wrong here
Full-batch training reduces gradient noise, letting the model converge tightly onto the 1,000 training samples and overfit further. Large batches are tempting because they speed training and stabilise updates, and would be correct for throughput on large datasets, not for regularisation.
- ✗
Use a neural network with more layers
Why it's wrong here
Adding layers increases model capacity, which with 10,000 features and only 1,000 samples deepens the overfitting. Deeper networks are tempting because they capture complex patterns, and would be correct when the model underfits and needs greater representational power.
- ✓
Perform feature selection to reduce the number of features
Why this is correct
With 10,000 features against only 1,000 samples, the model has far more parameters than observations, so it memorises noise. Feature selection removes irrelevant or redundant predictors, directly reducing the dimensionality that drives the severe overfitting described in the stem.
- ✓
Add L2 regularization to the loss function
Why this is correct
L2 regularization adds a penalty on squared weights to the loss function, shrinking coefficients and constraining model complexity. With 10,000 features and 1,000 samples, this directly limits the overfitting described in the stem without removing features.
- ✗
Train the model for more epochs
Why it's wrong here
More epochs let the model continue fitting the 1,000 training samples, worsening overfitting rather than reducing it. Training longer is tempting because it helps underfit models converge, and would be correct where training loss remains high and validation performance is still improving.
Go deeper
Related to this question
About these practice questions
This AIF-C01 question is part of Courseiva's 862-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AIF-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AIF-C01 exam.