mediumMultiple Select
MLA-C01 Practice Question: A data scientist is training a binary…
A data scientist is training a binary classification model using Amazon SageMaker. The dataset is highly imbalanced (95% negative class, 5% positive class). The model is evaluated on a held-out test set, and the F1 score is 0.12. The data scientist wants to improve the F1 score. Which two actions should the data scientist take? (Choose two.)
⚠ Common exam trap
Many candidates confuse threshold tuning with addressing imbalance directly, not realizing that adjusting the threshold without rebalancing the data or weighting classes typically fails to improve F1 score because it does not change the underlying model's learned distribution.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply SMOTE (Synthetic Minority Oversampling Technique) to the training data using a preprocessing script in SageMaker Processing.
Option B is correct because SMOTE generates synthetic examples of the minority (positive) class, rebalancing the 95/5 training distribution so the model learns the minority class decision boundary instead of being dominated by the negative class, which directly improves precision and recall (and thus F1) on the held-out test set; this can be implemented via a SageMaker Processing job with an imbalanced-learn script. Option E is correct because setting scale_pos_weight in the SageMaker XGBoost estimator to the ratio of negative to positive samples (roughly 95/5 = 19) upweights the minority class in the gradient boosting loss, making the model penalize minority-class misclassifications more heavily and improving F1 without altering the data. Option A is not appropriate because reducing neural network depth addresses overfitting, not class imbalance, and could even underfit and lower F1. Option C is wrong because raising the decision threshold reduces recall and typically lowers F1 in a heavily imbalanced setting. Option D is wrong because switching the evaluation metric to recall does not improve the model's F1 score; it only changes what is reported.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Reduce the model complexity by decreasing the number of layers in a deep neural network.
Why it's wrong here
Reducing layers addresses variance, not the 95:5 class imbalance driving the 0.12 F1 score; the model still predicts the majority class. It is tempting because pruning capacity genuinely helps when a network overfits a balanced dataset, but here the loss function and decision threshold need reweighting or resampling instead.
- ✓
Apply SMOTE (Synthetic Minority Oversampling Technique) to the training data using a preprocessing script in SageMaker Processing.
Why this is correct
SMOTE synthesises new minority-class examples by interpolating between existing positive samples and their nearest neighbours, rebalancing the training distribution. With only 5% positives, the model otherwise predicts the majority class, so oversampling directly lifts recall and therefore F1 on the held-out test set.
- ✗
Increase the decision threshold to reduce false positives.
Why it's wrong here
Raising the threshold suppresses positive predictions, cutting recall on a 5% minority class and lowering F1 further. Threshold tuning suits cost-weighted decisions where precision and recall trade off deliberately, not a model whose positive class is already under-predicted.
- ✗
Use recall as the primary evaluation metric instead of F1.
Why it's wrong here
Swapping the metric does not change the model's predictions or the F1 score; it only stops measuring it. Recall alone is the right focus when missing positives is the dominant cost and false positives are tolerable, not when both errors matter.
- ✓
Set the `scale_pos_weight` parameter in the SageMaker XGBoost estimator to the ratio of negative to positive samples.
Why this is correct
scale_pos_weight multiplies the positive class's gradient contribution by the negative-to-positive ratio, so XGBoost penalises minority-class errors more heavily during boosting. This reweights the loss without altering the data, directly countering the 95/5 imbalance that suppressed F1.
Go deeper
Related to this question
About these practice questions
This MLA-C01 question is part of Courseiva's 665-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.