mediumMultiple Choice
MLA-C01 Practice Question: A data scientist is building a binary…
A data scientist is building a binary classification model on a highly imbalanced dataset where the positive class represents only 1% of the data. The scientist needs to train the model using Amazon SageMaker's built-in XGBoost algorithm. Which strategy should be used to address the class imbalance?
⚠ Common exam trap
It's easy for candidates to confuse `scale_pos_weight` with resampling techniques (like SMOTE or undersampling) or with other hyperparameters like `max_delta_step`, assuming any imbalance-handling method is equally valid, but the exam expects knowledge of the specific built-in mechanism for SageMaker's XGBoost.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set the `scale_pos_weight` hyperparameter to `sum(negative cases) / sum(positive cases)`
XGBoost's `scale_pos_weight` hyperparameter is specifically designed to handle class imbalance by adjusting the weight of the positive class during training. Setting it to `sum(negative cases) / sum(positive cases)` (i.e., 99/1 = 99) tells the algorithm to penalize misclassifications of the minority class more heavily, effectively balancing the gradient updates. This is the recommended approach for built-in XGBoost in SageMaker, as it directly modifies the loss function without altering the dataset.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Undersample the majority class until the dataset is balanced
Why it's wrong here
Undersampling discards most majority-class records, leaving roughly 2% of the original data and causing severe information loss and overfitting. It is tempting because balancing class counts is intuitive, and it would be correct when the majority class is small enough that discarding rows is affordable.
- ✓
Set the `scale_pos_weight` hyperparameter to `sum(negative cases) / sum(positive cases)`
Why this is correct
Setting `scale_pos_weight` to the negative-to-positive ratio (roughly 99) reweights the positive class's gradient contributions during boosting, so the 1% minority class influences tree splits proportionally to its cost. This directly counteracts the imbalance in SageMaker's built-in XGBoost without resampling the training data.
- ✗
Use the `max_delta_step` hyperparameter to increase the learning rate for the majority class
Why it's wrong here
max_delta_step limits the maximum step size for leaf weight updates to stabilise logistic training; it does not reweight classes or raise the majority class's learning rate, so imbalance persists. It is tempting because it aids convergence on skewed data, and would be correct for numerical stability rather than correcting class imbalance.
- ✗
Use SMOTE to oversample the minority class before passing the data to XGBoost
Why it's wrong here
SMOTE oversamples the minority class before training, but XGBoost’s built-in `scale_pos_weight` parameter directly adjusts the loss function gradient to penalise misclassifications of the minority class during training, making external resampling redundant and potentially introducing synthetic noise. This option is tempting because SMOTE is a standard technique for balancing datasets in general machine learning workflows, and it would be correct if the algorithm lacked native imbalance handling—for example, when using a linear classifier without class-weight support.
Go deeper
Related to this question
About these practice questions
One of 665 original MLA-C01 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This MLA-C01 practice question is part of Courseiva's free Amazon Web Services certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the MLA-C01 exam.