A financial institution is training a fraud detection model using SageMaker. The dataset is highly imbalanced, with only 0.1% fraudulent transactions. The team wants to use SageMaker Automatic Model Tuning to find the best hyperparameters. They notice that the tuning job spends most of its time on configurations that predict all transactions as non-fraudulent. Which hyperparameter should they tune to directly address this issue?
In SageMaker's built-in XGBoost algorithm, scale_pos_weight controls the balance of positive and negative weights. Setting it to a higher value increases the weight of the positive class (fraudulent transactions), making the model pay more attention to them. This directly addresses the issue of the model predicting all transactions as non-fraudulent, as it penalizes misclassification of the minority class more heavily.
Why this answer
The scale_pos_weight hyperparameter in SageMaker's XGBoost algorithm adjusts the weight of the positive class, which is crucial for imbalanced datasets. By increasing this value, the model's loss function penalizes false negatives more, encouraging better detection of fraudulent transactions. This directly tackles the problem of the model predicting all instances as the majority class.
Exam trap
The trap here is focusing on general regularization or optimization hyperparameters, when the core issue is class imbalance that requires a weighting adjustment.