A financial services company operates a real-time inference endpoint for a fraud detection model on Amazon SageMaker. The model was trained on historical transaction data from 2023. Over the past month, the model's precision has dropped from 92% to 78%, while recall remains high at 95%. The data science team suspects data drift and has already enabled SageMaker Model Monitor with data capture and a baseline from the training data. The latest monitoring report indicates no statistically significant drift in any of the input features. The team also verified that the inference code and model artifact have not changed. Despite the stable feature distributions, the model is misclassifying an increasing number of legitimate transactions as fraudulent (false positives). The business is concerned about the impact on customer experience. What is the best course of action?
Label drift occurs when the underlying relationship between features and labels changes. Collecting and analyzing recent labels can confirm if the fraud criteria have shifted.
Why this answer
The model's precision dropped while recall remained high, meaning it is flagging more legitimate transactions as fraudulent (false positives). Since input feature distributions show no drift and the model artifact is unchanged, the issue likely lies in the target variable—label drift or a change in the definition of fraud. Investigating recent ground truth labels will reveal if the labels used for evaluation or retraining have shifted, which would explain the degradation.
This is the best course of action before retraining or changing the model.
Exam trap
MLA-C01 often tests the distinction between feature drift and label drift, so candidates must remember that stable feature distributions do not rule out label drift, which requires investigating ground truth labels.
How to eliminate wrong answers
Option A is wrong because replacing the algorithm does not address the root cause, which is likely label drift, and could introduce new issues. Option B is wrong because retraining with recent data without understanding the label shift may perpetuate the problem if the labels themselves are inconsistent. Option C is wrong because increasing data capture sampling does not address the cause of false positives; the monitoring already shows no feature drift, so more data capture is unnecessary.