A team is building a classifier to detect fraudulent transactions. The dataset has 99.9% legitimate transactions and 0.1% fraudulent. Which evaluation metric is most appropriate?
With only 0.1% fraud, accuracy is misleading because predicting all legitimate scores 99.9%. AUC-ROC measures ranking performance across all thresholds, remaining informative under such extreme class imbalance, so it distinguishes fraudulent from legitimate transactions regardless of the skewed prior.
Why this answer
AUC-ROC is the most appropriate metric for this highly imbalanced classification problem because it evaluates the model's ability to distinguish between fraudulent (positive) and legitimate (negative) transactions across all classification thresholds, without being biased by the overwhelming majority of legitimate transactions. Unlike accuracy, AUC-ROC focuses on the trade-off between true positive rate and false positive rate, making it robust to class imbalance where 99.9% of data is negative.
Exam trap
The AWS AI Practitioner exam often tests the misconception that accuracy is always a good metric, but the trap here is that accuracy is dangerously misleading for imbalanced datasets, and candidates must recognize that AUC-ROC or precision-recall metrics are required for such scenarios.
How to eliminate wrong answers
Option A is wrong because Mean Absolute Error (MAE) is a regression metric that measures average absolute differences between predicted and actual continuous values, not suitable for binary classification tasks like fraud detection. Option B is wrong because Root Mean Squared Error (RMSE) is also a regression metric that penalizes larger errors more heavily, and it cannot evaluate classification performance on imbalanced datasets. Option D is wrong because accuracy would be 99.9% even if the model predicts all transactions as legitimate, completely failing to detect any fraud; it is misleading for imbalanced datasets where the minority class is the one of interest.