Handling Class Imbalance in Machine Learning
A fraud detection model is trained on a dataset where only 0.1% of transactions are fraudulent. The model achieves 99.9% accuracy but fails to catch most frauds. Which metric should the team prioritize, and which technique could help?
⚠ Common exam trap
Test-takers often mistakenly believe that high accuracy always indicates a good model, and that techniques like PCA or regularization can fix class imbalance. In reality, only metrics and resampling methods designed for skewed distributions are effective.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Precision-Recall AUC; use oversampling like SMOTE
The dataset is highly imbalanced (0.1% fraud), so 99.9% accuracy is misleading because a model that predicts 'not fraud' for every transaction achieves it. Precision-Recall AUC focuses on the positive class (fraud) and is robust to class imbalance, unlike accuracy or ROC-AUC. Oversampling like SMOTE generates synthetic fraud samples to balance the dataset, helping the model learn the minority class patterns.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Mean Squared Error; use L2 regularization
Why it's wrong here
MSE suits regression, not classification, and L2 regularization shrinks weights without rebalancing the 0.1% fraud class, so the model still predicts the majority class. It tempts because MSE with L2 is standard for continuous prediction tasks, where it would be the right pairing.
- ✗
F1 score; use principal component analysis
Why it's wrong here
F1 score is the right metric for this imbalance, but principal component analysis is dimensionality reduction and does nothing for class imbalance. Resampling, class weighting, or anomaly detection techniques are needed. PCA would suit a wide feature set with correlated variables, not a skewed label distribution.
- ✗
Accuracy; collect more data
Why it's wrong here
Accuracy is precisely the misleading metric here: predicting every transaction as legitimate yields 99.9%, so collecting more data preserves the same class imbalance and the same blind spot. Accuracy is the correct metric only when classes are roughly balanced.
- ✓
Precision-Recall AUC; use oversampling like SMOTE
Why this is correct
With 0.1% positives, accuracy is misleading because predicting all negatives scores 99.9%. Precision-Recall AUC focuses on the minority fraud class, and SMOTE synthetically oversamples it, directly addressing the severe class imbalance that causes the model to miss most frauds.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on AI0-001
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. A financial institution uses a deep learning model for fraud detection. The model is a feedforward neural network with three hidden layers. It was trained on a balanced dataset of 100,000 transactions. During deployment, the model achieves high accuracy on the test set but the fraud detection rate (true positive rate) is only 40% while the false positive rate is 0.1%. The business requires a true positive rate of at least 80%. Which of the following actions is most likely to achieve the required true positive rate while minimizing the increase in false positives?
hard- A.Increase the number of hidden layers to five to capture more complex patterns
- B.Use synthetic minority oversampling (SMOTE) to rebalance the training set
- ✓ C.Change the threshold for classifying a transaction as fraud from the default 0.5 to a lower value
- D.Add L2 regularization to reduce overfitting
Why C: (increase hidden layers) may capture more complexity but does not directly increase TPR and could overfit. Option B (SMOTE) rebalances the training set, but the dataset is already balanced, so this is unlikely to improve TPR. Option D (L2 regularization) reduces overfitting but increases bias, which could lower TPR. Option C (change threshold) is the most direct approach: lowering the classification threshold increases the true positive rate, and by tuning, it can achieve 80% TPR with a minimal increase in false positives.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.