Courseiva

AI0-001 Machine Learning and Deep Learning Practice Question

A data scientist is training a binary classification model to detect fraudulent transactions. The dataset contains 99.9% legitimate transactions and 0.1% fraudulent transactions. After training a logistic regression model, the accuracy is 99.9%, but the recall for the fraud class is 0%. Which of the following is the MOST likely cause?

⚠ Common exam trap

CompTIA often tests the 'accuracy paradox' where candidates mistakenly attribute high accuracy to model quality, ignoring that in imbalanced datasets, a dummy classifier predicting the majority class can achieve the same accuracy, and the trap is to overlook recall or precision for the minority class.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The dataset is highly imbalanced, and the model predicts the majority class for all instances.

The dataset has a severe class imbalance (99.9% legitimate, 0.1% fraudulent). A logistic regression model that predicts the majority class (legitimate) for every instance will achieve 99.9% accuracy but 0% recall for the fraud class, because it never identifies any positive fraud cases. This is the classic 'accuracy paradox' in imbalanced classification.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The regularization parameter is too large, causing underfitting.

    Why it's wrong here

    Large L2 regularization shrinks coefficients toward zero, but that yields uniformly poor fits across both classes, not 99.9% accuracy with zero fraud recall. Regularisation is tuned to curb variance when a model overfits noisy features; here the failure is class imbalance, which needs resampling or class weights.

  • ✗

    The model is overfitting due to too many features.

    Why it's wrong here

    Overfitting produces high training accuracy and degraded validation performance, not a model that predicts the majority class perfectly on both. Feature-count reduction addresses variance when training and test metrics diverge; this stem shows a degenerate majority-class predictor caused by the 99.9/0.1 imbalance.

  • ✗

    The learning rate was too high.

    Why it's wrong here

    A high learning rate causes divergence or oscillation, not a model that predicts the majority class for every row; the 99.9% accuracy with 0% fraud recall reflects class imbalance, where the loss is dominated by legitimate examples. It tempts because learning-rate faults do destabilise training, but they would degrade accuracy too.

  • ✓

    The dataset is highly imbalanced, and the model predicts the majority class for all instances.

    Why this is correct

    With 0.1% fraud, a model predicting legitimate for every transaction scores 99.9% accuracy yet catches no fraud, giving 0% recall. The extreme class imbalance dominates the loss function, so the model never learns the minority pattern. Resampling or class weighting is required.

About these practice questions

One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.