AI0-001 Machine Learning and Deep Learning Practice Question
A data scientist is building a model to predict the likelihood of a patient having a rare disease. The dataset is highly imbalanced, with only 2% of patients having the disease. The data scientist trains a logistic regression model and achieves 98% accuracy, but the model predicts 'no disease' for all patients. Which evaluation metric should the data scientist use to better assess the model's performance?
⚠ Common exam trap
The trap here is relying on accuracy as the primary metric for imbalanced datasets, which can hide poor performance on the minority class.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Precision-Recall AUC (Area Under the Curve)
Precision-Recall AUC is designed for imbalanced classification problems, focusing on the performance of the positive class. It provides a more informative picture than accuracy, which can be misleadingly high when the negative class dominates. This metric helps the data scientist understand the trade-off between precision and recall for the rare disease detection.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Mean Squared Error (MSE)
Why it's wrong here
MSE is a regression metric that measures the average squared difference between predicted and actual values. It is not applicable to classification problems like predicting disease presence. Using MSE would be a category error and would not provide meaningful evaluation of the classification model's performance on imbalanced data.
- ✓
Precision-Recall AUC (Area Under the Curve)
Why this is correct
Precision-Recall AUC is particularly useful for imbalanced datasets because it focuses on the positive class. It plots precision against recall at various thresholds, providing a comprehensive view of the model's ability to identify the rare disease without being overwhelmed by the large number of true negatives. This metric helps assess how well the model distinguishes the minority class.
- ✗
Accuracy
Why it's wrong here
Accuracy is the ratio of correct predictions to total predictions. In this imbalanced dataset, a model that always predicts 'no disease' achieves 98% accuracy because 98% of patients do not have the disease. However, this model is useless for identifying patients with the disease. Accuracy is misleading for imbalanced datasets and should not be used to assess performance in this scenario.
- ✗
R-squared (R²)
Why it's wrong here
R-squared is a regression metric that indicates the proportion of variance in the dependent variable explained by the model. It is not used for classification tasks. Applying R² to a binary classification problem is inappropriate and would not help in evaluating the model's ability to detect the rare disease.
About these practice questions
One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.