AI0-001 AI Concepts and Techniques Practice Question
A data scientist is evaluating a binary classifier for a medical diagnosis task. The dataset is imbalanced with 5% positive cases. Which THREE metrics should the data scientist consider for a comprehensive evaluation?
⚠ Common exam trap
Test-takers frequently default to accuracy as a universal metric, but CompTIA AI tests the understanding that accuracy is unreliable for imbalanced datasets, and that metrics like precision, recall, and F1 score are required for a comprehensive evaluation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Precision
Precision (A) is correct because it measures the proportion of predicted positives that are truly positive, which is critical in an imbalanced medical diagnosis task where false positives carry real cost. Recall (E) is correct because it measures the proportion of actual positives that are correctly identified, ensuring the classifier does not miss the rare 5% positive cases. F1 score (B) is correct because it is the harmonic mean of precision and recall, providing a single balanced metric that is far more informative than accuracy on skewed class distributions. Accuracy (C) is not appropriate here because a trivial model predicting all negatives would achieve 95% accuracy while detecting zero positive cases. Perplexity (D) does not belong because it is a language-model metric measuring how well a probability distribution predicts a sample, not a classification performance measure.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Precision
Why this is correct
Precision measures the proportion of predicted positives that are truly positive, exposing how many false alarms the classifier raises. At a 5% positive rate, this matters because accuracy alone would look high while missing poor positive-class performance.
- ✓
F1 score
Why this is correct
The F1 score is the harmonic mean of precision and recall, giving a single balanced figure that penalises poor performance on either. For a 5% positive class, it prevents a model from appearing strong merely by favouring the majority negative class.
- ✗
Accuracy
Why it's wrong here
Accuracy reports the proportion of all predictions that are correct, so with 5% positives a model predicting every case as negative scores 95% while detecting no disease — it cannot expose missed positives. It is tempting because accuracy suits balanced datasets where classes carry equal cost, but here recall, precision and F1 are needed.
- ✗
Perplexity
Why it's wrong here
Perplexity measures how well a language model predicts token sequences, so it cannot assess a binary classifier's thresholded predictions. It is the right metric when evaluating generative or probabilistic language models, not classification on imbalanced tabular data.
- ✓
Recall
Why this is correct
Recall measures the proportion of actual positives correctly identified, which is critical in medical diagnosis where missing a true case is costly. With only 5% positives, high accuracy can coexist with low recall, so this metric exposes missed diagnoses.
About these practice questions
One of 962 original AI0-001 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.