AI0-001 AI Concepts and Techniques Practice Question
A data scientist is building a model to predict whether a transaction is fraudulent. The dataset has 99.9% legitimate transactions and 0.1% fraudulent ones. Which evaluation metric is MOST appropriate to assess model performance given this class imbalance?
⚠ Common exam trap
A common trap is that candidates default to accuracy as the universal metric, failing to recognize that in extreme class imbalance (e.g., 99.9% vs 0.1%), accuracy becomes meaningless and F1-score is the standard alternative.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
F1-score
With 99.9% legitimate transactions and only 0.1% fraudulent ones, accuracy would be misleadingly high (99.9%) even if the model never predicts fraud. The F1-score is the harmonic mean of precision and recall, making it robust to class imbalance by penalizing both false positives and false negatives. This makes it the most appropriate metric for evaluating fraud detection performance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
BLEU score
Why it's wrong here
BLEU compares n-gram overlap between generated and reference text, so it cannot consume the binary fraud labels this classifier produces. It is tempting because it is the standard metric for machine translation and text summarisation, where output is free-form prose rather than a class prediction.
- ✗
Accuracy
Why it's wrong here
With 99.9% legitimate transactions, a model predicting every transaction as legitimate scores 99.9% accuracy while catching zero fraud, so accuracy hides the failure that matters. It is tempting because it is the default metric for balanced classification, where classes occur at roughly equal rates.
- ✓
F1-score
Why this is correct
F1-score combines precision and recall into a single harmonic mean, so it penalises models that ignore the 0.1% fraudulent minority. Unlike accuracy, which reaches 99.9% by predicting "legitimate" always, F1-score reflects performance on the positive class, satisfying the stem's class-imbalance constraint.
- ✗
Perplexity
Why it's wrong here
Perplexity measures how well a language model predicts the next token in a sequence, so it cannot score a binary fraud classifier. It is tempting because it is the standard intrinsic metric for evaluating language models, where lower values indicate the model assigns higher probability to held-out text.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.