AI0-001 Machine Learning and Deep Learning Practice Question
A retail analytics team has a labeled dataset of 50,000 customer transactions where each record is tagged as either 'fraudulent' or 'legitimate.' They need a supervised learning approach that outputs a probability between 0 and 1 for the fraudulent class so it can be compared against a business threshold. Which algorithm is most appropriate for this task?
⚠ Common exam trap
The trap here is assuming any algorithm that groups or summarizes data can classify labeled records, when supervised classification requires a model that learns from the target label and emits a class probability.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Logistic regression
The task is supervised binary classification with a need for a 0-to-1 probability, which points to logistic regression because its sigmoid output is directly comparable to a decision threshold. The other techniques either lack supervision (K-means, Apriori) or do not model the target label (PCA). Logistic regression also scales well to 50,000 records and provides coefficients that help explain which transaction features drive fraud risk.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
K-means clustering
Why it's wrong here
K-means is an unsupervised algorithm that partitions unlabeled data into clusters based on distance to centroids. It does not use the fraud/legitimate labels and produces cluster assignments rather than class probabilities. Since the scenario explicitly has labeled records and requires a probability for a binary class, K-means cannot deliver the required calibrated output and would discard the available supervision signal entirely.
- ✗
Apriori association rule mining
Why it's wrong here
Apriori discovers frequent itemsets and association rules such as 'customers who buy X also buy Y.' It operates on transaction baskets without a target label and returns rules with support and confidence, not per-record class probabilities. The fraud detection requirement is a supervised binary classification problem, so association rule mining does not address the stated output or the labeled data available.
- ✗
Principal component analysis
Why it's wrong here
PCA is a dimensionality-reduction technique that projects features onto orthogonal components maximizing variance. It does not model a target label and produces no probability of fraud. While PCA could be used as a preprocessing step before a classifier, it cannot by itself output the 0-to-1 fraud probability the team needs, so it is the wrong tool for this supervised prediction requirement.
- ✓
Logistic regression
Why this is correct
Logistic regression is a supervised binary classifier that applies a sigmoid function to a linear combination of features, producing a probability between 0 and 1 for the positive class. This exactly matches the requirement to compare the fraud probability against a threshold. It also trains efficiently on 50,000 labeled records and yields interpretable coefficients, which is valuable for explaining fraud decisions to stakeholders.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.