AI0-001 AI Concepts and Techniques Practice Question
A natural language processing team is building a system to classify support tickets into categories. They have a large corpus of unlabeled ticket text and a small set of manually labeled tickets. They want to leverage both to improve classification performance. Which approach is MOST suitable?
⚠ Common exam trap
The trap here is assuming that any use of unlabeled data is clustering, when semi-supervised learning can incorporate unlabeled data while still training a supervised classifier for predefined categories.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use semi-supervised learning that combines the unlabeled corpus with the labeled set
Semi-supervised learning is the best fit because it explicitly uses both the large unlabeled corpus and the small labeled set. Methods like self-training or consistency regularization can improve classification by learning from unlabeled tickets while still respecting the predefined categories. The other options either ignore one of the data sources or fail to produce a classifier aligned with the desired categories.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Apply unsupervised clustering to the unlabeled tickets and treat clusters as categories
Why it's wrong here
Clustering groups similar tickets but does not map them to the predefined support categories the team needs. The resulting clusters may not align with business-defined labels, and there is no guarantee that cluster boundaries correspond to ticket types. This approach ignores the labeled set entirely and does not produce a classifier for the desired categories, so it fails the requirement.
- ✗
Train a supervised classifier only on the small labeled set
Why it's wrong here
Training only on the small labeled set ignores the large unlabeled corpus and is likely to overfit, yielding poor generalization. The team explicitly wants to leverage both data sources to improve performance. While this approach is simple, it wastes the available unlabeled text and does not address the limited label problem, making it suboptimal for the stated goal.
- ✓
Use semi-supervised learning that combines the unlabeled corpus with the labeled set
Why this is correct
Semi-supervised learning is designed for exactly this situation: a large amount of unlabeled data plus a small labeled set. Techniques such as self-training, co-training, or consistency regularization can use the unlabeled tickets to learn better representations and improve classification. This approach directly leverages both data sources, matching the team's objective and the data availability.
- ✗
Perform principal component analysis on the ticket text and train a classifier on the components
Why it's wrong here
Principal component analysis is a dimensionality reduction technique that does not incorporate label information or leverage unlabeled data in a semi-supervised manner. It may reduce noise but does not address the core challenge of limited labels. Training a classifier on principal components still requires sufficient labeled data and does not exploit the unlabeled corpus as semi-supervised methods do.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.