Evaluate Custom Text Classification Model Performance Using Test Set and Confusion Matrix
You are deploying an Azure AI Language custom text classification model. You need to ensure the model meets performance requirements before promoting it to production. Which two actions should you take? (Choose two.)
Quick Answer
The correct answer is to evaluate the model using a held-out test set and review the confusion matrix to identify frequently misclassified classes. This is because a held-out test set provides an unbiased estimate of real-world performance, while the confusion matrix reveals which specific classes the model confuses, allowing you to target improvements. On the Microsoft Azure AI Engineer Associate AI-102 exam, this tests your understanding of model evaluation best practices, often appearing as a trap where candidates mistakenly choose to evaluate on the training set, which overestimates accuracy. Remember that precision and recall are more informative than overall accuracy, and comparing to a baseline is good practice but not a required step before promotion. A helpful memory tip is "Test and Confuse": always test on unseen data, then use the confusion matrix to diagnose confusion between classes.
⚠ Common exam trap
It's easy for candidates to assume a fixed accuracy threshold (like 95%) is required for production promotion, but Microsoft Azure AI Language custom text classification does not mandate any specific metric value—the focus is on evaluating generalization via a held-out test set and analyzing misclassifications with the confusion matrix.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Evaluate the model on a held-out test set that was not used during training.
Option A is correct because evaluating the custom text classification model on a held-out test set that was never used during training provides an unbiased estimate of generalization performance, which is essential before promoting the model to production. Option B is correct because reviewing the confusion matrix reveals per-class precision/recall patterns and shows exactly which classes are frequently misclassified, enabling targeted data or label improvements. Option C is incorrect because Azure AI Language does not require a fixed 95% accuracy threshold on a cross-validation split; performance targets are scenario-specific and cross-validation is not the standard evaluation workflow for this service. Option D is incorrect because computing accuracy on the training set measures memorization rather than generalization and will be optimistically biased. Option E is incorrect because comparing to a majority-class baseline is a useful sanity check but is not one of the required actions for validating a custom text classification model before production.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Evaluate the model on a held-out test set that was not used during training.
Why this is correct
Evaluating on a held-out test set never seen during training yields unbiased precision, recall and F1 scores, confirming the model generalises rather than memorising training data. This validation gate must pass before promotion to production, satisfying the performance-verification requirement.
- ✓
Review the confusion matrix to understand which classes are frequently misclassified.
Why this is correct
The confusion matrix exposes per-class precision and recall, revealing which categories the model confuses most often. Reviewing it directly satisfies the stem's requirement to verify performance before promotion, since aggregate metrics alone hide class-level weaknesses that would affect production accuracy.
- ✗
Ensure the model achieves at least 95% accuracy on a cross-validation split.
Why it's wrong here
Accuracy alone can be misleading, especially with imbalanced classes; precision and recall are more important.
- ✗
Use the training set to compute accuracy and ensure it is above 90%.
Why it's wrong here
Performance on the training set is inflated and does not reflect generalization.
- ✗
Compare the model's performance to a baseline model that always predicts the most common class.
Why it's wrong here
While comparing to a baseline is useful, it is not a required action before promotion; the question asks for actions to ensure performance requirements.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AI-102 question from scratch — 761 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on AI-102
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. You are developing a custom text classification model using Azure AI Language. You have labeled 2000 documents across 10 categories. You need to evaluate the model's performance before deploying to production. Which THREE metrics should you examine?
hard- ✓ A.Recall
- B.Word Error Rate
- C.BLEU Score
- ✓ D.F1 Score
- ✓ E.Precision
Why A: Recall is correct because it measures the proportion of actual positive instances correctly identified by the model, which is critical in custom text classification to ensure that relevant documents are not missed. In Azure AI Language, recall helps assess how well the model captures all instances of each category, especially when class distribution is imbalanced across the 10 categories.
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.