AI-102 Practice Question: Implement knowledge mining and information extraction solutions
Your company deploys an Azure AI Document Intelligence solution to extract data from invoices. During testing, you notice that some fields are not being extracted correctly, especially for invoices from a specific vendor with a non-standard layout. You need to improve extraction accuracy for this vendor's invoices. What should you do?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Train a custom model using labeled samples of the vendor's invoices.
Training a custom model using labeled samples of the vendor's invoices allows Document Intelligence to learn the non-standard layout, improving extraction accuracy. Option A is incorrect because while OCR and regex can extract text, they are not effective for structured data extraction from variable layouts. Option B is incorrect because converting invoices to a standard format is time-consuming and may lose important data; it's better to train a model on the actual invoices. Option D is incorrect because the prebuilt invoice model is designed for standard invoice layouts; adjusting confidence thresholds won't fix extraction accuracy for non-standard formats.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable OCR on the documents and use regular expressions to extract fields.
Why it's wrong here
OCR plus regular expressions bypasses the prebuilt model's field understanding and breaks whenever the vendor's layout shifts. Regex extraction suits stable, rigidly structured text, whereas training a custom model on labelled vendor invoices teaches the service that non-standard layout directly.
- ✗
Convert the invoices to a standard format before processing.
Why it's wrong here
Reformatting invoices alters the input rather than teaching the model the vendor's layout, and it demands preprocessing effort outside the service. Standardising documents helps when formats are inconsistent, but custom model training on the vendor's own labelled invoices addresses the extraction failure directly.
- ✓
Train a custom model using labeled samples of the vendor's invoices.
Why this is correct
A custom model trained on labelled samples of that vendor's invoices learns its non-standard layout and field positions, which the prebuilt invoice model cannot capture. This directly targets the extraction accuracy problem for that specific vendor.
- ✗
Use the prebuilt invoice model with confidence threshold adjustment.
Why it's wrong here
Confidence thresholds only filter which extracted values are returned; they cannot teach the model a vendor's non-standard layout, so accuracy stays poor. Threshold tuning suits balancing precision against recall on well-supported layouts, not correcting fields the prebuilt model never learned to locate.
Go deeper
Related to this question
About these practice questions
This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.