AI-102 Practice Question: Implement knowledge mining and information extraction solutions
You are implementing an Azure AI Search enrichment pipeline that extracts text from PDF documents stored in Azure Blob Storage. The PDFs are scanned images with no embedded text layer. You need to ensure the extracted text is available for downstream skills. Which skill should you add to the skillset?
⚠ Common exam trap
The trap here is assuming that the built-in document extraction skill automatically handles scanned images, when in fact it only extracts text from documents with an embedded text layer.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Microsoft.Skills.Vision.OcrSkill
The OCR skill is specifically designed to extract text from images, including scanned PDFs, within an Azure AI Search enrichment pipeline. It leverages the Computer Vision Read API to recognize text and output it for further processing. Other skills like Split, Merge, or custom Web API do not provide built-in OCR capabilities. Therefore, adding the OCR skill is the correct approach to make the scanned PDF content searchable.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Microsoft.Skills.Text.SplitSkill
Why it's wrong here
The Split skill is used to divide large text into smaller chunks, typically for language detection or translation. It does not perform any recognition of text from images. Since the PDFs are scanned images without a text layer, the Split skill would receive no text to split. It is not suitable for extracting text from image-based documents.
- ✗
Microsoft.Skills.Text.MergeSkill
Why it's wrong here
The Merge skill combines text and optionally embedded images from multiple fields into a single consolidated field. It does not extract text from images. While it could be used after OCR to merge extracted text with other content, it cannot itself perform OCR. Therefore, it does not solve the problem of scanned PDFs lacking a text layer.
- ✗
Microsoft.Skills.Custom.WebApiSkill
Why it's wrong here
The Web API skill allows calling a custom HTTP endpoint to perform custom enrichment. While you could potentially build a custom service to perform OCR, this is not the built-in solution and adds unnecessary complexity. The scenario does not require custom logic; a built-in OCR skill is available and more appropriate. Using a custom skill would require additional infrastructure and maintenance.
- ✓
Microsoft.Skills.Vision.OcrSkill
Why this is correct
The OCR skill uses the Computer Vision Read API to extract text from images, including scanned PDF pages. It is designed for exactly this scenario where documents lack a text layer. The skill outputs text and layout information that can be mapped to index fields. This is the correct choice because it directly addresses the need to perform optical character recognition on image-based PDFs within the enrichment pipeline.
Quick reference
Azure Blob Storage Tier Comparison
| Tier | Storage Cost | Retrieval Cost | Latency | Use Case |
|---|---|---|---|---|
| Hot | Highest | Lowest | Immediate | Active data, frequent reads |
| Cool | Lower | Higher | Immediate | Data accessed < once / month |
| Cold | Lower still | Higher | Immediate | Data accessed < once / quarter |
| Archive | Lowest | Highest + rehydration delay | Hours | Long-term compliance retention |
Go deeper
Related to this question
About these practice questions
This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.