Courseiva

AI-102 Practice Question: Implement knowledge mining and information extraction solutions

You are implementing an Azure AI Search enrichment pipeline that extracts text from PDF documents stored in Azure Blob Storage. The PDFs are scanned images with no embedded text layer. You need to ensure the extracted text is available for downstream skills. Which skill should you add to the skillset?

⚠ Common exam trap

The trap here is assuming that the built-in document extraction skill automatically handles scanned images, when in fact it only extracts text from documents with an embedded text layer.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Microsoft.Skills.Vision.OcrSkill

The OCR skill is specifically designed to extract text from images, including scanned PDFs, within an Azure AI Search enrichment pipeline. It leverages the Computer Vision Read API to recognize text and output it for further processing. Other skills like Split, Merge, or custom Web API do not provide built-in OCR capabilities. Therefore, adding the OCR skill is the correct approach to make the scanned PDF content searchable.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Microsoft.Skills.Text.SplitSkill

    Why it's wrong here

    The Split skill is used to divide large text into smaller chunks, typically for language detection or translation. It does not perform any recognition of text from images. Since the PDFs are scanned images without a text layer, the Split skill would receive no text to split. It is not suitable for extracting text from image-based documents.

  • ✗

    Microsoft.Skills.Text.MergeSkill

    Why it's wrong here

    The Merge skill combines text and optionally embedded images from multiple fields into a single consolidated field. It does not extract text from images. While it could be used after OCR to merge extracted text with other content, it cannot itself perform OCR. Therefore, it does not solve the problem of scanned PDFs lacking a text layer.

  • ✗

    Microsoft.Skills.Custom.WebApiSkill

    Why it's wrong here

    The Web API skill allows calling a custom HTTP endpoint to perform custom enrichment. While you could potentially build a custom service to perform OCR, this is not the built-in solution and adds unnecessary complexity. The scenario does not require custom logic; a built-in OCR skill is available and more appropriate. Using a custom skill would require additional infrastructure and maintenance.

  • ✓

    Microsoft.Skills.Vision.OcrSkill

    Why this is correct

    The OCR skill uses the Computer Vision Read API to extract text from images, including scanned PDF pages. It is designed for exactly this scenario where documents lack a text layer. The skill outputs text and layout information that can be mapped to index fields. This is the correct choice because it directly addresses the need to perform optical character recognition on image-based PDFs within the enrichment pipeline.

Quick reference

Azure Blob Storage Tier Comparison

TierStorage CostRetrieval CostLatencyUse Case
HotHighestLowestImmediateActive data, frequent reads
CoolLowerHigherImmediateData accessed < once / month
ColdLower stillHigherImmediateData accessed < once / quarter
ArchiveLowestHighest + rehydration delayHoursLong-term compliance retention

About these practice questions

This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.