AI-102 Practice Question: Implement knowledge mining and information extraction solutions
You are building an Azure AI Search enrichment pipeline that processes PDF documents from Azure Blob Storage. The documents contain both text and images. You need to extract text from the images and also detect the language of the extracted text to route documents to language-specific processing. Which two built-in skills should you include in the skillset? (Choose two.)
⚠ Common exam trap
The trap here is assuming that language detection can work directly on images or that OCR also detects language, when in fact they are separate steps that must be chained.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
OCR skill
The OCR skill is required to extract text from images within the PDFs. The Language Detection skill is required to identify the language of the extracted text for routing. Together, they enable the pipeline to process image-based text and then determine its language. The other skills do not provide these capabilities: key phrase extraction, entity recognition, and translation serve different purposes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Text Translation skill
Why it's wrong here
Text Translation translates text from one language to another. It does not extract text from images or detect language. The requirement is to detect the language, not translate it. Using this skill would be premature because you first need to detect the language before deciding whether translation is necessary. It does not meet either of the two specified needs.
- ✓
OCR skill
Why this is correct
The OCR skill is a built-in cognitive skill that extracts text from images. It is essential for this scenario because the PDFs contain images with embedded text. The OCR skill can process image files or images extracted from PDFs (when using the Document Extraction skill). It outputs text that can then be used by other skills, such as language detection. Without OCR, the text in images would not be available for indexing or further enrichment.
- ✓
Language Detection skill
Why this is correct
The Language Detection skill identifies the language of the input text and returns a language code and name. It is needed here to determine the language of the text extracted by OCR, so that you can route documents to language-specific processing. This skill takes a single string input and outputs the detected language, which can be used to conditionally apply other skills or to populate a language field in the index.
- ✗
Entity Recognition skill
Why it's wrong here
Entity Recognition extracts named entities such as people and organizations from text. It does not extract text from images or detect language. While it could be used later to enrich the extracted text, it does not fulfill the stated requirements. Adding it would not help with OCR or language detection and would increase the complexity and cost of the skillset without providing the needed functionality.
- ✗
Key Phrase Extraction skill
Why it's wrong here
Key Phrase Extraction identifies main concepts in text but does not extract text from images or detect language. It would be useful for summarizing content, but the requirements are specifically to extract text from images and detect language. Including this skill would not address either requirement and would add unnecessary processing overhead to the enrichment pipeline.
Quick reference
Azure Blob Storage Tier Comparison
| Tier | Storage Cost | Retrieval Cost | Latency | Use Case |
|---|---|---|---|---|
| Hot | Highest | Lowest | Immediate | Active data, frequent reads |
| Cool | Lower | Higher | Immediate | Data accessed < once / month |
| Cold | Lower still | Higher | Immediate | Data accessed < once / quarter |
| Archive | Lowest | Highest + rehydration delay | Hours | Long-term compliance retention |
Go deeper
Related to this question
About these practice questions
This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.