AI-102 Practice Question: Implement knowledge mining and information extraction solutions
You are a solution architect at a news agency. The agency publishes thousands of articles daily. You need to build a knowledge mining solution that enables journalists to search for articles by topic, sentiment, key people, and locations mentioned. The articles are stored as HTML files in Azure Blob Storage. The solution must also provide a summary for each article. You plan to use Azure AI Search with cognitive skills and Azure OpenAI. Which combination of skills and features should you include to meet all requirements with the best performance and accuracy?
⚠ Common exam trap
AI-102 often tests whether candidates can map each stated requirement to a specific skill — the trap is picking an option with a plausible-sounding but irrelevant skill (Text Translation, Text Analytics for Health, Document Intelligence) while missing a required capability like summarization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Skillset with Entity Recognition skill, Sentiment skill, and Key Phrase Extraction skill. Use Azure OpenAI service to generate summaries via a custom skill that calls the GPT model. Enable semantic search.
The requirements are topic search, sentiment, key people, locations, and article summaries. Entity Recognition extracts people and locations, Sentiment provides sentiment, Key Phrase Extraction supports topic search, and a custom skill calling Azure OpenAI GPT generates summaries — all wired into an Azure AI Search skillset with semantic search for relevance. This combination directly maps to every requirement without extraneous skills.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Azure AI Document Intelligence to extract content from HTML, then use Azure AI Language to extract entities and sentiment. Index in Azure AI Search with semantic search.
Why it's wrong here
Document Intelligence targets scanned or PDF documents, not HTML, so it adds cost and latency without improving extraction; HTML parsing plus Azure AI Language and semantic search already meets the requirements. It is tempting because Document Intelligence is the standard OCR choice for image-based or PDF sources.
- ✗
Skillset with Entity Recognition skill, Sentiment skill, Key Phrase Extraction skill, and Text Translation skill. Enable semantic search.
Why it's wrong here
Text Translation converts content between languages, yet the articles are English HTML and no multilingual requirement exists, so it adds cost without satisfying topic, person or location extraction. It would be correct for a multilingual corpus needing cross-language search.
- ✓
Skillset with Entity Recognition skill, Sentiment skill, and Key Phrase Extraction skill. Use Azure OpenAI service to generate summaries via a custom skill that calls the GPT model. Enable semantic search.
Why this is correct
Entity Recognition extracts people and locations, Sentiment scores tone, and Key Phrase Extraction surfaces topics; a custom skill invoking Azure OpenAI generates article summaries. Semantic search then reranks results for relevance, meeting every stated requirement across HTML blob content.
- ✗
Skillset with Entity Recognition skill, Sentiment skill, and Text Analytics for Health skill to extract medical terms. Use Azure OpenAI for summarization as a custom skill.
Why it's wrong here
Text Analytics for Health extracts medical entities, which the stem never requests; topic, people and location extraction need Entity Recognition and Key Phrase Extraction instead. It is tempting because health skills genuinely suit clinical or biomedical corpora, where medical term extraction would be the correct inclusion.
Quick reference
Azure Blob Storage Tier Comparison
| Tier | Storage Cost | Retrieval Cost | Latency | Use Case |
|---|---|---|---|---|
| Hot | Highest | Lowest | Immediate | Active data, frequent reads |
| Cool | Lower | Higher | Immediate | Data accessed < once / month |
| Cold | Lower still | Higher | Immediate | Data accessed < once / quarter |
| Archive | Lowest | Highest + rehydration delay | Hours | Long-term compliance retention |
Go deeper
Related to this question
About these practice questions
Courseiva writes every AI-102 question from scratch — 761 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.