AI-102 Practice Question: Implement knowledge mining and information extraction solutions
You are designing an Azure AI Search solution that indexes documents from an Azure SQL Database. The documents include a field named 'content' that contains HTML markup. You need to strip the HTML tags and extract only the plain text before applying further enrichment. Which built-in skill should you use?
⚠ Common exam trap
The trap here is assuming that Text Split or Text Merger can also clean HTML, when they only restructure text without removing markup.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
HTML Strip skill
The HTML Strip skill is specifically designed to remove HTML markup and return plain text. It is the correct choice for cleaning HTML content before applying other enrichment skills. The Text Merger and Text Split skills manipulate text structure but do not remove HTML. Language Detection analyzes language but does not alter the text content.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Text Merger skill
Why it's wrong here
The Text Merger skill combines text from multiple fields into a single string. It does not remove HTML tags. While it could concatenate content from different fields, it would not strip HTML markup. Using it here would not achieve the goal of extracting plain text from HTML, and the HTML tags would remain in the output, potentially interfering with subsequent skills.
- ✗
Text Split skill
Why it's wrong here
The Text Split skill breaks large text into smaller chunks based on size or sentence boundaries. It does not remove HTML tags. It is useful for chunking text before sending to skills with input size limits, but it does not clean HTML. Applying it to HTML content would still leave the tags in the chunks, so it does not meet the requirement of stripping HTML.
- ✓
HTML Strip skill
Why this is correct
The HTML Strip skill is a built-in cognitive skill that removes HTML tags from a string and returns plain text. It is designed exactly for this scenario: cleaning HTML content before further processing. By using this skill, you ensure that subsequent enrichment skills receive clean text, improving the accuracy of language detection, entity recognition, and other NLP tasks.
- ✗
Language Detection skill
Why it's wrong here
Language Detection identifies the language of the text but does not modify the text or remove HTML tags. It would detect the language of the HTML content, possibly with reduced accuracy due to the tags, but it would not produce plain text. This skill is not suitable for stripping HTML and would not fulfill the requirement.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AI-102 question from scratch — 761 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.