Courseiva

AI-102 Practice Question: Implement knowledge mining and information extraction solutions

Your organization is building a knowledge base from technical manuals stored in multiple formats (PDF, Word, HTML). You need to extract text and images from these documents and create a searchable index. The solution must handle tables and preserve their structure. Which approach should you use?

⚠ Common exam trap

AI-102 often tests the distinction between OCR (text extraction from images) and document layout analysis (structure extraction), causing candidates to pick the Vision OCR skill when tables and structure are required.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Azure AI Document Intelligence layout model as a custom skill

The Azure AI Document Intelligence layout model is purpose-built to extract text, tables, and structure from documents in PDF, Word, and HTML formats, and it preserves table structure and reading order. When integrated as a custom skill in an Azure AI Search skillset, it enriches the pipeline with structured content that can be indexed and searched. This directly satisfies the requirement to handle tables while preserving their structure.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Upload documents directly to Azure AI Search

    Why it's wrong here

    Uploading directly to Azure AI Search indexes only plain text and basic metadata; it cannot extract embedded images or preserve table structure from PDF, Word and HTML. It is tempting for its simplicity, and would be correct for already-extracted plain text needing full-text search.

  • ✗

    Use Azure AI Language custom entity extraction

    Why it's wrong here

    Custom entity extraction labels named entities in existing text; it neither parses PDF, Word or HTML content nor preserves table structure. It is tempting because it extracts structured fields from unstructured prose, which suits scenarios like pulling invoice numbers or product names from free text, not document layout indexing.

  • ✓

    Use Azure AI Document Intelligence layout model as a custom skill

    Why this is correct

    The layout model outputs structured Markdown and JSON that preserves table rows, columns and headings, unlike plain OCR which flattens them. Wrapping it as a custom skill lets the indexer enrich each PDF, Word or HTML document, satisfying the requirement to retain table structure in the searchable index.

  • ✗

    Use Azure AI Vision OCR skill in the skillset

    Why it's wrong here

    The AI Vision OCR skill extracts printed text from images only; it does not parse document layouts, embedded images or table structure across PDF, Word and HTML. It is tempting because it performs optical character recognition, and would be correct for reading text within scanned image files.

About these practice questions

This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.