Courseiva

AI-102 · topic practice

Implement knowledge mining and information extraction solutions practice questions

This domain covers Azure AI Search pipelines for knowledge mining: ingesting documents from Blob Storage, using AI enrichment with built-in and custom skills, and extracting fields from scanned or unstructured content. Questions test index schema design, indexers, skillsets, and scheduled enrichment for searchable, filterable solutions.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Implement knowledge mining and information extraction solutions

What the exam tests

What to know about Implement knowledge mining and information extraction solutions

Candidates must design and implement Azure AI Search enrichment pipelines: create data sources, indexers, skillsets, and indexes that extract and enrich content. The critical skill is correctly mapping enriched outputs to searchable and filterable index fields for the required query behavior.

Configuring Azure AI Search indexers to ingest PDFs and scanned documents from Azure Blob Storage.

Building skillsets with built-in skills like OCR, key phrase extraction, and entity recognition.

Using the Custom Web API skill to call Azure Functions or external enrichment logic.

Creating indexes with searchable, filterable, and facetable fields for query scenarios.

Watch out for

Common Implement knowledge mining and information extraction solutions exam traps

  • ▸Forgetting that OCR skill is required for scanned invoices; assuming indexer alone extracts text from images.
  • ▸Confusing indexer schedules with skillset execution; enrichment runs during indexer runs, not independently.
  • ▸Overlooking that custom skills require a skill definition with a valid URI and that outputs must map to index fields.

Practice set

Implement knowledge mining and information extraction solutions questions

20 questions · select your answer, then reveal the explanation

You are building a knowledge mining solution to extract insights from a large set of PDF contracts. The solution must identify parties, dates, and monetary amounts. Which Azure AI service should you use as the primary extraction engine?

You are designing a knowledge mining solution that must extract entities from scanned handwritten forms. The forms contain signatures and checkboxes. Which combination of Azure AI services should you recommend?

Your organization has a knowledge base of technical manuals in PDF format. You need to enable users to ask natural language questions and get answers from the manuals. Which solution should you build?

Your knowledge mining pipeline uses Azure AI Search with a custom skillset that calls an Azure Function. The function sometimes times out for large documents. What is the best way to handle this?

Which TWO actions should you take to ensure that an Azure AI Search indexer can access data from an Azure Storage account that contains sensitive data?

Which THREE components are required to build a knowledge mining solution using Azure AI Search that extracts and enriches content from PDF files?

You are reviewing a skillset definition for an Azure AI Search indexer. The indexer is configured to index 1000 PDF documents. After running the indexer, you notice that only 500 documents have sentiment scores. What is the most likely cause?

Exhibit

Refer to the exhibit.

{
  "name": "my-skillset",
  "description": "Custom skillset",
  "skills": [
    {
      "@odata.type": "#Microsoft.Skills.Text.SplitSkill",
      "context": "/document",
      "textSplitMode": "pages",
      "maximumPageLength": 5000,
      "defaultLanguageCode": "en",
      "inputs": [
        { "name": "text", "source": "/document/content" }
      ],
      "outputs": [
        { "name": "textItems", "targetName": "pages" }
      ]
    },
    {
      "@odata.type": "#Microsoft.Skills.Text.V3.SentimentSkill",
      "context": "/document/pages/*",
      "defaultLanguageCode": "en",
      "inputs": [
        { "name": "text", "source": "/document/pages/*" }
      ],
      "outputs": [
        { "name": "sentiment", "targetName": "sentiment" }
      ]
    }
  ]
}

You review the configuration for an Azure AI Search indexer. The indexer runs successfully but no documents are indexed. What is the most likely cause?

Exhibit

Refer to the exhibit.

{
  "dataSource": {
    "name": "blob-datasource",
    "type": "azureblob",
    "credentials": {
      "connectionString": "DefaultEndpointsProtocol=https;AccountName=myaccount;AccountKey=...;EndpointSuffix=core.windows.net"
    },
    "container": {
      "name": "documents"
    }
  },
  "index": {
    "name": "docs-index",
    "fields": [
      {"name":"id","type":"Edm.String","key":true,"searchable":false},
      {"name":"content","type":"Edm.String","searchable":true},
      {"name":"metadata_storage_name","type":"Edm.String","searchable":true}
    ]
  },
  "indexer": {
    "name": "docs-indexer",
    "dataSourceName": "blob-datasource",
    "targetIndexName": "docs-index",
    "parameters": {
      "batchSize": 10,
      "maxFailedItems": -1
    }
  }
}

You are using Azure AI Search to index a set of PDF documents. The index includes a 'content' field with the extracted text. Users report that when they search for 'budget forecast', documents containing only 'budget' or 'forecast' are ranked lower than expected. Which configuration change would improve the ranking for multi-word queries?

You are building an Azure AI Search solution that indexes data from multiple sources, including SQL Database and Azure Blob Storage. The index must be updated within 15 minutes of any source change. Which approach should you use to achieve near-real-time indexing?

You are building a solution to extract customer feedback from PDF documents stored in Azure Blob Storage. The solution must extract key phrases and sentiment scores, but you cannot use any pre-built models from Azure AI Language. What should you use?

Your company has a large collection of legal contracts in PDF format stored in Azure Blob Storage. You need to extract key clauses, parties, and effective dates using a custom model in Azure AI Document Intelligence. The model must be retrained monthly as new contract templates are added. What is the recommended approach to handle model versioning and retraining?

You need to extract personally identifiable information (PII) from a set of text documents before indexing them in Azure AI Search. The PII must be redacted. Which Azure AI service and configuration should you use?

Which THREE components are essential when building a custom skill for Azure AI Search?

Which TWO Azure AI services can be used to extract text from images as part of an Azure AI Search enrichment pipeline?

You are reviewing the skillset definition for an Azure AI Search indexer. The SplitSkill splits the document content into pages of 5000 characters. The SentimentSkill is set to run on each page. However, the sentiment analysis is not producing correct results. What is the most likely cause?

Exhibit

Refer to the exhibit.

{
  "skills": [
    {
      "@odata.type": "#Microsoft.Skills.Text.SplitSkill",
      "name": "#1",
      "context": "/document",
      "inputs": [
        {"name": "text", "source": "/document/content"},
        {"name": "textSplitMode", "source": "pages"},
        {"name": "maximumPageLength", "source": 5000}
      ],
      "outputs": [
        {"name": "textItems", "targetName": "pages"}
      ]
    },
    {
      "@odata.type": "#Microsoft.Skills.Text.V3.SentimentSkill",
      "name": "#2",
      "context": "/document/pages/*",
      "inputs": [
        {"name": "text", "source": "/document/pages/*"}
      ],
      "outputs": [
        {"name": "sentiment", "targetName": "sentimentLabel"},
        {"name": "confidenceScores", "targetName": "confidenceScores"}
      ]
    }
  ]
}

You have an Azure AI Search indexer that is configured to index PDF files from Azure Blob Storage. The indexer is not extracting any text from the PDFs, and no errors are reported. You review the indexer definition as shown. What is the most likely cause?

Exhibit

Refer to the exhibit.

{
  "dataSourceName": "myblob",
  "skillsetName": "mypdfskillset",
  "targetIndexName": "myindex",
  "parameters": {
    "configuration": {
      "dataToExtract": "contentAndMetadata",
      "parsingMode": "json"
    }
  },
  "fieldMappings": [
    {"sourceFieldName": "metadata_storage_path", "targetFieldName": "path"},
    {"sourceFieldName": "content", "targetFieldName": "content"}
  ]
}

You executed the Azure CLI command shown to create an indexer. However, the indexer fails to run. The error indicates that the data source connection string is invalid. You have verified that the connection string is correct. What is the most likely issue?

Network Topology
az search indexer createname myindexerresource-group myrgservice-name mysearchdata-source-name myblobskillset-name mypdfskillsettarget-index-name myindexparameters '{"configuration": {"dataToExtract": "contentAndMetadata"query '{"status": "success"}'Refer to the exhibit.

You are designing a knowledge mining solution for a large legal firm. The solution must extract key clauses, parties, and dates from thousands of PDF contracts. You need to minimize manual labeling effort while achieving high extraction accuracy. Which Azure AI service should you use?

Your company uses Azure AI Search for an internal knowledge base. Users complain that searches for 'annual report 2023' return irrelevant results. You analyze the search index and find that the content field contains large blocks of text from PDFs. You need to improve relevance without re-indexing all documents. Which approach should you take?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Implement knowledge mining and information extraction solutions sessions

Start a Implement knowledge mining and information extraction solutions only practice session

Every question in these sessions is drawn from the Implement knowledge mining and information extraction solutions domain — nothing else.

Related practice questions

Related AI-102 topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the AI-102 exam test about Implement knowledge mining and information extraction solutions?
Candidates must design and implement Azure AI Search enrichment pipelines: create data sources, indexers, skillsets, and indexes that extract and enrich content. The critical skill is correctly mapping enriched outputs to searchable and filterable index fields for the required query behavior.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Implement knowledge mining and information extraction solutions questions in a focused session?
Yes — the session launcher on this page draws every question from the Implement knowledge mining and information extraction solutions domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other AI-102 topics?
Use the topic links above to move to related areas, or go back to the AI-102 question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the AI-102 exam covers. They are not copied from any real exam or dump site.