Courseiva

AI-102 · topic practice

Implement knowledge mining and document intelligence solutions practice questions

This domain covers Azure AI Search pipelines (data sources, indexers, skillsets, indexes) and Azure AI Document Intelligence custom and prebuilt models. Questions test choosing the right cognitive skill, wiring enrichment to an index, and diagnosing extraction or confidence issues in knowledge mining solutions.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Implement knowledge mining and document intelligence solutions

What the exam tests

What to know about Implement knowledge mining and document intelligence solutions

Be able to design an Azure AI Search enrichment pipeline and pick the correct Document Intelligence model type. The critical skill is mapping skillset outputs into index fields and knowing when to retrain a custom model with more labeled documents.

Selecting built-in skillset skills such as Key Phrase Extraction, Language Detection, and OCR

Configuring indexers, data sources, and output field mappings from enriched documents

Choosing Document Intelligence prebuilt models (invoice, receipt, ID) versus custom models

Training and improving custom Document Intelligence models with labeled samples

Watch out for

Common Implement knowledge mining and document intelligence solutions exam traps

  • ▸Adding Key Phrase Extraction without Language Detection, so multilingual documents are not normalized before key phrase extraction runs.
  • ▸Forgetting output field mappings or the knowledge store, so enriched skill outputs never reach the search index.
  • ▸Assuming a custom Document Intelligence model generalizes; low-confidence or failing extractions usually mean more labeled samples are needed.

Practice set

Implement knowledge mining and document intelligence solutions questions

20 questions · select your answer, then reveal the explanation

A company uses Azure Document Intelligence to extract data from invoices. They deploy the model to a container for on-premises processing. After deployment, they notice that the container consumes more memory than expected. What should they do to optimize memory usage?

A company uses Azure Document Intelligence to extract data from tax forms. They need to improve accuracy for a specific field. Which TWO actions should they take?

A company uses this skillset in an Azure AI Search enrichment pipeline. They notice that the enrichment pipeline fails when processing a document larger than 5000 characters. What is the most likely cause?

Exhibit

Refer to the exhibit.

```json
{
  "skills": [
    {
      "@odata.type": "#Microsoft.Skills.Text.SplitSkill",
      "name": "#1",
      "context": "/document",
      "defaultLanguageCode": "en",
      "textSplitMode": "pages",
      "maximumPageLength": 5000,
      "inputs": [
        {
          "name": "text",
          "source": "/document/content"
        }
      ],
      "outputs": [
        {
          "name": "textItems",
          "targetName": "pages"
        }
      ]
    }
  ]
}```

A company uses Azure Cognitive Search to index customer support emails. They need to implement a custom skill that extracts the sentiment of the email body and also identifies the primary product mentioned. The custom skill is a Python function deployed as an Azure Function. They want to ensure the skill can process multiple documents concurrently and handle errors gracefully. Which THREE configurations should they apply?

Refer to the exhibit. You have a skillset with two skills. You run the indexer and find that the output field 'organizations' is empty for documents that clearly contain organization names. The 'keyPhrases' output is populated correctly. What is the most likely cause of the issue?

Exhibit

Refer to the exhibit.
```json
{
  "skills": [
    {
      "@odata.type": "#Microsoft.Skills.Text.EntityRecognitionSkill",
      "name": "#1",
      "description": "Extract organizations",
      "context": "/document",
      "categories": ["Organization"],
      "inputs": [
        {
          "name": "text",
          "source": "/document/content"
        }
      ],
      "outputs": [
        {
          "name": "organizations",
          "targetName": "organizations"
        }
      ]
    },
    {
      "@odata.type": "#Microsoft.Skills.Text.KeyPhraseExtractionSkill",
      "name": "#2",
      "description": "Extract key phrases",
      "context": "/document",
      "inputs": [
        {
          "name": "text",
          "source": "/document/content"
        }
      ],
      "outputs": [
        {
          "name": "keyPhrases",
          "targetName": "keyPhrases"
        }
      ]
    }
  ]
}
```

You are a developer for a legal firm. The firm uses Azure Cognitive Search to index legal documents. They have a custom skill that performs OCR on scanned PDFs using Azure Form Recognizer. The skill is implemented as an Azure Function. Recently, the indexer has been failing with the error: "The request was canceled due to the configured HttpClient.Timeout of 100 seconds elapsing." The documents are large (up to 200 pages each). The skill calls the Form Recognizer API asynchronously. You need to resolve the timeout issue without losing the ability to process large documents. Current configuration: batchSize = 1, maxPageSize = 100, timeout = 100 seconds. You cannot change the execution time of the Form Recognizer API. What should you do?

A logistics company uses Azure AI Search with an indexer and skillset to enrich shipping manifests. The skillset includes a custom skill hosted on an Azure Function that enriches each document. During a full reindex, the indexer reports that the custom skill returns HTTP 429 responses for many documents, and enriched fields are missing. You need to make the pipeline resilient without changing the Function's logic. What should you do?

You are building a knowledge mining solution with Azure AI Search. Documents are stored in Azure Blob Storage and include scanned PDFs with embedded images and text. You need to extract both the text and the textual content of images during indexing. You create a data source, an index, and an indexer. You must configure the indexer to use the built-in OCR skill. What should you do first?

You are building a knowledge mining solution in Azure AI Search. Documents are ingested from Azure Blob Storage and must be enriched with a custom skill that calls an Azure Function. The Function requires an API key that must not be stored in the skillset definition or source code. The indexer runs on a schedule. You need to configure the custom skill so that the API key is supplied securely at runtime. What should you do?

A financial services firm uses an Azure AI Search indexer with a custom WebApiSkill that calls an internal scoring API. The API occasionally returns HTTP 500 responses due to transient load. The indexer currently fails the entire run whenever the custom skill returns an error. You need to make the pipeline more resilient to these transient failures. What should you do?

A research organization uses Azure AI Search with a skillset to enrich scientific papers. They want to extract named entities such as people, organizations, and locations from the paper text and store them in a collection field in the index. The entities must be searchable individually. Which configuration should they use in the index definition?

You are implementing a knowledge mining solution using Azure AI Search. You have an indexer that processes documents from Azure Blob Storage. You need to enrich the documents with entities extracted by Azure AI Language. You create a skillset with the Entity Recognition skill. You want to ensure that the extracted entities are stored in a complex collection field in the index. What should you do?

You are building a knowledge mining solution with Azure AI Search. The data source is an Azure Blob Storage container holding scanned PDF invoices. You configure a skillset that includes the built-in OCR skill followed by the built-in Entity Recognition skill. After running the indexer, you notice that entities are extracted from the OCR text but the original document content is not searchable. You need to ensure that both the raw text and the extracted entities are searchable. What should you do?

A financial services firm uses Azure AI Search to index earnings call transcripts stored in Azure Cosmos DB. They need to enrich each transcript with a custom skill that calls an external NLP API. The API returns a large JSON payload that must be split into separate fields for sentiment, entities, and topics. The indexer runs daily. You need to configure the skillset to map the API response into the index. What should you do?

A financial services company uses Azure AI Document Intelligence with a custom neural model to extract fields from loan applications. The model was trained on 200 labeled documents and performs well on the training set but poorly on new documents from a different branch office. The new documents use slightly different terminology and layout. You need to improve the model's accuracy on the new documents with minimal labeling effort. What should you do?

A company uses Azure AI Document Intelligence with a custom extraction model to process purchase orders. They need to extract line items, including product codes, quantities, and prices, from tables in the documents. The model currently returns line items but the product codes are frequently missing or misaligned. You need to improve the extraction of line items with minimal effort. What should you do?

A company is building a knowledge mining solution using Azure AI Search. They need to extract key phrases from a large set of documents in multiple languages. Which skill should they add to the skillset?

A healthcare organization uses Azure Document Intelligence to process patient intake forms. They notice that the confidence scores for field extraction are low. What is the most likely cause?

A company builds a knowledge mining solution using Azure AI Search with a custom skillset that includes an OCR skill. They want to ensure that images embedded in PDFs are processed. What should they configure?

A law firm uses Azure Document Intelligence to extract clauses from legal contracts. They have a custom model trained on 15 labeled contracts. The model extracts clauses with high confidence on similar documents but fails to extract correct clauses from a new batch of contracts that have a different font and layout. The firm needs to improve extraction accuracy without retraining the model from scratch. The solution must minimize manual effort and cost. What should they do?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Implement knowledge mining and document intelligence solutions sessions

Start a Implement knowledge mining and document intelligence solutions only practice session

Every question in these sessions is drawn from the Implement knowledge mining and document intelligence solutions domain — nothing else.

Related practice questions

Related AI-102 topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the AI-102 exam test about Implement knowledge mining and document intelligence solutions?
Be able to design an Azure AI Search enrichment pipeline and pick the correct Document Intelligence model type. The critical skill is mapping skillset outputs into index fields and knowing when to retrain a custom model with more labeled documents.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Implement knowledge mining and document intelligence solutions questions in a focused session?
Yes — the session launcher on this page draws every question from the Implement knowledge mining and document intelligence solutions domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other AI-102 topics?
Use the topic links above to move to related areas, or go back to the AI-102 question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the AI-102 exam covers. They are not copied from any real exam or dump site.