AI-102 · domain
Implement knowledge mining and information extraction solutions
This domain covers Azure AI Search pipelines for knowledge mining: ingesting documents from Blob Storage, using AI enrichment with built-in and custom skills, and extracting fields from scanned or unstructured content. Questions test index schema design, indexers, skillsets, and scheduled enrichment for searchable, filterable solutions.
Focused practice
Practice Implement knowledge mining and information extraction solutions questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Implement knowledge mining and information extraction solutions
Candidates must design and implement Azure AI Search enrichment pipelines: create data sources, indexers, skillsets, and indexes that extract and enrich content. The critical skill is correctly mapping enriched outputs to searchable and filterable index fields for the required query behavior.
Configuring Azure AI Search indexers to ingest PDFs and scanned documents from Azure Blob Storage.
Building skillsets with built-in skills like OCR, key phrase extraction, and entity recognition.
Using the Custom Web API skill to call Azure Functions or external enrichment logic.
Creating indexes with searchable, filterable, and facetable fields for query scenarios.
Watch out for
Common Implement knowledge mining and information extraction solutions exam traps
- ▸Forgetting that OCR skill is required for scanned invoices; assuming indexer alone extracts text from images.
- ▸Confusing indexer schedules with skillset execution; enrichment runs during indexer runs, not independently.
- ▸Overlooking that custom skills require a skill definition with a valid URI and that outputs must map to index fields.
Question index
All Implement knowledge mining and information extraction solutions questions (101)
Click any question to see the full explanation, or start a practice session above.
You are building a knowledge mining solution for legal documents stored in Azure Blob Storage. The solution must extract entities, key phrases, and relationships from the documents. Which Azure AI service should you use?
Medium2You are building a knowledge mining solution for a legal firm to extract clauses from contracts. The contracts are stored as PDFs in Azure Blob Storage. You need to design the solution to minimize cost while ensuring high accuracy for clause extraction. Which approach should you use?
Medium3You are creating an Azure AI Search index that will be populated from an enrichment pipeline. You need to ensure that the original content of each document is searchable. Which index field should you map the document content to?
Easy4Your company has a large repository of scanned invoices in PDF format. You need to extract invoice number, date, total amount, and vendor name from these PDFs. Which Azure AI service should you use?
Medium5You are deploying a knowledge mining solution using Azure AI Search and Azure AI Document Intelligence. The solution must extract text from scanned documents, identify named entities, and index the content. You need to configure the skillset. Which TWO built-in skills should you include in the skillset?
Medium6You are using Microsoft Purview to create a knowledge map of your organization's data assets. The solution must automatically scan and classify sensitive data in Azure Blob Storage. You need to configure the scanning and classification. Which THREE actions should you perform?
Hard7You are implementing a knowledge mining solution with Azure AI Search that ingests data from Azure Blob Storage. The pipeline includes a custom skill that calls an external API for specialized entity extraction. The custom skill sometimes returns HTTP 429 (Too Many Requests). How should you handle this to ensure reliable indexing?
Hard8You are designing a knowledge mining solution that ingests documents from SharePoint Online and makes them searchable using Azure AI Search. The solution must extract text from images and perform optical character recognition (OCR) on embedded images within PDFs. Which built-in skill should you include in the skillset?
Medium9You need to extract key-value pairs from a large set of invoices. The invoices have a consistent layout but vary in format (PDF, TIFF). Which Document Intelligence model should you use?
Easy10You need to extract entities such as dates, locations, and organization names from unstructured text documents. Which Azure AI service should you use?
Easy11You are designing a solution to extract structured data from a large number of handwritten forms. The forms are scanned and stored as images. Which Azure AI feature should you use?
Easy12A company uses Azure AI Search to index customer support tickets. They need to automatically extract key phrases from each ticket to improve search relevance. Which built-in skill should they add to the skillset?
Easy13You are building an Azure AI Search knowledge mining pipeline that enriches PDF documents with key phrases. The enrichment must be applied after text extraction and before the data is written to the index. You need to ensure the enriched key phrases are available for downstream skills and are mapped to an index field. Which component of the skillset defines the output of the Key Phrase Extraction skill and its mapping to the index?
Medium14Your company is building a knowledge base for customer support using Azure AI Search. You have a large dataset of customer emails stored in Azure Blob Storage. The solution must extract key phrases, detect sentiment, and identify customer intents (e.g., complaint, inquiry, feedback). You plan to use built-in AI skills for key phrase extraction and sentiment detection. For intent identification, you need a custom solution because the intents are specific to your business. You have trained a custom Language Understanding (LUIS) model and published it. How should you integrate the LUIS model into the Azure AI Search enrichment pipeline to extract intents?
Hard15You are a solution architect at a news agency. The agency publishes thousands of articles daily. You need to build a knowledge mining solution that enables journalists to search for articles by topic, sentiment, key people, and locations mentioned. The articles are stored as HTML files in Azure Blob Storage. The solution must also provide a summary for each article. You plan to use Azure AI Search with cognitive skills and Azure OpenAI. Which combination of skills and features should you include to meet all requirements with the best performance and accuracy?
Medium16You are using Azure AI Document Intelligence to process a large batch of PDF forms. The forms have varying layouts and handwriting. You need to extract text and key-value pairs. Which custom model type should you train?
Medium17Your company uses Azure AI Search to power a customer support portal. The search index includes product documentation and known issues. Recently, the portal's search performance has degraded, and users report slow response times. You need to identify the cause of the performance issue. What should you check first?
Easy18Which TWO features of Azure AI Search allow you to improve the relevance of search results for users?
Easy19Your organization has a large corpus of legal documents stored in Azure Blob Storage. You need to build a solution that allows lawyers to ask natural language questions and get answers directly from the documents, without moving data out of Azure. Which service should you use?
Medium20Which THREE Azure AI services can be used to extract text from images?
Easy21You are building an Azure AI Search enrichment pipeline that must extract text and layout information from scanned PDFs stored in Azure Blob Storage. The extracted content must include bounding boxes for each text line so that a downstream custom skill can associate key-value pairs spatially. You need to add a built-in skill to the skillset to perform this extraction. Which skill should you add?
Medium22You are creating an Azure AI Search solution that must extract named entities such as people, organizations, and locations from text documents. You want to use a built-in cognitive skill to perform this extraction during indexing. Which skill should you add to the skillset?
Easy23You are building an Azure AI Search solution that enriches documents by detecting the language of each document and then routing content to language-specific analyzers. You add a LanguageDetectionSkill to the skillset and want the detected language code to be available to downstream skills and to be stored in the index. The detected language must be mapped to a field named 'languageCode' in the index. What should you do?
Medium24Which TWO options are valid ways to index content from Azure SQL Database into Azure AI Search? (Select TWO.)
Medium25You are configuring an Azure AI Search indexer to process documents from Azure Blob Storage. The documents include PDFs and Microsoft Word files. You need to extract both text and metadata such as author and creation date. Which indexer configuration should you use?
Medium26You are building a question answering solution using Azure AI Language. You have a set of frequently asked questions (FAQs) in a Word document. You need to import the FAQs into a project. Which approach should you use?
Easy27You are creating an Azure AI Search index that will store documents enriched with key phrases and sentiment scores. You need to define the index fields to store these enriched values. The key phrases should be searchable and retrievable, and the sentiment score should be filterable and sortable. Which field definitions should you use?
Easy28You are designing a knowledge mining solution for a medical research organization. The solution must extract relationships between drugs, diseases, and genes from scientific articles. The data will be stored in a knowledge graph for querying. Which Azure AI service should you use for the extraction?
Medium29You are designing an enterprise search solution using Azure AI Search. The solution must index data from multiple sources: SQL Database, SharePoint Online, and custom REST APIs. The search index must support faceted navigation and filtering by metadata such as department and document type. You also need to ensure that updates to source data are reflected in the index within 5 minutes. Which approach should you use?
Hard30Which THREE components are required to build a custom skill for Azure AI Search enrichment?
Easy31You are building a knowledge mining solution using Azure AI Search with AI enrichment. Which TWO built-in skills can be used to extract information from images embedded in documents?
Medium32You are a developer at an e-commerce company. The company wants to build a product search feature that allows customers to search for products using natural language phrases like "red running shoes under $100". The product catalog is stored in Azure Cosmos DB and includes product descriptions, prices, and categories. The solution must use Azure AI Search and must extract entities from product descriptions to enable filtering (e.g., color, size, brand). The search must also support fuzzy matching for misspelled queries. You need to design the indexing pipeline. Which actions should you take?
Medium33Your organization has a large set of PDF invoices stored in Azure Blob Storage. You need to extract line-item details (product names, quantities, prices) and store them in Azure SQL Database for downstream reporting. The invoices have varied layouts. Which Azure AI service should you use?
Medium34You are designing a knowledge mining solution that must extract tables from scanned invoices stored in Azure Blob Storage and make the table cells searchable. The invoices are in PDF and JPEG formats. Which Azure AI service should you use to extract the tables before loading the data into Azure AI Search?
Hard35You are using Azure AI Language to perform entity recognition on customer feedback. You need to identify the sentiment expressed towards specific entities. Which feature should you use?
Easy36You are building an Azure AI Search enrichment pipeline that extracts key phrases from documents stored in Azure Blob Storage. The documents are plain text files in English. You need to add a built-in skill that identifies the main concepts in each document without writing custom code. Which skill should you use?
Medium37Your company uses Azure Cognitive Search to index millions of documents. Users report that search results include irrelevant documents. You need to improve search relevance by boosting documents that contain the search term in the title field. Which scoring profile configuration should you use?
Hard38You are building an Azure AI Search solution to index a collection of technical manuals. Users need to find documents by searching for specific terms and also have the ability to filter by document category. Which feature should you configure in the index to support filtering?
Easy39You run the Azure CLI command 'az search indexer list --search-service mysearch --query "[].{name:name, status:status, lastResult:lastResult}"' and get the above output. Your indexer shows 5 warnings. What should you do to investigate the warnings?
Medium40You are building a knowledge mining solution for legal documents using Azure AI Search. The solution must extract entities like dates, organizations, and persons from PDF files and index them. Which built-in skill should you add to the skillset to perform this extraction?
Medium41Your team is building a knowledge mining solution for research papers. You need to automatically categorize papers into topics and extract author names, publication dates, and references. The solution must use custom models because the papers are domain-specific. Which combination of Azure services should you use?
Medium42You are deploying an Azure AI Search solution that indexes medical research papers. The papers contain sensitive patient data that must be de-identified before indexing. You need to use Azure AI Services to detect and redact personal information. Which combination of skills should you include in a skillset?
Hard43Which THREE considerations are important when designing a custom skill for Azure AI Search that calls an external API for specialized data extraction?
Hard44Refer to the exhibit. You execute a search query on an Azure AI Search index and get these results. The query was 'brown fox'. Why is the first result scored higher than the second?
Medium45You need to implement a solution that searches through a collection of scanned invoices and extracts invoice numbers, dates, and total amounts. The solution must run on a schedule without manual intervention. Which Azure service should you use?
Easy46You have an Azure AI Search indexer that enriches documents with a custom skill that calls an external API. The custom skill returns a JSON object containing a list of product codes. You need to store these product codes in a collection field named 'productCodes' in the index, and you want the field to be searchable and filterable. What should you do?
Medium47You are using Azure AI Language to extract information from medical research papers. You need to identify terms like 'dosage', 'side effects', and 'contraindications' specific to the medical domain. Which capability should you use?
Medium48Refer to the exhibit. You have this Azure AI Search indexer configuration. The indexer is failing after processing 6 documents that contain errors. What should you do to ensure the indexer continues processing even if some documents fail?
Easy49You are building a knowledge mining solution for a legal firm that needs to extract key clauses from thousands of scanned contract PDFs. The solution must identify parties, effective dates, and termination conditions. Which Azure AI service should you use as the primary component?
Medium50You are designing a knowledge mining solution for a publishing company that needs to extract metadata from thousands of book manuscripts in various formats (PDF, Word, EPUB). The solution must identify authors, publication dates, and chapter titles. You are using Microsoft Foundry with Azure AI Search and Azure AI Document Intelligence. The manuscripts are stored in Azure Blob Storage. You need to ensure that the solution can handle all file formats. You have configured a skillset with a Document Intelligence skill for the PDFs and Word documents. However, the EPUB files are not being processed. What should you do to include EPUB files in the enrichment pipeline?
Medium51You are extracting text from scanned documents that are in French. Which capability of Azure AI Document Intelligence should you use?
Easy52You plan to use Azure AI Search to index a large number of text documents stored in Azure Blob Storage. The documents are in English. You want to automatically extract key phrases from the content during indexing. What should you add to the skillset?
Easy53You are building a solution to extract key information from scanned invoices. The invoices are in PDF format and contain both printed and handwritten fields. Which Azure AI service should you use?
Medium54Your organization has a large repository of technical manuals in PDF format. You need to build a chatbot that can answer questions about the content of these manuals. Which combination of Azure services should you use?
Easy55Your organization is building a knowledge base from technical manuals stored in multiple formats (PDF, Word, HTML). You need to extract text and images from these documents and create a searchable index. The solution must handle tables and preserve their structure. Which approach should you use?
Hard56Your knowledge mining solution ingests documents from multiple tenants. Each tenant's data must be isolated and searchable only by that tenant. You have a single Azure AI Search service. How should you implement multi-tenancy?
Hard57You have defined the custom WebApiSkill shown in the exhibit. The skill calls an Azure Function that can process up to 10 documents per second. However, you notice that the skill is failing with 429 errors. What is the most likely cause?
Medium58Your company has a large set of PDF documents stored in Azure Blob Storage. You need to index these documents in Azure Cognitive Search so that users can search the text content. What is the first step you should take?
Easy59You are a data scientist for Contoso Pharmaceuticals. The company has thousands of research documents in PDF format stored in Azure Blob Storage. You need to build an Azure Cognitive Search solution that enables researchers to search for documents based on chemical compound names, disease mentions, and experimental results. The solution must extract these entities using a custom AI model built in Azure AI Language. Additionally, the solution must support semantic search for natural language queries. The search index must be updated daily with new documents. You have an existing Azure AI Language custom entity extraction model that recognizes chemical compounds and diseases. The model is deployed as an endpoint. You need to configure the enrichment pipeline. What should you do?
Hard60Your knowledge mining solution uses Azure AI Search. Users complain that search results are not relevant. You have enabled semantic search but results still lack context. What should you do to improve relevance?
Medium61You have the above skillset in Azure AI Search. The indexer processes a document with 12,000 characters of content. How many entity recognition skill executions occur?
Hard62You are using Azure AI Search to index customer support tickets. You want to automatically extract the customer's sentiment and key phrases from each ticket. Which Azure AI service should you integrate as a skillset?
Easy63You are implementing a knowledge mining solution for a legal firm. The solution must ingest large volumes of legal documents (PDFs and Word files) stored in Azure Blob Storage. You need to extract text, recognize named entities (e.g., parties, judges, case numbers), and index the content for full-text search. The solution should also support redaction of sensitive information before indexing. Which combination of Azure AI services should you use?
Hard64You are a data engineer at a university. The university wants to digitize its historical student records (paper forms) to make them searchable. The records are scanned as images (JPEG) and stored in Azure Blob Storage. Each form contains handwritten fields: student name, ID number, date of birth, and degree. You need to extract these fields and index them in Azure AI Search. The solution must use Azure AI Services and minimize manual labeling effort. Which approach should you take?
Easy65You are implementing a knowledge mining solution using Azure AI Search. The data source is a large Azure Cosmos DB collection containing customer support tickets. Each ticket has fields: ticket_id, description, category, and resolution. You need to ensure that the search index can support fuzzy search and autocomplete suggestions. What should you configure in the index definition?
Hard66You are building a solution to extract key information from invoices using Azure AI Document Intelligence. The invoices contain fields such as invoice number, date, total amount, and line items. However, the model is not correctly extracting the line items. Which prebuilt model should you use?
Medium67You are building an Azure AI Search enrichment pipeline that processes PDF documents from Azure Blob Storage. The documents contain both text and images. You need to extract text from the images and also detect the language of the extracted text to route documents to language-specific processing. Which two built-in skills should you include in the skillset? (Choose two.)
Medium68You are building a knowledge mining solution that indexes technical manuals in multiple languages. The solution must enable users to search in their native language and retrieve results in the same language. Which approach should you use?
Medium69Your organization is using Azure AI Document Intelligence to process expense reports. The reports are submitted as images and need to be classified into categories (e.g., travel, office supplies) before extraction. Which feature of Document Intelligence should you use?
Medium70You are designing an Azure AI Search enrichment pipeline that extracts entities from text using the Entity Recognition skill. You need to ensure that the extracted entities are stored as a collection in the index so that users can filter and facet on them. Which index field type should you use?
Hard71You are building a knowledge mining solution using Azure AI Search. You need to ensure that sensitive information such as credit card numbers is automatically removed from the indexed content. Which built-in skill should you add to your skillset?
Easy72You are designing a knowledge mining solution that must handle sensitive customer data. The solution must ensure that personally identifiable information (PII) is not returned in search results. What should you do?
Hard73Which TWO Azure AI Search features should you enable to improve the relevance of search results for a knowledge mining solution that supports natural language queries?
Medium74Your company deploys an Azure AI Document Intelligence solution to extract data from invoices. During testing, you notice that some fields are not being extracted correctly, especially for invoices from a specific vendor with a non-standard layout. You need to improve extraction accuracy for this vendor's invoices. What should you do?
Easy75You are implementing a knowledge mining solution using Azure AI Search with a custom skillset. The custom skill is an Azure Function that enriches documents with additional metadata. You need to ensure that the custom skill receives the entire document content as input. How should you configure the skill's context and inputs?
Medium76You are reviewing an index definition created with PowerShell. The index is used for a knowledge mining solution that extracts people and organizations from documents. Users report that when they type partial names in the search bar, the suggester does not return suggestions. What is the most likely reason?
Medium77You are implementing an Azure AI Search enrichment pipeline that extracts text from PDF documents stored in Azure Blob Storage. The PDFs are scanned images with no embedded text layer. You need to ensure the extracted text is available for downstream skills. Which skill should you add to the skillset?
Medium78Which TWO Azure AI services are most appropriate for extracting text from images and recognizing handwritten text?
Easy79Your organization is using Azure AI Document Intelligence to process a mix of invoices and purchase orders. You need to ensure that documents are correctly classified before extraction. Which THREE steps should you take?
Hard80You are troubleshooting an Azure AI Search indexer that fails to index a PDF file stored in Azure Blob Storage. The error message indicates that the document is encrypted. What is the most likely cause and solution?
Medium81Your organization is implementing a knowledge mining solution for a research institute that needs to extract chemical compound names and reactions from scientific articles in PDF format. The solution must use a custom model because the scientific terminology is not covered by built-in skills. You have trained a custom model using Azure AI Language's custom entity recognition (NER) and deployed it as a REST endpoint. You are using Azure AI Search with a skillset. How should you integrate the custom NER model into the enrichment pipeline?
Medium82Which TWO configurations are required to enable Azure AI Search to index content from an Azure SQL database?
Medium83Which THREE factors should you consider when designing a knowledge mining solution that uses Azure AI Search and custom skills to extract insights from large volumes of documents?
Hard84You are designing a solution to extract customer names and addresses from scanned handwritten forms. The forms are stored as images in Azure Blob Storage. The extraction must achieve high accuracy with minimal manual review. Which combination of Azure AI services should you use?
Hard85You are using Azure AI Search to index a set of contracts. You need to extract named entities such as organizations, people, and dates from the contract text and store them as separate fields in the index. Which skill should you add to the skillset?
Easy86Your team is using Azure AI Search to index a large collection of technical manuals. Users report that searches for 'disk failure' do not return relevant results because the manuals use terms like 'hard drive crash'. Which feature should you implement to improve recall?
Hard87You are designing an Azure AI Search solution that indexes documents from an Azure SQL Database. The documents include a field named 'content' that contains HTML markup. You need to strip the HTML tags and extract only the plain text before applying further enrichment. Which built-in skill should you use?
Easy88You need to extract key-value pairs from scanned forms as part of a knowledge mining solution. Which Azure AI service should you use?
Easy89Which TWO Azure AI services can be used to extract text from images as part of a knowledge mining pipeline?
Easy90Your team has built a knowledge mining pipeline using Azure AI Search and Document Intelligence. After ingestion, you notice that some documents are not appearing in search results. What is the most likely cause?
Easy91Which TWO actions should you take to optimize the performance of an Azure AI Search solution that indexes large volumes of data?
Medium92You are building an Azure AI Search enrichment pipeline that enriches documents with key phrases detected by Azure AI Language. You need the key phrases to be returned as a structured collection that can be mapped to a Collection(Edm.String) field in the search index, and you must avoid storing the enriched text in the enrichment cache. Which output mapping configuration should you use in the indexer?
Medium93You are designing a knowledge mining solution for a manufacturing company that needs to extract information from equipment maintenance manuals. The manuals are in multiple languages (English, French, German). You need to ensure that the extracted content is searchable in English only. Which approach should you use?
Medium94An organization uses Azure AI Search to power an internal knowledge base. They notice that search results are returning irrelevant documents. The index includes a 'content' field with full text and a 'tags' field with metadata. Users often search for specific terms that appear in the 'tags' field. How should you configure the search index to improve relevance?
Hard95You have an Azure AI Search indexer that uses a custom skill hosted in an Azure Function to normalize product codes. The function occasionally returns HTTP 429 responses. You need the indexer to retry these calls automatically without failing the entire indexing run. What should you configure?
Medium96A company uses Azure AI Search to index customer support transcripts. They want to enable users to find relevant answers by asking natural language questions. Which feature should they enable in the search service?
Easy97You are using Azure AI Language Service to extract key phrases from customer reviews. You notice that for reviews containing the word 'not good', the service sometimes extracts 'good' as a key phrase. What is the most likely reason?
Easy98You are using Azure AI Search to build a knowledge base for a customer support portal. The index includes a 'sentiment' field that should be populated using the Sentiment skill. However, the sentiment scores are not being written to the index. The skillset runs successfully. What is the most likely cause?
Medium99You are a solution architect at a legal firm. The firm wants to build a copilot using Microsoft Foundry that answers questions about case law documents stored in Azure Blob Storage. The copilot should use the Retrieval Augmented Generation (RAG) pattern with Azure AI Search as the vector store. The documents are in PDF format and include complex tables and footnotes. The solution must ensure that the answers are grounded in the documents and that the copilot can handle follow-up questions. You need to design the ingestion pipeline. Which approach should you take?
Medium100Your organization is using Azure AI Search to index a large collection of PDF documents stored in Azure Blob Storage. The index currently returns search results, but users complain that the results are not relevant when they search using natural language phrases. You need to improve the relevance of search results without rewriting the application. What should you do?
Medium101You are creating an Azure AI Search indexer that processes documents from Azure Blob Storage. The documents include PDFs and images. You need to extract both text and image content, and then use the image content to generate captions via the Image Analysis skill. Which indexer configuration is required to enable image extraction and passing images to the skillset?
MediumOther domains
All AI-102 exam domains
Frequently asked questions
- What does the Implement knowledge mining and information extraction solutions domain cover on the AI-102 exam?
- Candidates must design and implement Azure AI Search enrichment pipelines: create data sources, indexers, skillsets, and indexes that extract and enrich content. The critical skill is correctly mapping enriched outputs to searchable and filterable index fields for the required query behavior.
- How many questions are in this domain?
- This page lists all 101 Implement knowledge mining and information extraction solutions questions in the AI-102 question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Implement knowledge mining and information extraction solutions questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.