Courseiva

CCNA Implement knowledge mining and information extraction solutions Questions

26 of 101 questions · Page 2/2 · Implement knowledge mining and information extraction solutions · Answers revealed

76
MCQmedium

You are reviewing an index definition created with PowerShell. The index is used for a knowledge mining solution that extracts people and organizations from documents. Users report that when they type partial names in the search bar, the suggester does not return suggestions. What is the most likely reason?

A.The people and organizations fields should be Edm.String instead of Collection(Edm.String)
B.The id field is not defined as a key in the index
C.The suggester sourceFields do not include people or organizations
D.The suggester searchMode should be 'analyzingInfixMatching' which is incorrect
AnswerC

The suggester only generates suggestions from the fields listed in its sourceFields collection. If people and organizations are absent from that list, partial-name queries return nothing, regardless of the analyser or searchMode settings. Adding those extracted fields to sourceFields satisfies the stem's requirement that partial names produce suggestions.

Why this answer

A suggester in Azure Cognitive Search only returns suggestions for fields explicitly listed in its `sourceFields` property. If the `people` and `organizations` fields are not included in `sourceFields`, the suggester cannot match partial names typed in the search bar, even if those fields are indexed and searchable. The suggester relies on prefix matching against the specified source fields to generate suggestions.

Exam trap

The trap here is that candidates may focus on data types or index keys instead of recognizing that the suggester's `sourceFields` property explicitly controls which fields participate in suggestion generation, a detail frequently tested in AI-102.

How to eliminate wrong answers

Option A is wrong because `Collection(Edm.String)` is the appropriate type for fields that contain multiple values (e.g., multiple people or organizations per document); changing them to `Edm.String` would lose multi-value support and is not related to suggester functionality. Option B is wrong because the `id` field being defined as a key is required for any index, but its presence or absence does not affect whether a suggester returns suggestions for other fields. Option D is wrong because `analyzingInfixMatching` is not a valid `searchMode` for a suggester; the correct `searchMode` is `analyzingInfixMatching` (note the typo in the option) but the real issue is that the suggester's `sourceFields` must include the target fields, not the search mode.

77
MCQmedium

You are implementing an Azure AI Search enrichment pipeline that extracts text from PDF documents stored in Azure Blob Storage. The PDFs are scanned images with no embedded text layer. You need to ensure the extracted text is available for downstream skills. Which skill should you add to the skillset?

A.Microsoft.Skills.Text.SplitSkill
B.Microsoft.Skills.Text.MergeSkill
C.Microsoft.Skills.Custom.WebApiSkill
D.Microsoft.Skills.Vision.OcrSkill
AnswerD

The OCR skill uses the Computer Vision Read API to extract text from images, including scanned PDF pages. It is designed for exactly this scenario where documents lack a text layer. The skill outputs text and layout information that can be mapped to index fields. This is the correct choice because it directly addresses the need to perform optical character recognition on image-based PDFs within the enrichment pipeline.

Why this answer

The OCR skill is specifically designed to extract text from images, including scanned PDFs, within an Azure AI Search enrichment pipeline. It leverages the Computer Vision Read API to recognize text and output it for further processing. Other skills like Split, Merge, or custom Web API do not provide built-in OCR capabilities.

Therefore, adding the OCR skill is the correct approach to make the scanned PDF content searchable.

Exam trap

The trap here is assuming that the built-in document extraction skill automatically handles scanned images, when in fact it only extracts text from documents with an embedded text layer.

78
Multi-Selecteasy

Which TWO Azure AI services are most appropriate for extracting text from images and recognizing handwritten text?

Select 2 answers
A.Azure AI Document Intelligence
B.Azure AI Speech
C.Azure AI Vision
D.Azure AI Search
E.Azure AI Language
AnswersA, C

Azure AI Document Intelligence reads text from images and PDFs, and its Read model handles both printed and handwritten content. That dual capability satisfies the stem's requirement to extract text from images and recognise handwriting in one service.

Why this answer

Azure AI Document Intelligence (option A) is correct because it provides prebuilt and custom models that perform OCR and extract text, key-value pairs, and tables from documents, including support for handwritten text via its Read and Layout models. Azure AI Vision (option C) is correct because its Image Analysis and OCR capabilities, including the Read API, are specifically designed to extract printed and handwritten text from images. Azure AI Speech (option B) is incorrect because it handles speech-to-text and text-to-speech, not text extraction from images.

Azure AI Search (option D) is incorrect because it is a search indexing and query service, not an OCR or handwriting recognition service. Azure AI Language (option E) is incorrect because it focuses on natural language processing tasks such as sentiment analysis, entity recognition, and translation, not image text extraction.

Exam trap

AI-102 often tests whether candidates conflate text-analytics services (Language, Search) with OCR services — the exam wants the two services that actually perform image and handwriting text extraction.

79
Multi-Selecthard

Your organization is using Azure AI Document Intelligence to process a mix of invoices and purchase orders. You need to ensure that documents are correctly classified before extraction. Which THREE steps should you take?

Select 3 answers
A.Train the classification model with one sample per type
B.Create a custom classification model in Document Intelligence
C.Label at least 5 samples for each document type
D.Chain the classification model with extraction models
E.Use the prebuilt invoice and purchase order models for classification
AnswersB, C, D

A custom classification model in Document Intelligence identifies each document's type before extraction, letting the pipeline route invoices and purchase orders to the appropriate extraction model. This directly satisfies the requirement to classify correctly beforehand.

Why this answer

Option B is correct because a custom classification model in Azure AI Document Intelligence is the required resource for identifying document type before extraction, since prebuilt models extract fields but do not perform custom document classification. Option C is correct because training a custom classification model requires a minimum of 5 labeled samples per document type, which provides the model with enough examples to distinguish invoices from purchase orders. Option D is correct because after classification, you must chain the classification model with the appropriate extraction models so each document is routed to the correct extraction model for field retrieval.

Option A is incorrect because one sample per type is insufficient; the minimum is 5 labeled samples per document type. Option E is incorrect because prebuilt invoice and purchase order models are extraction models, not classification models, and they do not classify documents before extraction.

Exam trap

The trap here is that candidates confuse prebuilt models (which perform extraction) with classification capabilities, assuming they can automatically identify document types without a dedicated classifier.

80
MCQmedium

You are troubleshooting an Azure AI Search indexer that fails to index a PDF file stored in Azure Blob Storage. The error message indicates that the document is encrypted. What is the most likely cause and solution?

A.The indexer is not configured with the PDF parser; set the parsing mode
B.The file format is unsupported; convert to PDF/A
C.The file is too large; split it into smaller parts
D.The PDF is encrypted; remove encryption before indexing
AnswerD

Azure AI Search's document cracking cannot decrypt password-protected or rights-managed PDFs, so the indexer reports the document as encrypted and skips it. Removing encryption before indexing, or supplying an unencrypted copy, lets the blob indexer extract text successfully.

Why this answer

The most likely cause is that the PDF is encrypted, and the solution is to remove encryption before indexing. Azure AI Search indexers cannot decrypt password-protected or encrypted PDFs; they require unencrypted content to extract text. Therefore, the file must be decrypted prior to indexing.

Exam trap

AI-102 often tests the assumption that indexer errors are due to configuration or format issues, but encrypted files require explicit decryption, which is a common oversight.

How to eliminate wrong answers

Option A is wrong because the error explicitly indicates encryption, not a missing PDF parser; the parser is likely configured correctly. Option B is wrong because PDF/A is a format for archiving, not a solution for encryption; the file format is supported. Option C is wrong because file size would produce a different error, not an encryption error.

81
MCQmedium

Your organization is implementing a knowledge mining solution for a research institute that needs to extract chemical compound names and reactions from scientific articles in PDF format. The solution must use a custom model because the scientific terminology is not covered by built-in skills. You have trained a custom model using Azure AI Language's custom entity recognition (NER) and deployed it as a REST endpoint. You are using Azure AI Search with a skillset. How should you integrate the custom NER model into the enrichment pipeline?

A.Create a custom skill that calls the custom NER endpoint and map the output to the index fields.
B.Use a Language Understanding (LUIS) app to extract entities and call it from a custom skill.
C.Use the built-in Entity Recognition skill and configure it with your custom model's endpoint.
D.Configure the indexer to call the custom NER endpoint directly during indexing.
AnswerA

A custom skill in the skillset invokes the deployed custom NER REST endpoint, passing enriched document text and writing returned entities into index fields. Built-in skills cannot cover the scientific terminology, so wrapping the endpoint as a custom skill is the only integration path.

Why this answer

To integrate a custom NER model into an Azure AI Search enrichment pipeline, you must create a custom skill that calls the custom NER endpoint. The custom skill is a web API that the skillset invokes, and it can map the JSON output to index fields. This allows the enrichment pipeline to use the custom model's predictions.

Exam trap

The trap is assuming that built-in skills can be customized with a custom endpoint; candidates may choose the built-in Entity Recognition skill, but it does not support custom models.

How to eliminate wrong answers

Option B is wrong because LUIS is for intent recognition and conversational language understanding, not for custom entity recognition from text; it is not suitable for extracting chemical compound names. Option C is wrong because the built-in Entity Recognition skill uses a pre-trained model and cannot be configured with a custom model's endpoint; it does not support custom models. Option D is wrong because the indexer cannot directly call a custom NER endpoint; it must go through a skillset with a custom skill.

82
Multi-Selectmedium

Which TWO configurations are required to enable Azure AI Search to index content from an Azure SQL database?

Select 2 answers
A.Create a custom skillset for data enrichment
B.Configure semantic ranking on the index
C.Enable change tracking on the Azure SQL table
D.Define a data source connection to the Azure SQL database
E.Enable high availability on the Azure SQL database
AnswersC, D

Change tracking lets the indexer detect which rows were inserted, updated, or deleted since the last run, so incremental re-indexing stays accurate without full reloads. Without it, the SQL indexer cannot identify changed rows, and the required high-water mark column for incremental indexing is unavailable.

Why this answer

Option C is correct because Azure AI Search's SQL indexer relies on change tracking (or a rowversion/timestamp column) to detect which rows have been inserted, updated, or deleted since the last indexing run, so enabling change tracking on the Azure SQL table is required for incremental indexing. Option D is correct because the indexer must be given a data source object that specifies the connection string, table or view, and change-tracking policy for the Azure SQL database, which is the mandatory link between the search service and the SQL data. Option A is not required because a skillset is only needed for AI enrichment (for example OCR, key phrase extraction, or embedding generation), not for basic SQL-to-index ingestion.

Option B is not required because semantic ranking is an optional query-time feature that improves relevance; it does not affect whether content can be indexed. Option E is not required because high availability is a resilience/uptime configuration for the SQL database and is unrelated to the indexer's ability to read and index data.

Exam trap

The trap here is that candidates often confuse optional enrichment features (like custom skillsets or semantic ranking) with mandatory infrastructure requirements for data ingestion in Azure AI Search, leading them to select those options instead of the core connectivity and change tracking configurations.

83
Multi-Selecthard

Which THREE factors should you consider when designing a knowledge mining solution that uses Azure AI Search and custom skills to extract insights from large volumes of documents?

Select 3 answers
A.The number of knowledge store projections affects indexing speed
B.The maximum execution time of the custom skill must fit within the indexer timeout
C.Incremental enrichment should be enabled to avoid reprocessing unchanged documents
D.Semantic ranking configuration must be included in the skillset
E.The custom skill should be stateless and idempotent to allow parallel execution
AnswersB, C, E

Custom skills execute synchronously inside the indexer's enrichment pipeline, so a skill exceeding the indexer timeout aborts the run and leaves documents partially enriched. Sizing skill execution against that timeout is therefore a core design constraint for large document volumes.

Why this answer

Option B is correct because custom skills run inside the indexer execution pipeline, and each skill invocation must complete within the indexer's timeout limits (for example, the default HTTP timeout of 3 minutes 30 seconds for a WebApiSkill); a long-running custom skill will cause the indexer to fail or time out. Option C is correct because enabling incremental enrichment (by setting the indexer's data source change detection and using a high-water mark) lets Azure AI Search skip documents whose content has not changed, avoiding unnecessary re-invocation of expensive custom skills and reducing cost and processing time. Option E is correct because the indexer can invoke skills concurrently across documents, so a custom skill must not depend on shared mutable state or on a specific call order; making it stateless and idempotent ensures consistent, repeatable enrichment results under parallel execution.

Option A is not a primary design factor here because knowledge store projections are an output concern and do not fundamentally constrain the design of the custom-skill enrichment pipeline in the way timeout, incremental enrichment, and statelessness do. Option D is incorrect because semantic ranking is configured on the search index/query side (semantic configuration), not as a required element of the skillset, so it is not a factor in designing the custom-skill extraction pipeline.

Exam trap

The trap here is that candidates often confuse knowledge store projections (output storage) with indexing performance, or assume semantic ranking is a mandatory skillset component, when in fact it is an optional query-time feature.

84
MCQhard

You are designing a solution to extract customer names and addresses from scanned handwritten forms. The forms are stored as images in Azure Blob Storage. The extraction must achieve high accuracy with minimal manual review. Which combination of Azure AI services should you use?

A.Azure AI Document Intelligence with prebuilt invoice and receipt models
B.Azure AI Document Intelligence with a custom model trained on handwritten forms
C.Azure AI Language Service with custom Named Entity Recognition (NER)
D.Azure AI Computer Vision with OCR and Azure AI Search
AnswerB

Azure AI Document Intelligence custom models learn your forms' specific layout and handwriting variations from labelled samples, directly satisfying the high-accuracy, minimal-review constraint. Unlike the prebuilt read model, a custom neural model handles the unpredictable field positions and cursive styles typical of scanned handwritten forms stored in Blob Storage.

Why this answer

Azure AI Document Intelligence's custom model capability allows you to train a model specifically on handwritten forms, enabling it to learn the unique handwriting patterns and layout structures present in your scanned documents. This tailored approach achieves high accuracy with minimal manual review, as the model is optimized for your specific form type rather than generic invoice or receipt templates.

Exam trap

The trap here is that candidates often confuse prebuilt models (which work well for printed documents) with custom models (which are necessary for handwritten forms), or they assume OCR alone is sufficient without considering the need for structured field extraction.

How to eliminate wrong answers

Option A is wrong because prebuilt invoice and receipt models are designed for structured, printed documents and cannot reliably extract handwritten text with high accuracy, leading to increased manual review. Option C is wrong because Azure AI Language Service with custom NER extracts entities from text but does not perform OCR or handle image-based handwritten input, so it cannot process scanned forms directly. Option D is wrong because Azure AI Computer Vision with OCR provides raw text extraction but lacks the document understanding and field-level extraction capabilities needed to accurately parse structured fields like customer names and addresses from forms, and Azure AI Search is for indexing and querying, not extraction.

85
MCQeasy

You are using Azure AI Search to index a set of contracts. You need to extract named entities such as organizations, people, and dates from the contract text and store them as separate fields in the index. Which skill should you add to the skillset?

A.Text Translation skill
B.Language Detection skill
C.Key Phrases skill
D.Entity Recognition skill
AnswerD

The Entity Recognition skill uses Azure AI Language to detect named entities in text and returns them with type and subtype information, such as Organization, Person, and DateTime. This allows each entity category to be mapped to its own index field. It is the correct skill for extracting typed entities from contract text.

Why this answer

Entity Recognition is the built-in Azure AI Search skill that calls Azure AI Language to detect named entities with type and subtype metadata. Its output includes categorized entities such as organizations, people, and dates, which can be projected into separate index fields. The other language skills perform different functions and do not produce typed entity output.

Exam trap

The trap here is confusing Key Phrases, which returns a flat list of salient phrases, with Entity Recognition, which returns categorized entities with type information.

86
MCQhard

Your team is using Azure AI Search to index a large collection of technical manuals. Users report that searches for 'disk failure' do not return relevant results because the manuals use terms like 'hard drive crash'. Which feature should you implement to improve recall?

A.Apply a filter
B.Configure a scoring profile
C.Enable semantic search
D.Add a synonym map to the index
AnswerD

A synonym map expands queries so 'disk failure' also matches 'hard drive crash', raising recall for terminology mismatches. It applies at query time against the index, directly addressing the vocabulary gap between user phrasing and manual wording.

Why this answer

A synonym map in Azure AI Search allows you to define equivalent terms (e.g., 'disk failure' = 'hard drive crash') so that queries automatically expand to include synonyms. This directly addresses the vocabulary mismatch between user queries and indexed content, improving recall without requiring changes to the documents or queries.

Exam trap

The trap here is that candidates often confuse semantic search (which improves ranking via language models) with synonym expansion (which directly addresses vocabulary mismatch by broadening the query), leading them to choose option C instead of D.

How to eliminate wrong answers

Option A is wrong because a filter narrows results based on structured field criteria (e.g., date range, category) and does not expand query terms to match synonyms. Option B is wrong because a scoring profile boosts relevance ranking based on fields or functions (e.g., freshness, magnitude) but does not alter which documents match the query. Option C is wrong because semantic search re-ranks results using language understanding to improve relevance, but it does not expand the query to include synonymous terms; it still relies on the original query tokens for matching.

87
MCQeasy

You are designing an Azure AI Search solution that indexes documents from an Azure SQL Database. The documents include a field named 'content' that contains HTML markup. You need to strip the HTML tags and extract only the plain text before applying further enrichment. Which built-in skill should you use?

A.Text Merger skill
B.Text Split skill
C.HTML Strip skill
D.Language Detection skill
AnswerC

The HTML Strip skill is a built-in cognitive skill that removes HTML tags from a string and returns plain text. It is designed exactly for this scenario: cleaning HTML content before further processing. By using this skill, you ensure that subsequent enrichment skills receive clean text, improving the accuracy of language detection, entity recognition, and other NLP tasks.

Why this answer

The HTML Strip skill is specifically designed to remove HTML markup and return plain text. It is the correct choice for cleaning HTML content before applying other enrichment skills. The Text Merger and Text Split skills manipulate text structure but do not remove HTML.

Language Detection analyzes language but does not alter the text content.

Exam trap

The trap here is assuming that Text Split or Text Merger can also clean HTML, when they only restructure text without removing markup.

88
MCQeasy

You need to extract key-value pairs from scanned forms as part of a knowledge mining solution. Which Azure AI service should you use?

A.Azure AI Vision
B.Azure AI Language
C.Azure AI Search
D.Azure AI Document Intelligence
AnswerD

Document Intelligence provides prebuilt and custom models that return structured key-value pairs from scanned forms, unlike pure OCR which yields unstructured text. This directly satisfies the requirement to extract key-value pairs from scanned forms within the knowledge mining pipeline.

Why this answer

Azure AI Document Intelligence (formerly Form Recognizer) is the correct service because it is specifically designed to extract key-value pairs, tables, and structured data from scanned forms and documents using prebuilt and custom models. This aligns directly with the requirement for knowledge mining from scanned forms.

Exam trap

The trap here is that candidates often confuse Azure AI Vision's OCR capability with form-specific extraction, not realizing that Document Intelligence is the dedicated service for key-value pair extraction from scanned forms, while Vision only provides raw text coordinates without semantic understanding.

How to eliminate wrong answers

Option A is wrong because Azure AI Vision provides image analysis capabilities like OCR, object detection, and captioning, but it does not have native support for extracting key-value pairs from forms; it would require additional processing to structure the data. Option B is wrong because Azure AI Language focuses on text analytics, sentiment analysis, and entity recognition from written text, not from scanned forms or document layouts. Option C is wrong because Azure AI Search is a search indexing and query service that can index extracted data but does not perform the extraction itself; it relies on other services like Document Intelligence to provide the structured input.

89
Multi-Selecteasy

Which TWO Azure AI services can be used to extract text from images as part of a knowledge mining pipeline?

Select 2 answers
A.Azure AI Language
B.Azure AI Document Intelligence
C.Azure AI Computer Vision
D.Azure AI Video Indexer
E.Azure AI Custom Vision
AnswersB, C

Document Intelligence's Read model performs OCR, returning printed and handwritten text from images and PDFs, and integrates into enrichment pipelines via the built-in Document Intelligence skill. It satisfies the requirement to extract text from images as part of knowledge mining.

Why this answer

Azure AI Document Intelligence (option B) is correct because its Read and Layout models perform OCR on documents and images, extracting printed and handwritten text along with structure for downstream knowledge mining enrichment. Azure AI Computer Vision (option C) is correct because its Read OCR feature (Image Analysis / Read API) extracts printed and handwritten text from images, a core skill in Azure AI Search cognitive skillsets. Azure AI Language (option A) is wrong because it handles text analytics such as entity recognition, sentiment, and key phrase extraction, not image OCR.

Azure AI Video Indexer (option D) is wrong because it targets video and audio insights like transcription and face tracking, not still-image text extraction. Azure AI Custom Vision (option E) is wrong because it trains image classification and object detection models, not text recognition.

Exam trap

The trap here is that candidates often confuse Azure AI Computer Vision's OCR capabilities with Azure AI Document Intelligence, but Document Intelligence is the dedicated service for structured document extraction in knowledge mining, while Computer Vision provides general-purpose image analysis and OCR without the same level of document-specific parsing.

90
MCQeasy

Your team has built a knowledge mining pipeline using Azure AI Search and Document Intelligence. After ingestion, you notice that some documents are not appearing in search results. What is the most likely cause?

A.The indexer encountered errors and marked the documents as failed
B.The index does not have a semantic configuration
C.The search service has insufficient replicas
D.The search service is throttled due to high query volume
AnswerA

Indexers record per-document status during enrichment; documents failing skill execution or field mapping are marked failed and omitted from the index, so they never surface in queries. Checking indexer execution history and error details identifies the specific failing documents.

Why this answer

When an indexer runs, it processes each document and can encounter errors such as unsupported file formats, corrupt content, or permission issues. If a document fails during indexing, the indexer records the error and does not add that document to the index, making it invisible to search queries. This is the most direct cause of missing documents after ingestion.

Other options affect search behavior but not whether documents are indexed in the first place.

Exam trap

AI-102 often tests the misconception that search service configuration (like replicas or semantic ranker) affects document indexing, when in fact indexing failures are the primary cause of missing documents.

How to eliminate wrong answers

Option B is wrong because a semantic configuration enhances ranking and relevance but is not required for documents to appear in search results; without it, documents are still indexed and searchable via full-text search. Option C is wrong because insufficient replicas affect query throughput and high availability, not whether documents are indexed; indexing is handled by indexers and the number of replicas does not determine document inclusion. Option D is wrong because throttling due to high query volume impacts query performance and may cause rate-limiting, but it does not prevent documents from being indexed; indexing and querying are separate workloads.

91
Multi-Selectmedium

Which TWO actions should you take to optimize the performance of an Azure AI Search solution that indexes large volumes of data?

Select 2 answers
A.Use the appropriate search tier (S1, S2, etc.) based on document size
B.Use the free tier for production workloads
C.Increase the replica count for better indexing throughput
D.Batch documents in groups of up to 1000 per index operation
E.Disable scoring profiles to speed up indexing
AnswersA, D

Larger document volumes demand more storage, processing power and replica/partition capacity. Selecting a tier such as S1 or S2 sized to document volume directly addresses the stem's large-scale indexing constraint, preventing throttling and slow enrichment throughput.

Why this answer

Option A is correct because Azure AI Search tiers (Basic, S1, S2, S3, etc.) differ in storage size, document limits, partitions, and indexer throughput, so selecting a tier that matches your document size and volume is essential for indexing performance. Option D is correct because the Azure AI Search indexing API supports batching up to 1000 documents (or about 16 MB of payload) per index operation, which amortizes network and service overhead and dramatically improves indexing throughput. Option B is wrong because the Free tier is for evaluation only, with strict limits on storage, indexes, and indexers, and is not suitable for production workloads.

Option C is wrong because replicas provide high availability and increased query throughput, not indexing throughput; indexing scale is driven by partitions. Option E is wrong because scoring profiles affect query-time ranking, not indexing speed, and disabling them would not optimize indexing performance.

Exam trap

The trap here is confusing replicas (which scale query performance) with partitions (which scale indexing throughput), leading candidates to incorrectly select option C as a way to improve indexing speed.

92
MCQmedium

You are building an Azure AI Search enrichment pipeline that enriches documents with key phrases detected by Azure AI Language. You need the key phrases to be returned as a structured collection that can be mapped to a Collection(Edm.String) field in the search index, and you must avoid storing the enriched text in the enrichment cache. Which output mapping configuration should you use in the indexer?

A.Define an outputFieldMapping with a source of /document/keyPhrases and a targetField of keyPhrases.
B.Define an outputFieldMapping with a source of /document/pages/*/keyPhrases and a targetField of keyPhrases.
C.Define an outputFieldMapping with a source of /document/keyPhrases and a targetField of keyPhrases, and set the indexer cache to disabled.
D.Define an outputFieldMapping with a source of /document/keyPhrases and a targetField of keyPhrases, and mark the field as retrievable and filterable.
AnswerA

The Key Phrases skill emits its result at the enrichment tree path /document/keyPhrases as a JSON array of strings. Mapping that source directly to a Collection(Edm.String) index field stores the phrases as a structured collection. Because the mapping targets an index field rather than a cache-only enrichment, the correct configuration preserves the intended behavior without persisting the enriched text in the cache.

Why this answer

The Key Phrases skill writes its output to the enrichment tree at /document/keyPhrases as an array of strings. To persist that array into a Collection(Edm.String) field, the indexer needs an outputFieldMapping whose source is that exact path and whose target is the index field. Cache behavior is configured separately on the indexer and is not part of output mapping.

Exam trap

The trap here is assuming the Key Phrases skill emits per-page results under /document/pages/*/keyPhrases, when it actually writes a single array at /document/keyPhrases for the whole document.

93
MCQmedium

You are designing a knowledge mining solution for a manufacturing company that needs to extract information from equipment maintenance manuals. The manuals are in multiple languages (English, French, German). You need to ensure that the extracted content is searchable in English only. Which approach should you use?

A.Use the Entity Recognition skill to extract entities and then index entities only.
B.Use the Language Detection skill to identify language and then index all content as-is.
C.Use the Text Translation skill to translate all content to English during indexing.
D.Use the Key Phrase Extraction skill to extract key phrases and then index them.
AnswerC

The Text Translation skill translates French and German content into English during indexing, so all extracted text is stored in English. This makes the index searchable in English only, regardless of the manuals' original languages.

Why this answer

Azure AI Search's Text Translation skill (backed by Azure AI Translator) translates non-English content into a target language during the enrichment pipeline, so all indexed content is normalized to English and searchable in English only. This directly satisfies the requirement to make multilingual manuals searchable in English. The skill runs at indexing time, so the index contains only English text.

Exam trap

AI-102 often tests the difference between enrichment skills that transform content (Translation) versus those that only extract metadata (Entity Recognition, Key Phrase Extraction), causing candidates to pick extraction skills when translation is required.

How to eliminate wrong answers

Option A is wrong because Entity Recognition extracts entities (people, places, organizations) but does not translate content, so French and German text would remain untranslated and unsearchable in English. Option B is wrong because Language Detection only tags the language; indexing content as-is leaves non-English text in the index, failing the English-only search requirement. Option D is wrong because Key Phrase Extraction pulls salient phrases but does not translate, so German/French phrases would still be indexed in their original language.

94
MCQhard

An organization uses Azure AI Search to power an internal knowledge base. They notice that search results are returning irrelevant documents. The index includes a 'content' field with full text and a 'tags' field with metadata. Users often search for specific terms that appear in the 'tags' field. How should you configure the search index to improve relevance?

A.Add a custom scoring profile based on freshness.
B.Configure a scoring profile with a higher weight for the 'tags' field.
C.Set the 'tags' field to use the 'keyword' analyzer.
D.Enable semantic search on the 'content' field.
AnswerB

Scoring profiles apply field weights during query evaluation, so boosting the 'tags' field raises documents whose metadata matches the user's terms. This directly addresses the relevance problem, since tags carry the specific terms users search, whereas the full-text 'content' field dilutes matching scores.

Why this answer

Configuring a scoring profile with a higher weight for the 'tags' field increases the relevance score of documents where search terms match the tags, thereby prioritizing those results. Option A (freshness-based scoring) would favor newer documents but does not address matching on tags. Option C sets the 'tags' field to use the 'keyword' analyzer, which changes tokenization but does not adjust field weighting.

Option D enables semantic search on the 'content' field, which enhances understanding of natural language queries but does not specifically boost the weight of the tags field.

95
MCQmedium

You have an Azure AI Search indexer that uses a custom skill hosted in an Azure Function to normalize product codes. The function occasionally returns HTTP 429 responses. You need the indexer to retry these calls automatically without failing the entire indexing run. What should you configure?

A.Set the batchSize property on the indexer to a smaller value.
B.Implement retry logic inside the Azure Function and return a success response after retries.
C.Add a retryPolicy to the custom skill definition in the skillset.
D.Configure the indexer's maxFailedItems and maxFailedItemsPerBatch to tolerate failures.
AnswerB

Because the throttling originates from the custom skill's downstream dependency, the function itself should handle transient failures using retry policies such as exponential backoff. Returning a successful response after internal retries prevents the indexer from seeing 429 errors. This is the supported and reliable way to make custom skills resilient to intermittent throttling.

Why this answer

Azure AI Search does not provide a retryPolicy on custom skills. When a custom skill returns transient errors such as HTTP 429, the correct pattern is to implement retry logic inside the skill implementation, for example using exponential backoff in the Azure Function, so the indexer receives a successful response once the downstream call succeeds.

Exam trap

The trap here is looking for a retryPolicy property on the custom skill, which Azure AI Search does not support; retries must be handled inside the skill code.

96
MCQeasy

A company uses Azure AI Search to index customer support transcripts. They want to enable users to find relevant answers by asking natural language questions. Which feature should they enable in the search service?

A.Semantic search
B.Synonym maps
C.Cognitive skills
D.Knowledge mining
AnswerA

Semantic search adds a reranking layer over results using Microsoft's language models, matching natural-language questions to the most relevant passages in the transcripts. This directly satisfies the stem's requirement to find answers by asking questions, since keyword search alone cannot interpret query intent or rank by semantic relevance.

Why this answer

Semantic search in Azure AI Search enhances the ranking of results by using language understanding models to re-rank matches based on semantic relevance to the query, enabling users to ask natural language questions and get more relevant answers. It is the feature designed to improve relevance for natural language queries.

Exam trap

AI-102 often tests the distinction between semantic search and other Azure AI Search features — candidates may confuse semantic search with synonym maps or cognitive skills, which serve different purposes.

How to eliminate wrong answers

Option B is wrong because synonym maps expand queries with equivalent terms but do not provide semantic understanding or natural language question answering. Option C is wrong because cognitive skills are used during indexing to enrich content (e.g., entity recognition, OCR), not to enable natural language querying at search time. Option D is wrong because knowledge mining is a broader solution pattern for extracting insights from large content sets, not a specific search feature for natural language question answering.

97
MCQeasy

You are using Azure AI Language Service to extract key phrases from customer reviews. You notice that for reviews containing the word 'not good', the service sometimes extracts 'good' as a key phrase. What is the most likely reason?

A.The language detection model misidentified the language
B.You need to set a confidence threshold to exclude negative phrases
C.Key phrase extraction does not consider negation
D.The service is not trained on your specific domain
AnswerC

Key phrase extraction identifies statistically significant terms without parsing negation, so 'good' is extracted as a standalone phrase from 'not good'. The model treats words independently rather than understanding that the negation inverts the sentiment.

Why this answer

Key phrase extraction in Azure AI Language Service uses a statistical model that identifies significant terms based on frequency and context, but it does not inherently understand negation. When the phrase 'not good' appears, the model may still extract 'good' as a key phrase because it recognizes 'good' as a high-value term, ignoring the negation. This is a known limitation of the feature, as it focuses on noun phrases and important terms rather than sentiment or negated constructs.

Exam trap

The trap here is that candidates often assume Azure AI Language Service handles negation across all features, but key phrase extraction explicitly does not consider negation, unlike sentiment analysis which does.

How to eliminate wrong answers

Option A is wrong because language detection is a separate step that identifies the language of the text; misidentification would cause incorrect processing but would not specifically cause 'good' to be extracted from 'not good'. Option B is wrong because confidence thresholds filter out low-confidence phrases, not negative phrases; the service does not have a built-in mechanism to exclude negated terms via threshold settings. Option D is wrong because while domain-specific training can improve accuracy, the core issue here is a fundamental limitation of the key phrase extraction model's handling of negation, not a lack of domain adaptation.

98
MCQmedium

You are using Azure AI Search to build a knowledge base for a customer support portal. The index includes a 'sentiment' field that should be populated using the Sentiment skill. However, the sentiment scores are not being written to the index. The skillset runs successfully. What is the most likely cause?

A.The output field mapping for 'sentiment' is missing or incorrectly defined in the indexer.
B.The Sentiment skill is not correctly configured in the skillset.
C.The indexer is in a failed state and not processing documents.
D.The sentiment field in the index is of type 'Collection(Edm.String)' but the skill outputs a double.
AnswerA

Skillset execution writes enriched values into the enrichment tree, not the index. Without an output field mapping linking the sentiment skill output to the target index field, scores are discarded, so the index remains empty despite successful skillset execution.

Why this answer

The Sentiment skill outputs a 'double' value for sentiment score, but the indexer requires an explicit output field mapping to write that value into the index's 'sentiment' field. Even when a skillset runs successfully, without a correct output field mapping in the indexer definition, the skill's output is not transferred to the index. The indexer's field mappings control how enriched data flows from the skillset's output nodes to the index fields.

Exam trap

The trap here is that candidates assume a successful skillset execution guarantees data is written to the index, but Azure AI Search requires explicit output field mappings in the indexer to bridge skill outputs to index fields, and this step is often overlooked.

How to eliminate wrong answers

Option B is wrong because the question states the skillset runs successfully, meaning the Sentiment skill itself is correctly configured and executed without errors. Option C is wrong because the indexer is explicitly described as running successfully, not in a failed state, so it is processing documents. Option D is wrong because the Sentiment skill outputs a double (a numeric score between 0 and 1), and if the index field were of type 'Collection(Edm.String)', the mismatch would cause an indexer error or warning, but the question says the skillset runs successfully — the issue is the missing mapping, not a type conflict.

99
MCQmedium

You are a solution architect at a legal firm. The firm wants to build a copilot using Microsoft Foundry that answers questions about case law documents stored in Azure Blob Storage. The copilot should use the Retrieval Augmented Generation (RAG) pattern with Azure AI Search as the vector store. The documents are in PDF format and include complex tables and footnotes. The solution must ensure that the answers are grounded in the documents and that the copilot can handle follow-up questions. You need to design the ingestion pipeline. Which approach should you take?

A.Use Azure AI Vision OCR to extract text, split by page, and use Azure AI Search keyword search
B.Use Azure AI Document Intelligence prebuilt-read model, chunk by character count, and use Azure AI Search with semantic ranking
C.Use Azure AI Document Intelligence to extract content, then chunk by headings and paragraphs, generate embeddings using Azure OpenAI, and index in Azure AI Search with vector search
D.Use Azure AI Language to extract key phrases, create a non-vector index, and use simple search
AnswerC

Document Intelligence's layout model preserves complex tables and footnotes that plain PDF text extraction loses, satisfying the grounding requirement. Heading and paragraph chunking keeps semantic units intact for retrieval, and embeddings indexed in Azure AI Search with vector search let the copilot retrieve relevant passages for follow-up questions.

Why this answer

It uses Azure AI Document Intelligence to accurately extract content from PDFs (including complex tables and footnotes), then chunks by headings and paragraphs to preserve document structure, generates embeddings via Azure OpenAI for semantic understanding, and indexes in Azure AI Search with vector search to enable RAG-based, grounded answers with follow-up support.

Exam trap

Microsoft often tests the misconception that simple OCR or keyword search is sufficient for complex documents, but the trap here is that legal documents with tables and footnotes require structure-aware extraction and vector search to support grounded, conversational RAG.

How to eliminate wrong answers

Option A is wrong because Azure AI Vision OCR is designed for image-based text extraction and lacks the ability to handle complex tables and footnotes in PDFs; splitting by page ignores document structure, and keyword search alone cannot support semantic understanding or follow-up questions. Option B is wrong because the prebuilt-read model extracts raw text without preserving table/footnote structure, chunking by character count breaks logical content boundaries, and semantic ranking on keyword search does not provide the vector-based retrieval needed for RAG. Option D is wrong because key phrase extraction loses document context and structure, a non-vector index cannot support semantic similarity search, and simple search cannot ground answers in document content or handle follow-up questions effectively.

100
MCQmedium

Your organization is using Azure AI Search to index a large collection of PDF documents stored in Azure Blob Storage. The index currently returns search results, but users complain that the results are not relevant when they search using natural language phrases. You need to improve the relevance of search results without rewriting the application. What should you do?

A.Increase the number of replicas for the search service to improve query performance.
B.Create a new index with a blob indexer that uses the 'content' field only.
C.Enable semantic search on the index and configure a semantic configuration.
D.Configure a custom analyzer on the index to handle stop words and synonyms.
AnswerC

Enabling semantic search and defining a semantic configuration adds L2 reranking and captions over the existing index, so natural-language phrase queries return more relevant results. No application rewrite is needed because the same query endpoint is used, just with the semantic query type.

Why this answer

Semantic search in Azure AI Search uses advanced language models to understand the intent behind natural language queries, re-ranking results based on semantic relevance rather than just keyword matching. Enabling semantic search and configuring a semantic configuration directly addresses the user complaint about poor relevance for natural language phrases without requiring application changes.

Exam trap

The trap here is that candidates often confuse improving query performance (replicas) or basic text processing (custom analyzers) with the semantic understanding needed for natural language queries, leading them to pick options that address performance or tokenization rather than relevance.

How to eliminate wrong answers

Option A is wrong because increasing replicas only improves query throughput and availability, not the relevance or semantic understanding of search results. Option B is wrong because creating a new index with only the 'content' field would reduce the available data for matching, likely worsening relevance rather than improving it. Option D is wrong because custom analyzers handle tokenization, stop words, and synonyms at indexing time, but they do not provide the deep semantic understanding needed to interpret natural language phrases; semantic search is required for that.

101
MCQmedium

You are creating an Azure AI Search indexer that processes documents from Azure Blob Storage. The documents include PDFs and images. You need to extract both text and image content, and then use the image content to generate captions via the Image Analysis skill. Which indexer configuration is required to enable image extraction and passing images to the skillset?

A.Set the parsingMode to json.
B.Set the allowSkillsetToReadFileData parameter to true.
C.Set the parsingMode to delimitedText and specify a delimiter.
D.Set the imageAction to generateNormalizedImages in the indexer's parameters.
AnswerD

The imageAction parameter in the indexer configuration controls image extraction. Setting it to generateNormalizedImages extracts images from documents and normalizes them, making them available in the enriched document for skills like Image Analysis. This is the correct configuration to enable image content to be passed to the skillset for caption generation.

Why this answer

To extract images from documents and make them available for enrichment, you must set the imageAction parameter to generateNormalizedImages in the indexer definition. This extracts images and places them in the enriched document under /document/normalized_images/*. The Image Analysis skill can then use those images as input to generate captions.

Other parsing modes or parameters do not enable image extraction for skills.

Exam trap

The trap here is confusing parameters that allow access to file data with those that actually extract and normalize images for skills.

← PreviousPage 2 of 2 · 101 questions total

Ready to test yourself?

Try a timed practice session using only Implement knowledge mining and information extraction solutions questions.