Courseiva

CCNA Implement knowledge mining and document intelligence solutions Questions

26 questions · Implement knowledge mining and document intelligence solutions · All types, answers revealed

1
MCQeasy

A financial services firm needs to extract structured fields such as invoice date, vendor name, and total amount from thousands of PDF invoices. The documents vary in layout across vendors. The firm wants a pretrained model that requires no custom training and can return field-level confidence scores. Which Azure AI Document Intelligence model should they use?

A.Read model
B.Custom template model
C.Prebuilt invoice model
D.General document model
AnswerC

The prebuilt invoice model is trained to extract invoice-specific fields including invoice date, vendor name, and total amount, and it returns confidence scores per field. Because it is pretrained, no custom labeling or training is required, which matches the firm's requirement to process varied vendor layouts quickly.

Why this answer

The firm needs invoice-specific structured fields with confidence scores and no custom training. The prebuilt invoice model in Azure AI Document Intelligence is designed exactly for that, recognizing common invoice fields across varied layouts. General document, custom template, and Read models either lack invoice-specific fields, require training, or return only raw text.

Exam trap

The trap here is choosing a custom model when a pretrained invoice model already covers the required fields without training.

2
MCQmedium

A company uses Azure Document Intelligence to process purchase orders. They have trained a custom model with 10 labeled samples and deployed it as 'purchaseOrderModel'. When analyzing a new purchase order, the extracted 'TotalAmount' field is often incorrect. The company wants to improve the model's accuracy for this field. What should they do?

A.Retrain the model with additional labeled samples that include variations of the 'TotalAmount' field.
B.Switch to the prebuilt invoice model, which automatically extracts total amounts from purchase orders.
C.Increase the model's confidence threshold for the 'TotalAmount' field in the project settings.
D.Add a labeled sample where the 'TotalAmount' field is left blank to teach the model to ignore missing values.
AnswerA

This is correct because custom model accuracy improves with more diverse labeled data. Adding samples that cover different formats, locations, and contexts of the TotalAmount field helps the model generalize better. Azure Document Intelligence learns from labeled examples, so increasing the quantity and variety of training data directly addresses the field's extraction accuracy.

Why this answer

Custom model accuracy in Azure Document Intelligence depends heavily on the quality and quantity of labeled training data. When a specific field like TotalAmount is frequently misidentified, the most effective action is to retrain with more labeled examples that capture the field's variations. This helps the model learn the patterns and contexts associated with that field, leading to better generalization on new documents.

Exam trap

The trap here is assuming that confidence thresholds or prebuilt models can be tweaked to improve a custom model's field accuracy, when the real fix is more and better training data.

3
Drag & Dropmedium

Drag and drop the steps to set up Azure AI Content Safety for content moderation into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

The correct order for setting up Azure AI Content Safety is: first create the Content Safety resource, then obtain its endpoint and key, call the Content Safety API with the content, analyze the response for moderation results, and finally act on the analysis (e.g., block or flag content). This sequence ensures proper authentication and logical flow from setup to action.

4
MCQmedium

You are designing an Azure AI Search enrichment pipeline that processes scanned PDF invoices stored in Azure Blob Storage. You need to extract both printed text and handwritten notes from the documents before sending the content to an Azure AI Language entity recognition skill. The solution must minimize development effort and cost. Which skill should you add to the skillset?

A.Microsoft.Skills.Vision.OcrSkill
B.Microsoft.Skills.Text.MergeSkill
C.Microsoft.Skills.Custom.WebApiSkill
D.Microsoft.Skills.Text.KeyPhraseExtractionSkill
AnswerA

The built-in OcrSkill extracts printed and handwritten text from image files and PDFs, producing a text output that downstream skills can consume. It is a first-party Azure AI Search skill, so no custom code is required. Using it minimizes development effort while supporting the handwriting requirement, and it is billed through the Azure AI services attached to the search service.

Why this answer

The built-in OCR skill reads printed and handwritten text from images and PDFs and writes the extracted text into the enrichment tree. Because it is a first-party Azure AI Search skill, it requires no custom code, which satisfies the goals of low effort and predictable cost. Other text skills assume machine-readable input and cannot perform image recognition.

Exam trap

The trap here is assuming that any text-analysis skill can process scanned documents directly, when image content must first be converted to text by an OCR skill.

5
MCQhard

A logistics company uses an Azure AI Search indexer to process bills of lading stored in Azure Blob Storage. The indexer uses a skillset with a ShaperSkill that builds a complex object named 'shipment' containing nested fields for carrier, origin, and destination. After a full index run, queries for the carrier field return no results even though the source documents contain the data. You need to make the carrier value searchable. What should you do?

A.Change the ShaperSkill to output a string instead of an object.
B.Increase the indexer batch size so that all documents are processed in a single run.
C.Enable the indexer's cache and rerun the indexer.
D.Add an outputFieldMapping that maps the nested carrier value to a top-level index field.
AnswerD

Values produced inside a ShaperSkill object exist only in the enrichment tree. To persist them in the index, the indexer needs an outputFieldMapping that targets a field defined in the index. Mapping the nested carrier node to a searchable top-level field makes the value retrievable and queryable, resolving the empty query results.

Why this answer

The ShaperSkill constructs a complex object inside the enrichment tree, but that tree is transient. Persisting any part of it requires an outputFieldMapping from the enrichment node to a field in the index. Mapping the nested carrier node to a searchable top-level field is what actually makes the value appear in query results.

Exam trap

The trap here is believing that shaping an object automatically indexes its members, when shaped nodes must be explicitly mapped to index fields.

6
MCQmedium

A company uses Azure AI Search to index documents from an Azure SQL Database. They have configured an indexer with a skillset that includes a custom skill hosted in an Azure Function. The custom skill enriches each document with a 'category' field. After running the indexer, they notice that the 'category' field is missing in the index for all documents. The Azure Function logs show that it is receiving requests and returning responses. What is the most likely cause?

A.The indexer's data source connection string is invalid, so documents are not being retrieved.
B.The output field mapping for the custom skill is not correctly configured in the indexer.
C.The custom skill's context is set to /document/pages/*, but the documents do not have a 'pages' node.
D.The custom skill's Azure Function is returning a 200 OK status but with an empty response body.
AnswerB

This is correct because if the custom skill's output is not mapped to an index field, the enriched data will not appear in the index. Even though the skill runs successfully, the indexer needs an outputFieldMapping to send the skill's output to the target field. Without it, the category data is discarded, resulting in missing values in the index.

Why this answer

The most likely cause is a missing or incorrect output field mapping. The custom skill executes and returns a response, but if the indexer is not configured to map that output to an index field, the enriched data never gets stored. This is a common configuration oversight: the skill definition includes an output, but the indexer's outputFieldMappings must explicitly associate it with a target field.

Exam trap

The trap here is focusing on the skill's execution or function code, but the missing field is often due to a missing output field mapping in the indexer.

7
MCQeasy

You are creating an Azure AI Search index for a knowledge mining solution. The index must support searching for documents by a required category field that can have one of five predefined values, and you want to enable faceted navigation on that field. Which index field configuration should you use?

A.Edm.String with searchable set to true and filterable set to false.
B.Edm.String with retrievable set to false and filterable set to true.
C.Edm.String with filterable and facetable set to true, and searchable set to false.
D.Edm.Int32 with filterable set to true and facetable set to false.
AnswerC

For a field used for filtering and faceting on exact values, setting it as Edm.String with filterable and facetable true and searchable false is optimal. This configuration allows exact-match filtering and facet counts without incurring the overhead of full-text analysis, which is unnecessary for predefined category values.

Why this answer

A category field with predefined values is best modeled as a filterable and facetable string field that is not searchable, because full-text analysis is unnecessary and can interfere with exact matching. Making it searchable could tokenize values, and omitting facetable would break faceted navigation. The retrievable attribute should remain true if the value needs to be displayed.

Exam trap

The trap here is enabling searchable on a category field, which tokenizes values and breaks exact filtering and faceting.

8
MCQmedium

A law firm uses Azure Document Intelligence to extract clauses from legal contracts. They have a custom model trained on 15 labeled contracts. The model extracts clauses with high confidence on similar documents but fails to extract correct clauses from a new batch of contracts that have a different font and layout. The firm needs to improve extraction accuracy without retraining the model from scratch. The solution must minimize manual effort and cost. What should they do?

A.Use the prebuilt-layout model to extract clauses instead
B.Increase the OCR confidence threshold in the analysis request
C.Label 15 more contracts with the original layout and retrain the model
D.Create a composed model that includes the existing model and a new model trained on 5 contracts with the new layout
AnswerD

A composed model can handle multiple layouts by combining models.

Why this answer

Creating a composed model in Azure Document Intelligence allows you to combine the existing model (trained on the original layout) with a new model trained on just 5 labeled contracts from the new layout. This approach improves accuracy on the new layout without retraining from scratch, minimizing manual effort and cost by leveraging the composed model's ability to route documents to the appropriate sub-model based on layout similarity.

Exam trap

The trap here is that candidates often assume retraining with more data (Option C) is always the best solution, but they overlook the composed model feature which is specifically designed to handle layout variations with minimal additional labeling and cost.

How to eliminate wrong answers

Option A is wrong because the prebuilt-layout model is designed for extracting text and structure (like tables and selection marks), not for custom clause extraction from legal contracts, and it would not leverage the firm's existing labeled data. Option B is wrong because increasing the OCR confidence threshold only filters out low-confidence text recognition results; it does not improve the model's ability to correctly classify or extract clauses from a different font and layout. Option C is wrong because labeling 15 more contracts with the original layout and retraining the model would not address the new layout variation; it would only reinforce the existing model's performance on the original layout, wasting effort and cost.

9
MCQmedium

A company is building a knowledge mining solution using Azure AI Search. They need to extract key phrases from a large set of documents in multiple languages. Which skill should they add to the skillset?

A.Key Phrase Extraction skill
B.Sentiment Analysis skill
C.Language Detection skill
D.Entity Recognition skill
AnswerA

The Key Phrase Extraction skill uses natural language processing to identify salient terms and phrases, and it supports multiple languages within Azure AI Search skillsets. It satisfies the requirement to extract key phrases from a large multilingual document set.

Why this answer

The Key Phrase Extraction skill is the correct choice because it is specifically designed to identify and extract the most important phrases from text, which directly supports the requirement to extract key phrases from documents. Azure AI Search's built-in Key Phrase Extraction skill leverages natural language processing to analyze text and return a list of key phrases, making it the appropriate skill for this knowledge mining solution.

Exam trap

The trap here is that candidates may confuse Entity Recognition (which extracts single-word entities like 'Microsoft') with Key Phrase Extraction (which extracts multi-word phrases like 'Azure AI Search'), leading them to choose Option D instead of A.

How to eliminate wrong answers

Option B (Sentiment Analysis skill) is wrong because it evaluates the emotional tone or sentiment (positive, negative, neutral) of text, not the extraction of key phrases. Option C (Language Detection skill) is wrong because it identifies the language of the text but does not extract key phrases from the content. Option D (Entity Recognition skill) is wrong because it identifies and categorizes named entities (e.g., people, organizations, locations) rather than extracting multi-word key phrases that summarize the document's main topics.

10
Multi-Selectmedium

A company uses Azure Document Intelligence to process custom forms. They have trained a custom model using labeled data. They need to improve the model's accuracy for a specific field that is frequently misrecognized. The field appears in a consistent location but has varying formats. Which two actions should they take? (Choose two.)

Select 2 answers
A.Add more labeled samples that include variations of the field's format to the training dataset.
B.Use the prebuilt model for that field type instead of the custom model.
C.Use the model's confidence scores to identify low-confidence instances and correct their labels in the training set.
D.Retrain the model using the same dataset but with a different neural network architecture.
E.Increase the model's confidence threshold to reduce false positives for that field.
AnswersA, C

Adding more labeled samples with format variations helps the model learn to generalize and recognize the field regardless of format changes. The custom model learns from the labeled examples, so increasing diversity in the training data directly improves accuracy for that field. This is a fundamental step in improving model performance.

Why this answer

Improving a custom model's accuracy requires enhancing the training data. Adding labeled samples with format variations exposes the model to diverse representations, while correcting low-confidence instances ensures the training set is accurate. These actions directly address the model's learning process, unlike threshold adjustments or architectural changes.

Exam trap

The trap here is thinking that adjusting confidence thresholds or changing the model type can fix recognition errors, when the real solution lies in improving the training data.

11
Matchingmedium

Match each Azure AI tool to its purpose.

Drag a concept onto its matching description — or click a concept then click the description.

Concepts
Matches

Drag-and-drop ML model building

Interactive code development

Command-line management of Azure resources

Programmatic access to Azure services

Run AI services on-premises

Why these pairings

Correct matches: Azure Cognitive Search is an AI search service; Azure Bot Service is for building bots; Azure Cognitive Services offers pre-built AI APIs. Common confusions include swapping Cognitive Search with Machine Learning, and Bot Service with Cognitive Services.

12
Multi-Selectmedium

You are designing an Azure AI Search solution that uses an AI enrichment pipeline to extract text and key phrases from scanned documents. You need to ensure the pipeline can process image-heavy PDFs and produce searchable text and key phrases. Which two actions should you include in the skillset? (Choose two.)

Select 2 answers
A.Add the built-in Key Phrase Extraction skill to identify important phrases from the OCR text.
B.Add the built-in Language Detection skill to determine the language of each document.
C.Add the built-in Text Merge skill to combine multiple text inputs into a single field.
D.Add the built-in Entity Recognition skill to extract people, organizations, and locations.
E.Add the built-in OCR skill to extract text from images embedded in the PDFs.
AnswersA, E

The Key Phrase Extraction skill analyzes text and returns a list of key phrases. Applied after OCR, it operates on the extracted text to surface salient terms, which can be mapped to a collection field in the index. This directly fulfills the requirement to produce key phrases from the scanned documents.

Why this answer

To process image-heavy PDFs, the OCR skill extracts text from embedded images. Once text is available, the Key Phrase Extraction skill identifies important phrases. Together, these skills enable the pipeline to produce searchable text and key phrases.

Other skills like Language Detection, Text Merge, or Entity Recognition serve different purposes and do not fulfill both requirements.

Exam trap

The trap here is choosing skills that seem related to text analysis but do not actually extract text from images or generate key phrases, such as Language Detection or Entity Recognition.

13
MCQeasy

A company uses Azure Document Intelligence to analyze prebuilt invoices. They need to extract the invoice total and the due date from each invoice. They call the Analyze Invoice operation and receive a result. Which part of the JSON response contains the extracted fields?

A.The documents array, where each document has a fields object with named fields such as InvoiceTotal and DueDate.
B.The tables array, where each table has cells that contain the extracted values.
C.The pages array, where each page has a lines array containing text and bounding boxes.
D.The analyzeResult object, which contains a keyValuePairs array with keys and values.
AnswerA

The Analyze Invoice operation returns a result with a documents array. Each document contains a fields object where each field (e.g., InvoiceTotal, DueDate) includes its value, confidence, and bounding regions. This is the standard structure for prebuilt models, allowing you to access extracted fields by name.

Why this answer

The prebuilt invoice model returns a structured JSON with a documents array. Each document includes a fields object that contains named fields such as InvoiceTotal and DueDate, along with their values and confidence scores. This is the correct location to extract specific invoice data.

Exam trap

The trap here is assuming that all extracted data is in a generic key-value format, when the prebuilt invoice model uses a specialized fields object.

14
MCQhard

A company uses Azure Document Intelligence with a custom neural model to extract data from purchase orders. The model was trained on 50 labeled samples and performs well on similar documents. However, when processing new purchase orders from a different supplier, the model fails to extract the 'TotalAmount' field accurately. What should you do to improve the model's performance on the new supplier's documents?

A.Add labeled samples from the new supplier's purchase orders and retrain the model.
B.Increase the model's confidence threshold for the 'TotalAmount' field.
C.Use the model's composed model feature to combine it with a prebuilt model.
D.Switch to the prebuilt invoice model and map the 'TotalAmount' field to 'InvoiceTotal'.
AnswerA

Custom neural models learn from the training data. Adding labeled samples that represent the new supplier's document variations helps the model generalize to those layouts and field positions. Retraining incorporates this new data, improving accuracy for the 'TotalAmount' field on similar documents from that supplier.

Why this answer

To improve extraction for documents from a new supplier, the custom neural model needs exposure to that supplier's document variations. Adding labeled samples and retraining is the direct method to enhance the model's field extraction accuracy. Other options do not address the model's knowledge gap.

Exam trap

The trap here is thinking that adjusting confidence thresholds or switching to prebuilt models can fix extraction errors when the model lacks training data for the new layout.

15
Multi-Selectmedium

You are building an Azure AI Search solution that uses a custom skill to enrich documents. The custom skill is implemented as an Azure Function. You need to ensure that the custom skill can access the documents' content and output enriched data. Which two actions should you perform? (Choose two.)

Select 2 answers
A.Add the Azure Function's output to the index's field mappings.
B.Ensure the Azure Function accepts a JSON payload with the expected input and returns a JSON response.
C.Set the custom skill's batch size to 1 to ensure sequential processing.
D.Configure the indexer to use a managed identity to access the Azure Function.
E.Define the custom skill in the skillset with the correct input source and output mappings.
AnswersB, E

Azure AI Search sends a JSON payload containing the input values to the custom skill and expects a JSON response with the output values. The Azure Function must be implemented to parse the request and return the appropriate structure, otherwise the skill will fail or produce no output.

Why this answer

The custom skill must be defined in the skillset with proper input and output mappings to integrate with the enrichment pipeline. Additionally, the Azure Function must handle the JSON payload correctly. These two actions ensure the skill can access document content and return enriched data for indexing.

Exam trap

The trap here is focusing on security or performance settings like managed identity or batch size, which are not required for basic functionality of a custom skill.

16
MCQhard

A company uses Azure AI Search to index a large collection of scanned invoices stored in Azure Blob Storage. They have a skillset that includes an OCR skill to extract text from the invoices. The indexer is configured to run every night. They notice that the indexer takes a long time to complete and sometimes times out. They want to optimize the indexer performance without reducing the quality of the extracted text. What should they do?

A.Switch from the OCR skill to the Text Merge skill to combine text from multiple pages.
B.Enable incremental enrichment and cache the enriched data to avoid reprocessing unchanged documents.
C.Increase the indexer's batch size and reduce the number of parallel indexers.
D.Reduce the OCR skill's image resolution by setting the 'imageAction' to 'none' and relying on the document's embedded text.
AnswerB

This is correct because incremental enrichment with caching allows the indexer to skip documents that have not changed since the last run, reusing previously enriched data. This reduces the amount of OCR processing required, significantly improving performance for subsequent runs. It maintains text quality because unchanged documents are not reprocessed, and only new or modified documents are enriched.

Why this answer

Enabling incremental enrichment and caching is the best approach to optimize indexer performance for recurring runs. It avoids reprocessing documents that have not changed, thus reducing the workload on the OCR skill and other enrichments. This maintains the quality of extracted text because only new or modified documents are processed, while unchanged documents reuse cached enrichments.

Other options either reduce quality or do not address the performance bottleneck.

Exam trap

The trap here is thinking that reducing image resolution or disabling OCR will speed up the indexer, but that sacrifices the required text extraction quality.

17
MCQeasy

A healthcare organization uses Azure Document Intelligence to process patient intake forms. They notice that the confidence scores for field extraction are low. What is the most likely cause?

A.The document resolution is too low
B.The document layout is not analyzed
C.The custom model was trained with only 10 labeled forms
D.The batch processing size is too large
AnswerC

Custom models require at least 5 labeled forms; more samples improve confidence.

Why this answer

Custom models in Azure Document Intelligence require a minimum of five labeled forms for training, but low confidence scores typically indicate insufficient training data. With only 10 labeled forms, the model lacks enough examples to generalize well across variations in handwriting, formatting, and field values, leading to poor extraction confidence.

Exam trap

The trap here is that candidates often confuse low confidence with OCR or resolution issues, but the exam tests the specific requirement for sufficient labeled training data in custom models, not generic document quality problems.

How to eliminate wrong answers

Option A is wrong because low resolution can reduce OCR accuracy, but Azure Document Intelligence handles a wide range of resolutions and the question specifically points to field extraction confidence, not OCR failure. Option B is wrong because layout analysis is automatically performed by the prebuilt layout model and is not a prerequisite for custom extraction models; the issue is with training data quantity, not layout processing. Option D is wrong because batch processing size affects throughput and latency, not the confidence scores of individual field extractions; confidence is determined by the model's training and the input document quality, not batch size.

18
MCQeasy

A company is building a knowledge mining solution using Azure AI Search. They need to extract text from handwritten notes stored as images in Azure Blob Storage. They want to use the built-in OCR skill in a skillset. Which cognitive service does the OCR skill rely on?

A.Azure AI Language
B.Azure AI Document Intelligence
C.Azure AI Translator
D.Azure AI Vision
AnswerD

This is correct because the OCR skill in Azure AI Search uses the Computer Vision OCR API from Azure AI Vision. It extracts text from images, including handwritten text, depending on the language and model. The skill sends image data to the Azure AI Vision service and receives recognized text, which is then enriched into the search index.

Why this answer

The built-in OCR skill in Azure AI Search is powered by the Computer Vision OCR API, which is part of Azure AI Vision. This service extracts text from images, including handwritten text, and returns it as structured data. The skill then makes this text available for indexing.

Other cognitive services like Language, Document Intelligence, and Translator serve different purposes and are not used by the OCR skill.

Exam trap

The trap here is confusing Azure AI Document Intelligence with the OCR skill, but Document Intelligence is a separate service not used by the built-in OCR skill.

19
Multi-Selecthard

You are building a knowledge mining solution with Azure AI Search. The solution indexes scanned PDF reports stored in Azure Blob Storage. Each report contains multiple embedded images with text that must be searchable. You have already created a data source, index, and indexer. You need to ensure that text from the embedded images is extracted and mapped to the 'content' field in the index. Which two actions should you perform? (Choose two.)

Select 2 answers
A.Enable image extraction by setting the 'imageAction' parameter to 'generateNormalizedImages' in the indexer definition.
B.Configure the indexer to use the 'text' parsing mode instead of the default 'json' mode.
C.Add an OCR skill to the skillset and set its context to /document/normalized_images/*.
D.Add an Entity Recognition skill to detect and extract text from images.
E.Add a Shaper skill that merges all image text into a single string before indexing.
AnswersA, C

This is correct because imageAction must be set to generateNormalizedImages to create normalized images from embedded images in the document. Without this, the OCR skill has no image inputs to process. The normalized images become available at /document/normalized_images/*, which is the required input path for the OCR skill.

Why this answer

To make text from embedded images searchable, you must first extract the images from the documents by setting imageAction to generateNormalizedImages. Then, you add an OCR skill with context /document/normalized_images/* to process each image and output the recognized text. This text can then be mapped to the content field.

Without both actions, the OCR skill either has no input or does not run per image.

Exam trap

The trap here is assuming that adding an OCR skill alone is sufficient, but without image extraction the OCR skill has no normalized images to process.

20
MCQmedium

A company builds a knowledge mining solution using Azure AI Search with a custom skillset that includes an OCR skill. They want to ensure that images embedded in PDFs are processed. What should they configure?

A.Set the 'defaultLanguageCode' to 'en'
B.Set the 'textExtractionAlgorithm' to 'printed'
C.Set the 'imageAction' parameter to 'generateNormalizedImages'
D.Set the 'lineEnding' parameter to 'space'
AnswerC

This parameter enables extraction of images from documents.

Why this answer

The 'imageAction' parameter in Azure AI Search's OCR skill controls whether images embedded in documents (including PDFs) are extracted and processed. Setting it to 'generateNormalizedImages' ensures that images within PDFs are normalized and passed to the OCR skill for text extraction, which is essential for processing embedded images.

Exam trap

The trap here is that candidates may confuse parameters that affect OCR output formatting (like 'lineEnding' or 'defaultLanguageCode') with the parameter that actually enables image extraction from PDFs, leading them to overlook the 'imageAction' setting.

How to eliminate wrong answers

Option A is wrong because 'defaultLanguageCode' specifies the language for text recognition, not whether images are extracted from PDFs; it does not enable image processing. Option B is wrong because 'textExtractionAlgorithm' determines the OCR algorithm (e.g., 'printed' or 'handwritten') but does not control the extraction of images from PDFs; it only affects how text is recognized once images are available. Option D is wrong because 'lineEnding' parameter controls the line break character in OCR output (e.g., 'space', 'carriageReturn'), which is irrelevant to enabling image extraction from PDFs.

21
MCQmedium

You are building an Azure AI Search knowledge mining pipeline that enriches scanned PDF invoices. The PDFs are stored in Azure Blob Storage, and you need to extract text from each page before running downstream entity recognition. The solution must minimize development effort and rely on a built-in cognitive skill. Which skill should you add to the skillset?

A.OcrSkill
B.EntityRecognitionSkill
C.ImageAnalysisSkill
D.KeyPhraseExtractionSkill
AnswerA

OcrSkill is the built-in cognitive skill that extracts text from image files and embedded images in PDFs, producing a text output that downstream skills can consume. In this scenario, the scanned invoices contain image-based text, so OCR is required before entity recognition. It minimizes development effort because no custom code or external endpoint is needed.

Why this answer

The pipeline must convert image-based invoice content into text before any language skill can run. OcrSkill is the built-in Azure AI Search cognitive skill designed for that conversion and outputs a text field that downstream skills consume. Image analysis, entity recognition, and key phrase extraction all assume text already exists, so they cannot satisfy the extraction requirement on their own.

Exam trap

The trap here is assuming any cognitive skill that produces text can read scanned images, when only the OCR skill performs image-to-text extraction.

22
Multi-Selectmedium

A media company is building a knowledge mining solution with Azure AI Search. They need to enrich video assets by extracting spoken words from the audio track and then indexing that transcript for search. Which two components must be included in the enrichment pipeline? (Choose two.)

Select 2 answers
A.A custom skill that calls Azure AI Video Indexer or the Speech service to transcribe the audio.
B.A scoring profile that boosts video documents by duration.
C.An indexer configured for Azure Blob Storage with the video files in a container.
D.A synonym map that maps spoken words to their text equivalents.
E.A skillset entry that calls the built-in OCR skill on the video file.
AnswersA, C

Azure AI Search does not include a built-in skill that transcribes video audio, so a custom skill is required to invoke a speech-to-text service such as Azure AI Video Indexer or the Speech service. The custom skill returns the transcript as enriched text that can then be mapped into an index field for search.

Why this answer

To index spoken words from video, the pipeline needs a data source and indexer to pull video assets from Blob Storage, plus a custom skill that calls a speech-to-text service because Azure AI Search has no built-in audio transcription skill. OCR, synonym maps, and scoring profiles operate on images, queries, or ranking respectively and cannot produce a transcript.

Exam trap

The trap here is assuming Azure AI Search has a built-in skill for audio transcription, when video or speech transcription must be handled by a custom skill.

23
MCQhard

A company uses Azure AI Search to index documents from an Azure SQL Database. They need to ensure that deleted rows in the database are also removed from the search index during incremental indexing. They have configured the data source with change detection policies. What should they do to enable deletion detection?

A.Configure the indexer to use the high water mark change detection policy and set the deletion detection policy to 'none'.
B.Create a SQL trigger that calls the Azure AI Search REST API to delete the document when a row is deleted.
C.Use a view that filters out deleted rows and configure the indexer to use that view as the data source.
D.Add a soft delete column to the table that indicates when a row is deleted, and configure the data source with a soft deletion policy that references that column.
AnswerD

Azure AI Search supports soft delete detection for Azure SQL Database. You must add a column (e.g., IsDeleted) to the table and set its value to true or a timestamp when a row is deleted. Then, in the data source definition, configure a soft deletion policy that specifies this column. The indexer will then remove corresponding documents from the index during incremental runs.

Why this answer

To enable deletion detection for Azure SQL Database in Azure AI Search, you must implement a soft delete column in the table and configure the data source with a soft deletion policy that references that column. The indexer then uses this column to identify and remove deleted documents from the index during incremental indexing.

Exam trap

The trap here is assuming that change detection policies automatically handle deletions, when in fact a separate soft delete policy is required.

24
Multi-Selectmedium

You are building an Azure AI Search knowledge mining solution over a repository of scanned product manuals. You need to extract structured entities such as product names and part numbers from the OCR text and store them in an index field. (Choose two.)

Select 2 answers
A.Add Microsoft.Skills.Text.LanguageDetectionSkill to the skillset.
B.Define an outputFieldMapping that writes the entity recognition output to an index field.
C.Add Microsoft.Skills.Text.KeyPhraseExtractionSkill to the skillset.
D.Configure the indexer to use a JSON parsing mode.
E.Add Microsoft.Skills.Text.EntityRecognitionSkill to the skillset and set its categories to include the needed entity types.
AnswersB, E

Skill outputs live only in the enrichment tree until they are mapped. An outputFieldMapping connects the entity recognition output node to a field defined in the index. Without this mapping the extracted entities would be discarded after enrichment, so this step is required to make the data queryable.

Why this answer

Extracting structured entities from OCR text requires a skill that performs entity recognition, and the results must be persisted through an output field mapping. The entity recognition skill with configured categories identifies the relevant named entities, while the field mapping writes those values into the index so they can be queried. Skills that only detect language or key phrases do not produce the required structured output.

Exam trap

The trap here is assuming that enriching text with any language skill automatically stores results in the index, when a field mapping is also required.

25
MCQmedium

You are developing an Azure AI Search solution that indexes scanned PDF invoices. The indexer must extract text from the PDFs and also recognize entities such as organization names and dates. You want to use built-in cognitive skills to minimize custom code. Which combination of skills should you include in the skillset?

A.Key Phrase Extraction skill and Language Detection skill
B.OCR skill and Entity Recognition skill
C.Text Merge skill and Text Split skill
D.Image Analysis skill and Sentiment skill
AnswerB

The OCR skill extracts text from image-based PDF content, and the Entity Recognition skill identifies entities like organizations and dates from that text. This combination directly addresses the requirement to extract text and recognize entities without writing custom code, leveraging built-in cognitive skills in Azure AI Search.

Why this answer

The OCR skill extracts text from scanned PDFs, and the Entity Recognition skill identifies organizations and dates from the extracted text. Using these built-in skills avoids custom code and satisfies both requirements efficiently. Other combinations either lack OCR or entity recognition, or provide unrelated functionality.

Exam trap

The trap here is assuming that Image Analysis performs OCR, when in Azure AI Search the dedicated OCR skill is needed for text extraction from images.

26
MCQeasy

A retail company wants to build a knowledge mining solution that indexes product descriptions stored in an Azure SQL Database and makes them searchable through a web application. The descriptions are already plain text. You need to configure Azure AI Search to pull the data into an index with the least effort. What should you create first?

A.A cognitive skillset that applies OCR to each product description.
B.A custom skill that queries the SQL database and writes documents to the index.
C.An indexer with a data source connection to Azure SQL Database.
D.An Azure AI Document Intelligence model trained on the product descriptions.
AnswerC

An indexer connects to a supported data source, reads documents, and populates an index. Creating the data source connection to Azure SQL Database and an indexer that targets the index is the standard low-effort way to ingest plain text records. No enrichment skills are needed because the content is already machine-readable.

Why this answer

Azure AI Search integrates directly with Azure SQL Database through indexers and data source connections. When the source content is already plain text, the indexer can read rows and populate the index without any enrichment skills. This is the least-effort approach and avoids custom code or document extraction models.

Exam trap

The trap here is reaching for enrichment skills or document intelligence when the source data is already structured text that a built-in indexer can ingest directly.

Ready to test yourself?

Try a timed practice session using only Implement knowledge mining and document intelligence solutions questions.

CCNA Implement knowledge mining and document intelligence solutions Questions | Courseiva