Courseiva

CCNA Implement computer vision solutions Questions

75 of 90 questions · Page 1/2 · Implement computer vision solutions · Answers revealed

1
MCQeasy

A developer needs to generate a descriptive caption for an image using Azure AI Vision. The image contains a dog catching a frisbee in a park. Which feature of Image Analysis should they use?

A.Caption
B.Read
C.Objects
D.Tags
AnswerA

The caption feature in Image Analysis generates a human-readable sentence that describes the image content, such as 'a dog catching a frisbee in a park'. It is specifically designed to produce descriptive captions. This directly meets the developer's requirement for a descriptive caption.

Why this answer

The caption feature of Azure AI Vision Image Analysis generates a natural language description of an image's content. It is the only feature that produces a sentence like 'a dog catching a frisbee in a park'. Tags, objects, and read serve different purposes: tags provide keywords, objects detect and locate items, and read extracts text.

Therefore, caption is the correct choice.

Exam trap

The trap here is confusing tags with captions, as both provide descriptive information, but only captions produce a full sentence.

2
MCQhard

A developer is using the Azure AI Vision Image Analysis API to extract text from a photo of a street sign. The sign contains text in both English and Japanese arranged in multiple columns. The developer needs the response to include the detected language and the bounding box for each text line. Which feature should they use?

A.The 'read' feature in Image Analysis
B.The 'detectObjects' feature in Image Analysis
C.The 'caption' feature in Image Analysis
D.The 'tags' feature in Image Analysis
AnswerA

The read feature in Image Analysis performs OCR and returns extracted text along with bounding boxes for lines and words. It also detects the language of the text. This directly provides the language and bounding box information required, and it handles mixed-language and multi-column text in a single call.

Why this answer

The read feature in Azure AI Vision Image Analysis is specifically designed for OCR. It returns extracted text with bounding boxes for lines and words, and it detects the language of the text. This matches the need to capture English and Japanese text in multiple columns and provide bounding boxes for each line.

Other features like caption, detectObjects, and tags do not perform OCR.

Exam trap

The trap here is confusing general image analysis features like captioning or tagging with OCR, when only the read feature extracts text and provides bounding boxes and language detection.

3
MCQhard

You are a data scientist at a healthcare startup. You have deployed a custom object detection model using Azure Custom Vision to detect tumors in MRI scans. The model was trained on 10,000 labeled scans from a single hospital. After deployment, the model performs well on scans from that hospital but poorly on scans from a different hospital with a different MRI machine. The new hospital's scans have slightly different contrast and resolution. The model's precision drops from 0.92 to 0.65, and recall drops from 0.88 to 0.50. You have access to 500 labeled scans from the new hospital. You need to improve the model's performance on the new hospital's data as quickly as possible with minimal effort. What should you do?

A.Collect more labeled scans from the new hospital and train a new model from scratch.
B.Create a new Custom Vision project and train only on the 500 new scans.
C.Apply image preprocessing to normalize the new hospital's scans to match the old hospital's style, then use the existing model.
D.Use the existing model as a starting point and retrain it with the 500 labeled scans from the new hospital.
AnswerD

Domain shift from differing contrast and resolution causes the precision and recall drop. Retraining the existing Custom Vision model with the 500 labelled new-hospital scans performs transfer learning, adapting features to the new distribution quickly and with minimal effort.

Why this answer

Azure Custom Vision supports transfer learning, allowing you to take an existing trained model and retrain it with new labeled data. By using the 500 labeled scans from the new hospital as a training set, you can fine-tune the model to adapt to the different contrast and resolution characteristics without starting from scratch. This approach is the fastest and requires minimal effort, leveraging the previously learned features while incorporating domain-specific adjustments.

Exam trap

The trap here is that candidates may overestimate the need for large datasets or manual preprocessing, failing to recognize that Azure Custom Vision's built-in transfer learning is designed to efficiently adapt models with minimal new data.

How to eliminate wrong answers

Option A is wrong because collecting more labeled scans and training a new model from scratch is time-consuming and resource-intensive, not the quickest or minimal-effort solution. Option B is wrong because creating a new Custom Vision project and training only on 500 scans ignores the valuable knowledge from the original 10,000 scans, leading to a model with insufficient data and likely poor generalization. Option C is wrong because applying image preprocessing to normalize the new hospital's scans to match the old hospital's style is a manual, error-prone process that may not fully address the underlying domain shift and does not leverage the labeled data for supervised adaptation.

4
MCQmedium

A security firm wants to analyze live video from cameras at a warehouse gate to count the number of people entering and leaving. They require a ready-to-use service that provides a real-time count and does not require training a custom model. Which Azure AI service should they use?

A.Azure AI Custom Vision
B.Azure AI Face
C.Azure AI Vision Spatial Analysis
D.Azure Video Indexer
AnswerC

Azure AI Vision Spatial Analysis is designed for real-time video analytics, including people counting, and does not require custom model training. It ingests RTSP streams and provides counts of people crossing a defined line or zone. This matches the requirement for a ready-to-use service that counts people entering and leaving.

Why this answer

The security firm needs a ready-to-use service for real-time people counting from live video. Azure AI Vision Spatial Analysis is purpose-built for this scenario, offering pre-trained models that analyze RTSP streams and count people crossing lines or zones without custom training. Other services either focus on static images or require significant custom development, making them unsuitable for immediate deployment.

Exam trap

The trap here is assuming that any Azure AI vision service can process live video streams for people counting, when only Spatial Analysis provides that out-of-the-box capability.

5
MCQhard

You are designing a solution that uses Azure AI Vision to extract text from scanned invoices. The invoices vary in layout and include both printed and handwritten fields. The solution must achieve high accuracy with minimal manual labeling. Which approach should you recommend?

A.Use the Read API to extract all text and then use a custom regex to parse fields.
B.Use the Azure AI Document Intelligence prebuilt invoice model.
C.Train a Custom Vision object detection model to locate fields.
D.Label hundreds of invoices and train a custom Azure AI Document Intelligence model.
AnswerB

The prebuilt invoice model in Azure AI Document Intelligence is trained on varied invoice layouts and extracts printed and handwritten fields with high accuracy, requiring no custom labelling. This satisfies the minimal manual labelling constraint for the scanned invoices.

Why this answer

The Azure AI Document Intelligence prebuilt invoice model is specifically designed to extract common fields from invoices with high accuracy, handling both printed and handwritten text without requiring manual labeling. It uses advanced OCR and deep learning models trained on thousands of invoices, making it ideal for varied layouts and minimal setup.

Exam trap

The trap here is that candidates often overestimate the need for custom training or regex-based parsing, not realizing that Azure provides a prebuilt, high-accuracy model specifically for invoices that requires zero manual labeling.

How to eliminate wrong answers

Option A is wrong because the Read API only extracts raw text without understanding document structure or field semantics, requiring complex and brittle regex patterns that fail with varied invoice layouts. Option C is wrong because Custom Vision object detection is designed for image classification and object localization, not for extracting structured text fields from documents. Option D is wrong because labeling hundreds of invoices to train a custom model is unnecessary and inefficient when a prebuilt model already exists for this specific use case, violating the requirement for minimal manual labeling.

6
Multi-Selecthard

Which THREE factors should you consider when choosing between Azure AI Custom Vision and Azure AI Vision pre-built models for an image classification task? (Choose three.)

Select 3 answers
A.Availability of labeled training data specific to the domain
B.Image format support (JPEG, PNG)
C.Whether the required labels are covered by the pre-built model
D.Need for real-time inference latency
E.Need to iterate and retrain the model over time
AnswersA, C, E

Custom Vision requires labeled data; pre-built models do not.

Why this answer

Azure AI Custom Vision is specifically designed for scenarios where you have domain-specific labeled training data that is not covered by pre-built models. Custom Vision allows you to upload your own labeled images and train a custom model tailored to your unique classification needs, which is essential when off-the-shelf models fail to recognize your target classes.

Exam trap

The trap here is that candidates often confuse image format support or latency as key differentiators, when in fact both services handle these similarly, and the core distinction lies in the availability of custom labeled data and the need for iterative retraining.

7
MCQhard

You are developing a solution that uses Azure AI Video Indexer to analyze surveillance videos for suspicious activity. The solution must generate alerts when a person is detected in a restricted area. Which feature should you use?

A.Azure AI Video Indexer sentiment analysis
B.Azure AI Video Indexer people detection and tracking
C.Azure AI Face identify API
D.Object detection in Azure AI Vision
AnswerB

People detection and tracking identifies and follows individuals across video frames, enabling detection of a person entering a defined restricted area. This satisfies the requirement to raise alerts on unauthorised presence, unlike face or emotion features.

Why this answer

Azure AI Video Indexer's people detection and tracking feature is specifically designed to detect and track individuals across video frames, making it ideal for generating alerts when a person enters a restricted area. This feature provides bounding boxes, timestamps, and tracking IDs that enable real-time monitoring and alerting based on spatial rules.

Exam trap

The trap here is that candidates confuse generic object detection (which can detect people as objects) with the dedicated people tracking capability that provides persistent IDs and temporal continuity needed for restricted-area alerts.

How to eliminate wrong answers

Option A is wrong because sentiment analysis in Azure AI Video Indexer detects emotional tone (e.g., happiness, sadness) from audio or text, not physical presence or location of people. Option C is wrong because the Azure AI Face Identify API matches detected faces against a known person group (identification), but does not track movement or detect entry into restricted zones. Option D is wrong because object detection in Azure AI Vision identifies generic objects (e.g., car, dog) and does not specialize in people tracking or spatial alerting within video streams.

8
MCQmedium

A retail company uses Azure Computer Vision to analyze customer traffic in stores. They process images from security cameras using the OCR API to detect product labels. Recently, the OCR accuracy has decreased for images with poor lighting. Which pre-processing step should the company implement to improve OCR accuracy?

A.Convert images to grayscale before sending to OCR API.
B.Increase the image resolution before calling OCR API.
C.Adjust brightness and contrast of images using image processing.
D.Reduce image size to decrease noise.
AnswerC

Adjusting brightness and contrast directly compensates for the poor lighting that is degrading OCR accuracy. Microsoft Entra ID is irrelevant here; the constraint is image quality, not authentication. Normalising luminance and contrast before sending images to the OCR API improves character segmentation and recognition, satisfying the stem's requirement to raise accuracy for poorly lit camera feeds.

Why this answer

Poor lighting directly reduces the contrast between text and background, which is critical for OCR accuracy. Adjusting brightness and contrast improves the signal-to-noise ratio of the text regions, making character edges more distinct for the Azure Computer Vision OCR engine. This pre-processing step compensates for the lighting deficiency without altering the fundamental image content that the API relies on.

Exam trap

The trap here is that candidates confuse image quality improvements (like resolution or noise reduction) with the specific need to correct lighting-induced contrast loss, which is a distinct pre-processing requirement for OCR in poor illumination.

How to eliminate wrong answers

Option A is wrong because converting to grayscale removes color information that can help distinguish text from similarly-toned backgrounds, and it does not address the root cause of low contrast due to poor lighting. Option B is wrong because increasing resolution does not fix the underlying contrast problem; it may even amplify noise and increase API processing time without improving text legibility. Option D is wrong because reducing image size discards pixel detail, which can make small or thin text characters unreadable for OCR, and it does not mitigate the effects of poor lighting.

9
MCQeasy

You are building a solution to detect if a person is wearing a hard hat in construction site images. You have a small dataset of labeled images. Which Azure service should you use?

A.Azure AI Vision Image Analysis
B.Azure Video Indexer
C.Azure AI Document Intelligence
D.Azure Custom Vision
AnswerD

Custom Vision trains a classification or object detection model on your own labelled images, which suits a small domain-specific dataset of hard-hat photos. The prebuilt Computer Vision service cannot reliably detect site-specific safety equipment without custom training.

Why this answer

Azure Custom Vision is the correct choice because it allows you to train a custom image classification model with your own small dataset of labeled construction site images to detect whether a person is wearing a hard hat. Unlike pre-built services, Custom Vision specializes in fine-tuning models for specific visual concepts that are not covered by general-purpose APIs, making it ideal for niche object detection tasks like hard hat detection.

Exam trap

The trap here is that candidates assume Azure AI Vision Image Analysis can handle any visual detection task because of its broad 'Image Analysis' name, but it cannot be customized for niche objects like hard hats, which requires a custom training service like Custom Vision.

How to eliminate wrong answers

Option A is wrong because Azure AI Vision Image Analysis provides pre-trained models for general image analysis (e.g., objects, tags, celebrities) but cannot be retrained on custom datasets like hard hat detection; it lacks the capability to learn new, specific classes from your labeled images. Option B is wrong because Azure Video Indexer is designed for analyzing video content (e.g., extracting insights, speech, faces) and is not suited for static image classification or custom object detection with a small dataset of images. Option C is wrong because Azure AI Document Intelligence is purpose-built for extracting text, tables, and key-value pairs from documents (e.g., invoices, forms) and has no capability for visual object detection or custom image classification.

10
MCQmedium

You are building a web app that allows users to upload images and receive a list of content tags, such as 'outdoor', 'tree', and 'person'. You want to use a prebuilt Azure AI service and minimize custom development. Which service should you use?

A.Azure AI Vision, using the Image Analysis API with the Tags visual feature
B.Azure AI Document Intelligence, using the prebuilt-read model
C.Azure AI Custom Vision, using a classification model
D.Azure AI Language, using key phrase extraction
AnswerA

The Image Analysis API provides a Tags feature that returns a list of content tags relevant to the image, such as 'outdoor', 'tree', or 'person'. It is prebuilt and requires no custom training, making it ideal for this scenario. You can call the API with the Tags feature enabled to get the desired output.

Why this answer

Azure AI Vision's Image Analysis API includes a Tags feature that returns content tags for an image without any custom training. This prebuilt capability directly satisfies the need for a list of tags such as 'outdoor', 'tree', and 'person'. Other services either require custom training, operate on text, or are document-focused.

Exam trap

The trap here is thinking that Custom Vision is needed for any tagging scenario, but prebuilt Vision tagging covers general concepts without training.

11
MCQhard

You work for a manufacturing company that uses Azure AI services to automate quality inspection on a production line. You have a Custom Vision object detection model that identifies defects on metal parts. The model was trained on images captured under ideal lighting conditions. However, when deployed in the factory, the model's accuracy drops significantly due to inconsistent lighting and glare. You need to improve the model's robustness without collecting new images from the factory floor. What should you do?

A.Increase the number of training iterations to force the model to learn more features.
B.Apply data augmentation techniques such as brightness, contrast, and blur adjustments to the existing training images.
C.Use higher resolution images for training.
D.Change the model type from object detection to classification.
AnswerB

Augmentation synthesises lighting variation—brightness, contrast and blur—directly from the existing dataset, so the detector learns glare-invariant features without any new factory-floor capture. This satisfies the stem's constraint of improving robustness under inconsistent lighting while collecting no additional images.

Why this answer

Using data augmentation techniques like brightness and contrast adjustments, rotation, and noise injection can simulate various lighting conditions and improve robustness. Option A is wrong because increasing training iterations may overfit to the existing data. Option C is wrong because higher resolution does not address lighting variation.

Option D is wrong because changing the model type does not address the data issue.

12
MCQmedium

You are developing an application that processes images of handwritten forms. The forms contain checkboxes that may be checked or unchecked. Which Azure AI service should you use to detect the state of the checkboxes?

A.Azure AI Custom Vision
B.Azure AI Language
C.Azure AI Document Intelligence
D.Azure AI Computer Vision
AnswerC

Azure AI Document Intelligence includes a prebuilt read and custom form models that detect checkbox selection state via the selectionMark field, returning checked or unchecked. This directly satisfies the requirement to determine checkbox state on handwritten forms.

Why this answer

Azure AI Document Intelligence (formerly Form Recognizer) is the correct service because it is specifically designed to extract structured data from documents, including detecting the state of checkboxes (checked or unchecked) in forms. It uses prebuilt models like the 'prebuilt-document' or custom extraction models to analyze form fields and checkbox selections, making it the optimal choice for this task.

Exam trap

The trap here is that candidates often confuse Azure AI Computer Vision's OCR capabilities with Document Intelligence's form-specific extraction, leading them to choose Computer Vision even though it cannot reliably detect checkbox states without additional custom logic.

How to eliminate wrong answers

Option A is wrong because Azure AI Custom Vision is used for training custom image classification and object detection models, not for extracting structured data like checkbox states from forms. Option B is wrong because Azure AI Language focuses on natural language processing tasks such as sentiment analysis, key phrase extraction, and entity recognition, not on visual document analysis or checkbox detection. Option D is wrong because Azure AI Computer Vision provides general image analysis capabilities like OCR and object detection, but it lacks the specialized form understanding and field extraction features needed to reliably detect checkbox states in structured documents.

13
MCQeasy

A university is developing an app for students to take photos of handwritten notes and convert them to digital text. The app must support multiple languages including English and Spanish. The solution should use a pre-built AI service. Which Azure service should you use?

A.Azure AI Document Intelligence with a custom model
B.Azure AI Vision Read API (OCR)
C.Azure AI Language with custom entity recognition
D.Custom Vision with a custom handwriting recognition model
AnswerB

Azure AI Vision's Read API performs optical character recognition on images, extracting handwritten and printed text. It supports multiple languages including English and Spanish, and is a pre-built service, meeting the university's requirements without custom model training.

Why this answer

Azure AI Vision's Read API (OCR) is a pre-built service specifically designed to extract printed and handwritten text from images, and it supports multiple languages including English and Spanish out of the box. It requires no training and is the correct choice for a general-purpose handwriting-to-text scenario.

Exam trap

AI-102 often tests the confusion between OCR services (Vision Read API) and document-understanding services (Document Intelligence), tricking candidates into picking Document Intelligence for simple handwriting OCR.

How to eliminate wrong answers

Option A is wrong because Azure AI Document Intelligence with a custom model requires labeled training data and is intended for structured document extraction (forms, invoices), not general handwriting OCR. Option B is correct. Option C is wrong because Azure AI Language with custom entity recognition extracts entities from text, not text from images — it does not perform OCR.

Option D is wrong because Custom Vision is an image classification/object detection service, not a handwriting recognition service, and training a custom model for handwriting is unnecessary when the Read API already handles it.

14
Multi-Selectmedium

Which TWO actions should you take to reduce the latency of an Azure AI Computer Vision OCR call on a large image?

Select 2 answers
A.Use a CPU-bound compute instance.
B.Resize the image to a smaller resolution before calling the API.
C.Increase the API timeout value.
D.Use the Read API asynchronously.
E.Deploy the Cognitive Services container on-premises.
AnswersB, D

Smaller images process faster.

Why this answer

Options B and D are correct. B: Resizing reduces processing time. D: Using the Read API asynchronously allows the client to poll, avoiding timeout.

A: Increasing timeout doesn't reduce latency. C: Using CPU doesn't help. E: Cognitive Services container on-premises might add network latency.

15
Multi-Selectmedium

Which THREE are valid uses of Azure AI Vision Image Analysis 4.0? (Select three.)

Select 3 answers
A.Extract printed text from an image using OCR
B.Transcribe spoken audio from a video file
C.Detect objects in an image and return bounding boxes
D.Generate a human-readable caption for an image
E.Translate text found in an image to another language
AnswersA, C, D

Image Analysis 4.0 includes a Read OCR feature that extracts printed and handwritten text from images via the Read API, satisfying the OCR use case. It returns text lines and words with bounding polygons, so printed text extraction is a supported capability of this service.

Why this answer

Option A is correct because Image Analysis 4.0 includes a Read OCR feature that extracts printed and handwritten text from images and documents, returning lines and words. Option C is correct because the object detection capability identifies objects in an image and returns their coordinates as bounding boxes along with confidence scores. Option D is correct because Image Analysis 4.0 can generate captions (and dense captions) describing the content of an image in natural language.

Option B is not valid because audio transcription is a speech service capability (e.g., Azure AI Speech), not Image Analysis. Option E is not valid because translating text is performed by Azure AI Translator; Image Analysis only extracts the text via OCR, it does not translate it.

Exam trap

The trap here is that candidates often confuse the capabilities of Azure AI Vision with those of Azure AI Speech or Azure AI Translator, assuming Image Analysis can handle audio transcription or text translation when it strictly processes visual content only.

16
MCQmedium

You are designing a solution that reads handwritten notes from patient intake forms. The solution must handle various handwriting styles. Which Azure AI capability should you use?

A.Azure AI Document Intelligence Read model
B.Azure AI Custom Vision
C.Azure AI Vision OCR
D.Azure AI Language
AnswerA

The Read model handles handwriting and printed text.

Why this answer

Azure AI Document Intelligence Read model is specifically designed to extract printed and handwritten text from documents, including patient intake forms. It uses advanced OCR capabilities optimized for varied handwriting styles and document layouts, making it the correct choice for this scenario.

Exam trap

The trap here is that candidates often confuse Azure AI Vision OCR (which is for printed text) with the Document Intelligence Read model (which is specialized for handwriting and document structure), leading them to choose the wrong service for handwriting recognition tasks.

How to eliminate wrong answers

Option B is wrong because Azure AI Custom Vision is used for image classification and object detection, not for extracting text from documents or handwriting. Option C is wrong because Azure AI Vision OCR is a general-purpose OCR that works well for printed text but is not optimized for handwriting recognition. Option D is wrong because Azure AI Language is focused on natural language processing tasks like sentiment analysis and entity recognition, not on extracting text from images or documents.

17
Multi-Selecteasy

Which TWO Azure AI services can be used to extract text from images?

Select 2 answers
A.Azure AI Face
B.Azure AI Document Intelligence
C.Azure AI Video Indexer
D.Azure AI Computer Vision
E.Azure AI Custom Vision
AnswersB, D

Azure AI Document Intelligence performs OCR plus layout analysis, extracting text from images embedded in documents and returning structured content. This satisfies the requirement to extract text from images, particularly where layout and form structure matter.

Why this answer

Azure AI Document Intelligence (formerly Form Recognizer) includes the Read OCR engine that extracts printed and handwritten text from images and documents. Azure AI Computer Vision provides the OCR API (optical character recognition) which can extract text from images, including both printed and handwritten text, and supports multiple languages. Both services are designed specifically for text extraction from visual content.

Exam trap

The trap here is that candidates may confuse Azure AI Video Indexer's ability to extract text from video frames as a primary image text extraction service, but it is designed for video analysis and indexing, not standalone image text extraction.

18
MCQhard

You are deploying a computer vision model using Azure AI Custom Vision with a small dataset of 200 images per class. The model shows high accuracy on training data but low accuracy on test data. Which action should you take to reduce overfitting?

A.Increase the learning rate
B.Reduce the image size to lower resolution
C.Increase the number of training epochs
D.Increase the training dataset size with more varied images
AnswerD

Overfitting arises when the model memorises limited training examples, so adding more varied images increases feature diversity and improves generalisation to unseen test data. This directly addresses the small dataset constraint of 200 images per class named in the stem.

Why this answer

Increase the training dataset size with more varied images. Overfitting occurs when the model learns noise from a small dataset. Adding more varied images helps the model generalize.

Option A (increase learning rate) is wrong because it may cause divergence or unstable training, not directly reduce overfitting. Option B (reduce image size) is wrong because it can lose important features and may even increase overfitting. Option C (increase training epochs) is wrong because it can actually increase overfitting by allowing the model to memorize more noise.

19
MCQeasy

Your application needs to determine whether two photos of the same person are of the same individual, even if they are from different angles. Which Azure AI service should you use?

A.Azure AI Video Indexer
B.Azure AI Custom Vision
C.Azure AI Vision OCR
D.Azure AI Face
AnswerD

Azure AI Face provides face verification and identification, comparing facial features across images to determine whether two photos show the same individual. Its detection and matching models tolerate pose and angle variation, satisfying the stem's requirement to compare photos taken from different angles.

Why this answer

Azure AI Face provides face verification APIs that compare two faces and return a confidence score indicating whether they belong to the same person. It uses deep learning models trained to handle variations in pose, lighting, and expression, making it ideal for matching photos of the same individual from different angles.

Exam trap

The trap here is that candidates may confuse the generic 'detect faces in video' capability of Video Indexer with the dedicated face verification API of the Face service, or assume Custom Vision can be trained for face matching without realizing it lacks built-in pose-invariant comparison.

How to eliminate wrong answers

Option A is wrong because Azure AI Video Indexer is designed for extracting insights from video content (e.g., speech, faces, objects) and does not provide a direct face comparison API for still images. Option B is wrong because Azure AI Custom Vision requires training a custom model with labeled images for specific classification or object detection tasks, not for out-of-the-box face verification across pose variations. Option C is wrong because Azure AI Vision OCR (Optical Character Recognition) extracts text from images and has no capability to analyze or compare facial features.

20
MCQhard

You are using Azure AI Custom Vision to detect defects in fabric rolls. After training an object detection model, you notice that it often misses small tears. You have a large dataset of labeled images, but the tears are very small relative to the image size. What should you do to improve the model's detection of small tears?

A.Train a new model with images where the tears are larger in the frame
B.Increase the number of training iterations
C.Use the 'General' domain instead of 'Product' domain
D.Increase the model's confidence threshold
AnswerA

To improve detection of small objects, you can crop or zoom in on the regions containing the tears so they occupy a larger portion of the image. This gives the model more pixels to learn from and makes the features more salient. Retraining with such images helps the model detect small tears more reliably, as it focuses on the relevant details.

Why this answer

Small object detection is challenging because the objects occupy few pixels. By training with images where the tears are larger in the frame—through cropping or zooming—you provide the model with more detailed features to learn from. This approach directly addresses the issue and improves the model's ability to detect small tears.

Exam trap

The trap here is thinking that increasing iterations or changing domains will solve small object detection, but the key is to make the objects larger in the training images.

21
Multi-Selectmedium

Which TWO Azure AI services can be used to detect objects in images?

Select 2 answers
A.Video Indexer
B.Face API
C.Custom Vision
D.Azure AI Document Intelligence
E.Computer Vision Object Detection API
AnswersC, E

Custom Vision trains a bespoke object detection model on your own labelled images, returning bounding boxes per class. It satisfies the stem's requirement to detect objects, unlike classification-only or face-based services, because you control the labelled dataset and exported model.

Why this answer

Custom Vision (C) is correct because it is an Azure AI service that lets you train and deploy custom image classification and object detection models, returning bounding boxes for detected objects in images. Computer Vision Object Detection API (E) is correct because the Azure AI Vision (Computer Vision) service provides a prebuilt object detection capability that identifies common objects and their bounding box coordinates in an image. Video Indexer (A) is not the right fit here since it focuses on analyzing and indexing video and audio content rather than being an image object detection service.

Face API (B) only detects and analyzes human faces (and related attributes), not general objects. Azure AI Document Intelligence (D) extracts text, key-value pairs, and structured data from documents, so it does not perform general object detection in images.

Exam trap

The trap here is that candidates often confuse the general-purpose Computer Vision Object Detection API (which is pre-trained on common objects) with Custom Vision (which requires custom training), but both are valid for object detection depending on the scenario, and the question asks for two services that can detect objects, making both C and E correct.

22
MCQmedium

A company uses the Computer Vision Image Analysis API to generate captions for images. The captions are often too generic. How can they improve the descriptiveness of captions?

A.Use the Object Detection API instead.
B.Train a custom model with Custom Vision.
C.Increase the confidence threshold for captions.
D.Use the Dense Captioning feature.
AnswerD

Dense Captioning generates multiple descriptions for distinct regions within an image, not one caption for the whole frame. This satisfies the stem's requirement for richer, more descriptive output, since generic captions stem from whole-image summarisation. It also returns bounding boxes, adding spatial detail the standard caption endpoint cannot provide.

Why this answer

The Dense Captioning feature in the Computer Vision Image Analysis API generates one-sentence descriptions for each of up to 10 regions detected in an image, providing more specific and detailed captions than the single generic caption. This directly addresses the problem of captions being too generic by breaking the image into meaningful areas and describing each one individually.

Exam trap

The trap here is that candidates confuse increasing the confidence threshold with improving descriptiveness, when in reality it only reduces the number of captions returned without adding detail.

How to eliminate wrong answers

Option A is wrong because the Object Detection API identifies and locates objects with bounding boxes but does not generate descriptive captions or improve caption descriptiveness. Option B is wrong because training a custom model with Custom Vision requires labeled images and is designed for classification or object detection, not for generating richer captions from the existing Image Analysis API. Option C is wrong because increasing the confidence threshold for captions only filters out lower-confidence results, making captions less frequent or more conservative, not more descriptive.

23
MCQeasy

You need to build a solution that detects whether a person is wearing a hard hat in a construction site image. Which Azure AI service should you use?

A.Azure AI Video Indexer
B.Azure AI Face
C.Azure AI Custom Vision
D.Azure AI Document Intelligence
AnswerC

Custom Vision trains an image classification or object detection model on your own labelled hard-hat images, satisfying the stem's specific detection requirement. Prebuilt Computer Vision models lack a hard-hat class, so custom training is required.

Why this answer

Azure AI Custom Vision is the correct service because it allows you to train a custom object detection model to identify specific objects—such as a hard hat—in images. Unlike pre-built services, Custom Vision lets you upload labeled images of workers with and without hard hats, train a model, and then use it to detect hard hat presence in construction site photos. This tailored approach is necessary because hard hat detection is a specialized use case not covered by generic vision APIs.

Exam trap

The trap here is that candidates often confuse Azure AI Custom Vision with pre-built vision services like Azure AI Face or Video Indexer, assuming they can be repurposed for custom object detection, but only Custom Vision allows training on your own labeled images for specific items like hard hats.

How to eliminate wrong answers

Option A is wrong because Azure AI Video Indexer is designed for analyzing video content (e.g., extracting transcripts, faces, and scenes) and does not support custom object detection for specific items like hard hats in static images. Option B is wrong because Azure AI Face is specialized for detecting and analyzing human faces (e.g., attributes like age, emotion, or identity) and cannot identify objects such as hard hats. Option D is wrong because Azure AI Document Intelligence (formerly Form Recognizer) is built for extracting text, tables, and key-value pairs from documents, not for object detection in images.

24
Multi-Selectmedium

You are designing a solution that uses Azure AI Vision to analyze images uploaded by users. You need to extract text and also generate a descriptive caption for each image. You want to use the Image Analysis API. Which two capabilities should you enable? (Choose two.)

Select 2 answers
A.Read
B.Tags
C.Brands
D.Caption
E.Objects
AnswersA, D

The Read capability in Image Analysis extracts printed and handwritten text from images and returns it as structured lines and words. This directly satisfies the requirement to extract text from user-uploaded images. It is the correct choice because it provides OCR functionality within the same API call, avoiding the need to call a separate OCR endpoint.

Why this answer

The Read capability extracts text, and the Caption capability generates a descriptive sentence for the image. Together they satisfy both requirements. Tags, Objects, and Brands provide different types of analysis that are not needed here.

Enabling Read and Caption allows a single API call to return both OCR results and a caption, streamlining the solution.

Exam trap

The trap here is confusing Tags with Caption; Tags provide keywords, whereas Caption produces a full descriptive sentence, which is what the scenario requires.

25
MCQmedium

You are building an Azure AI Vision solution that analyzes live video from a camera mounted on a delivery truck. The solution must read street signs in real time and return the recognized text with bounding box coordinates. You need to minimize latency and cost. Which Azure AI Vision feature should you use?

A.Azure AI Document Intelligence prebuilt-read model
B.Custom Vision object detection model
C.Read API with asynchronous processing
D.Optical character recognition (OCR) synchronous API
AnswerD

The synchronous OCR API is designed for near real-time, single-image text extraction and returns lines and words with bounding box coordinates. It is ideal for live video frames because you can call it per frame with low latency and pay only for the images you submit. It supports printed text in many languages and gives the positional data required to overlay results on the video.

Why this answer

The synchronous OCR API is the correct choice because it extracts printed text with bounding boxes in a single call, which suits low-latency, real-time video frame analysis. The Read API and Document Intelligence are better for documents and asynchronous processing, and Custom Vision detects objects rather than reading text. Using synchronous OCR minimizes latency and cost for live street sign recognition.

Exam trap

The trap here is assuming that the Read API is always the best OCR choice, when its asynchronous, document-oriented design makes it unsuitable for real-time video frames.

26
MCQhard

A logistics company uses Azure AI Vision to analyze images of packages on conveyor belts. They need to detect damaged packages and read tracking numbers. The solution must process high throughput (1000 images per minute) with low latency (<500ms per image). The images are captured by fixed cameras. Which approach should you recommend?

A.Use Azure AI Document Intelligence to process package labels
B.Train a single Custom Vision model that detects damage and reads tracking numbers using OCR
C.Use Azure AI Video Indexer to analyze the video stream from cameras
D.Use Azure AI Vision OCR Read API for tracking numbers and a separate Custom Vision model for damage detection
AnswerD

This approach splits the tasks: the OCR Read API reads tracking numbers, and a Custom Vision model detects damage. Both can run in parallel or be deployed at the edge to meet latency and throughput. This is the only technically feasible option.

Why this answer

It combines Azure AI Vision OCR Read API (for text extraction) with a separate Custom Vision model (for damage detection). This leverages the specialized strengths of each service: the OCR Read API is optimized for text reading with high accuracy, while Custom Vision excels at object detection. The latency requirement (<500ms per image) can be met through parallel processing or edge deployment, and throughput can be achieved by scaling API calls.

Option B is incorrect because Custom Vision does not include built-in OCR capability; it cannot read tracking numbers. Options A and C are less suitable: Document Intelligence is designed for structured documents, not real-time package analysis, and Video Indexer is for video streams, not still images.

Exam trap

The trap is believing that Custom Vision can perform OCR natively. Custom Vision only handles image classification and object detection; text reading requires a dedicated OCR service. Candidates may think a single model is simpler, but it's not technically possible.

How to eliminate wrong answers

Option A is wrong because Azure AI Document Intelligence is designed for structured document extraction (e.g., invoices, forms), not for real-time damage detection on conveyor belt images, and its latency is typically higher than 500ms per image. Option C is wrong because Azure AI Video Indexer is optimized for analyzing pre-recorded video streams with indexing and metadata extraction, not for real-time, low-latency processing of individual images at 1000 images per minute. Option D is wrong because using two separate services (Azure AI Vision OCR Read API for tracking numbers and a separate Custom Vision model for damage detection) introduces additional network round-trips and processing overhead, likely exceeding the 500ms latency budget per image.

27
MCQhard

You have an Azure AI Vision resource named MyVisionService. You run the above Azure CLI command and get the keys. Your application uses key1 for authentication. You need to rotate the keys without downtime. What should you do?

A.Delete and recreate the Cognitive Services resource
B.Regenerate key1 immediately and update the application to use the new key1
C.Regenerate both keys at the same time
D.Update the application to use key2, then regenerate key1
AnswerD

Both keys are valid simultaneously, so switching the application to key2 first keeps authentication working while key1 is regenerated. This ordering avoids downtime, whereas regenerating key1 while still in use would immediately break authentication for the running application.

Why this answer

It enables key rotation without downtime. By first updating the application to use key2 (the secondary key), you ensure that authentication continues to work while key1 is being regenerated. After key1 is regenerated, you can optionally update the application back to key1 at a later time.

This pattern is standard for Azure Cognitive Services to maintain continuous access.

Exam trap

The trap here is that candidates may think regenerating the key currently in use is acceptable if done quickly, but Azure explicitly requires using the secondary key to avoid any period of invalid credentials.

How to eliminate wrong answers

Option A is wrong because deleting and recreating the Cognitive Services resource would cause a complete loss of service and all associated configuration, resulting in significant downtime. Option B is wrong because regenerating key1 immediately would invalidate the key currently used by the application, causing authentication failures and downtime until the application is updated with the new key1. Option C is wrong because regenerating both keys at the same time would invalidate all active keys, leaving no valid key for the application to use, causing immediate downtime.

28
Multi-Selectmedium

A company is building a computer vision solution using Azure AI Vision to analyze images of retail shelves. The solution must detect product presence and read expiration dates. Which TWO Azure AI Vision features should be used?

Select 2 answers
A.Face detection
B.Brand detection
C.Object detection
D.Optical Character Recognition (OCR)
E.Image captioning
AnswersC, D

Object detection returns bounding boxes with labels for each product on the shelf, directly satisfying the product-presence requirement. Unlike image classification, which assigns a single label per image, detection localises multiple distinct items, so the solution can confirm which products are present and where before reading their expiration dates.

Why this answer

Object detection (C) is correct because it locates and classifies multiple objects within an image, which is exactly what is needed to determine whether specific products are present on retail shelves. Optical Character Recognition (D) is correct because OCR extracts printed or handwritten text from images, enabling the solution to read expiration dates printed on product packaging. Face detection (A) only identifies human faces and their attributes, which is irrelevant to detecting products or reading dates.

Brand detection (B) identifies known company logos but does not determine product presence or read expiration dates. Image captioning (E) generates a natural-language description of an image and cannot reliably detect specific products or extract date text.

Exam trap

Microsoft Azure often tests the distinction between object detection and image classification or captioning, where candidates mistakenly choose image captioning for product presence instead of object detection, which provides precise localization and identification.

29
MCQmedium

You are developing a solution to detect defects on a manufacturing assembly line using computer vision. The solution must classify images as 'defective' or 'non-defective'. You have a limited set of labeled images (500 per class). Which approach should you recommend?

A.Use Azure AI Vision Image Analysis with a pre-built model
B.Use Azure AI Custom Vision with image classification
C.Use Azure AI Custom Vision with object detection
D.Train a deep learning model from scratch using Azure Machine Learning
AnswerB

Azure AI Custom Vision image classification handles small labelled datasets through transfer learning, fine-tuning a pre-trained model on your 500 images per class. This satisfies the limited-data constraint directly, unlike object detection, which needs bounding boxes, or training from scratch, which would overfit.

Why this answer

Azure AI Custom Vision with image classification is the best choice because it allows you to fine-tune a pre-trained deep learning model on your limited dataset (500 images per class) to classify images as 'defective' or 'non-defective'. This approach requires minimal data and expertise compared to training from scratch, and it is specifically designed for custom classification tasks with small datasets.

Exam trap

The trap here is that candidates may confuse image classification (assigning a single label to the whole image) with object detection (locating objects), or assume that a pre-built model can be retrained for custom classes, when in fact Azure AI Custom Vision is the correct service for custom classification with limited data.

How to eliminate wrong answers

Option A is wrong because Azure AI Vision Image Analysis pre-built models are designed for general-purpose tasks (e.g., describing images, detecting common objects) and cannot be retrained on custom classes like 'defective' vs 'non-defective'. Option C is wrong because object detection identifies and locates multiple objects within an image, which is overkill for a simple binary classification task where only the presence of a defect matters, not its location. Option D is wrong because training a deep learning model from scratch with only 500 images per class would likely result in poor generalization and overfitting, requiring significantly more data and computational resources.

30
MCQmedium

You are building a solution to automatically tag images uploaded to an Azure Storage blob container using Azure AI Vision. The solution must process images as soon as they are uploaded. Which service should you use to trigger the image analysis?

A.Azure Functions with a timer trigger
B.Azure Event Grid with an Azure Function trigger
C.Azure Batch with a job schedule
D.Azure Logic Apps with a recurrence trigger
AnswerB

Event Grid delivers blob-created events the instant an upload completes, and its Azure Function trigger invokes analysis immediately. This satisfies the requirement to process images as soon as they are uploaded, unlike polling or scheduled batch approaches.

Why this answer

Azure Event Grid is the correct choice because it provides a serverless event-driven architecture that can react to blob storage events (e.g., BlobCreated) in near real-time. By configuring an Event Grid subscription on the storage account, you can trigger an Azure Function that uses Azure AI Vision to analyze the image as soon as it is uploaded, without polling or scheduled checks.

Exam trap

The trap here is that candidates often confuse scheduled triggers (timer/recurrence) with event-driven triggers, assuming any automated trigger will work, but the requirement for 'as soon as they are uploaded' demands an event-driven service like Event Grid, not a polling-based scheduler.

How to eliminate wrong answers

Option A is wrong because a timer trigger runs on a fixed schedule (e.g., every 5 minutes), which introduces latency and cannot react immediately to uploads; it would require polling the container for new blobs. Option C is wrong because Azure Batch is designed for large-scale parallel compute jobs with job schedules, not for real-time event-driven triggers on individual blob uploads. Option D is wrong because a recurrence trigger in Logic Apps also runs on a schedule, not event-driven, and would similarly require polling, missing the immediate processing requirement.

31
MCQeasy

A company wants to build a mobile app that recognizes and tags landmarks in photos taken by tourists. They need a prebuilt model that requires no training and can identify thousands of famous places worldwide. Which Azure AI Vision feature should they use?

A.Image Analysis with the Landmarks feature
B.Face API with the Identify feature to match faces to famous people
C.Computer Vision with the Describe feature to generate captions of the photos
D.Custom Vision with an object detection project trained on landmark images
AnswerA

The Image Analysis service includes a Landmarks feature that can identify thousands of famous landmarks from around the world. It is a prebuilt model, so no training is required, and it returns the name and confidence score for recognized landmarks. This directly meets the requirement for a no-training solution to tag tourist photos.

Why this answer

The Landmarks feature in Image Analysis is a prebuilt model that recognizes thousands of famous landmarks and returns their names. It requires no training and is designed exactly for scenarios like tagging tourist photos. Other options either require training, are for different purposes, or do not provide specific landmark names.

Exam trap

The trap here is confusing the Landmarks feature with general image description, which may mention a landmark but does not provide a structured landmark identifier.

32
Multi-Selectmedium

You are developing a solution that uses Azure AI Vision to analyze images of products on an assembly line. You need to detect whether each product has a specific logo and also read the serial number printed on it. Which two Azure AI Vision services should you use? (Choose two.)

Select 2 answers
A.Azure AI Custom Vision
B.Azure AI Vision Read API
C.Azure AI Document Intelligence
D.Azure AI Vision Spatial Analysis
E.Azure AI Face API
AnswersA, B

Custom Vision allows you to train a custom image classifier or object detector to recognize specific logos. You can label images of products with and without the logo, train a model, and then use it to detect the presence of the logo on new images. This meets the requirement to detect whether each product has the specific logo.

Why this answer

To detect a specific logo and read a serial number, you need two distinct capabilities: custom object detection and text extraction. Custom Vision can be trained to recognize the logo, while the Read API extracts the serial number from the image. Together, they fulfill both requirements without unnecessary complexity.

Exam trap

The trap here is assuming that a single service like Document Intelligence can handle both tasks, but it is not designed for logo detection on products.

33
MCQhard

A security company uses Azure AI Face API to analyze surveillance footage. They need to detect faces in low-light images and obtain face bounding boxes, but they do not need to identify individuals. They also want to minimize cost and avoid unnecessary features. Which Face API operation should they call?

A.Face - Detect with returnFaceId=true and returnFaceLandmarks=true
B.Face - Verify with two face IDs
C.Face - Detect with returnFaceId=false and returnFaceLandmarks=false
D.Face - Identify with a person group
AnswerC

This operation returns face bounding boxes without generating face IDs or landmarks, which is exactly what is required for detection only. It minimizes cost because face ID generation is a billable feature. By setting both parameters to false, the response includes only the rectangle coordinates and basic attributes if requested, aligning with the need to detect faces in low-light images without identifying individuals.

Why this answer

To detect faces without identification, the Face - Detect operation should be called with returnFaceId and returnFaceLandmarks set to false. This returns bounding boxes and avoids the cost of generating face IDs. Identify and Verify are for matching or comparing identities, which are not needed here, and enabling face IDs or landmarks adds unnecessary expense.

Exam trap

The trap here is thinking that face detection always requires face IDs, when in fact face IDs are only needed for identification or verification and can be disabled to reduce cost.

34
MCQeasy

You need to analyze videos stored in Azure Blob Storage to detect objects and generate timestamps. Which Azure service should you use?

A.Azure Custom Vision
B.Azure Form Recognizer
C.Azure Computer Vision
D.Azure Video Indexer
AnswerD

Azure Video Indexer extracts insights from video, including object detection with timestamps, and can ingest content directly from Azure Blob Storage. This satisfies the stem's requirements for both object detection and timestamp generation on stored videos.

Why this answer

Azure Video Indexer (D) is the correct choice because it is specifically designed to analyze videos, extracting insights such as object detection, scene segmentation, and timestamps. It uses AI models to process video content stored in Azure Blob Storage and generates a timeline of detected objects, making it ideal for this scenario.

Exam trap

The trap here is that candidates often confuse Azure Computer Vision (image analysis) with video analysis, overlooking that Computer Vision lacks native video processing and timestamp generation, while Video Indexer is the dedicated service for end-to-end video insights.

How to eliminate wrong answers

Option A is wrong because Azure Custom Vision is a service for training custom image classification and object detection models on images, not for analyzing pre-recorded videos with timestamp generation. Option B is wrong because Azure Form Recognizer is designed to extract text and structure from documents (e.g., invoices, forms), not for video analysis or object detection. Option C is wrong because Azure Computer Vision provides image analysis APIs (e.g., object detection in static images) but lacks native video processing capabilities and timestamp generation; it would require additional custom logic to handle video frames sequentially.

35
MCQhard

You are designing a solution to detect brand logos in social media images. The logos vary in size and orientation. You need to achieve high accuracy with minimal false positives. Which approach should you recommend?

A.Use Azure Computer Vision Describe API to generate captions and filter by logo mentions.
B.Train an Azure Custom Vision object detection model with labeled logo images.
C.Use Azure Computer Vision Analyze API with domain-specific models.
D.Use Azure Form Recognizer to extract logo positions from images.
AnswerB

Object detection returns bounding boxes, so it handles logos that vary in size and orientation, unlike classification which only labels the whole image. Training on your labelled logo images tunes the model to your brands, satisfying the high-accuracy, low-false-positive constraint.

Why this answer

Azure Custom Vision allows you to train a custom object detection model with your own labeled dataset of brand logos, enabling high accuracy for specific logo shapes, sizes, and orientations. This approach directly addresses the need for minimal false positives by learning the exact visual features of the logos, unlike generic pre-built models.

Exam trap

The trap here is that candidates confuse Azure Computer Vision's pre-built domain-specific models (which cover only landmarks, celebrities, and general objects) with the ability to detect custom logos, leading them to choose option C instead of recognizing that Custom Vision is required for custom object detection.

How to eliminate wrong answers

Option A is wrong because the Describe API generates natural language captions and is not designed for precise object detection or localization; filtering by logo mentions would be unreliable and produce many false positives. Option C is wrong because the Analyze API with domain-specific models (e.g., landmarks, celebrities) does not include a pre-built model for brand logos, so it cannot detect arbitrary logos with high accuracy. Option D is wrong because Azure Form Recognizer is specialized for extracting text and structured data from documents (e.g., invoices, forms), not for detecting or localizing visual objects like logos in images.

36
MCQhard

A company uses Azure AI Vision Image Analysis to generate captions for product photos. The solution must return a caption in English and a confidence score for each image. You call the Image Analysis API with the caption feature. The response does not include a confidence score. What should you do to obtain confidence scores for the captions?

A.Use the older Computer Vision v3.2 Analyze Image API with the 'description' visual feature.
B.Call the Image Analysis API twice and compare the captions to derive a confidence score.
C.Use the denseCaptions feature instead of captions.
D.Set the language parameter to 'en' and include the 'confidence' query parameter.
AnswerA

The legacy Computer Vision v3.2 Analyze Image API with the description visual feature returns a caption along with a confidence score between 0 and 1. The newer Image Analysis 4.0 caption feature does not include a confidence score. To meet the requirement, you must call the v3.2 endpoint, which still provides the confidence value for the generated description.

Why this answer

The legacy Computer Vision v3.2 Analyze Image API with the description feature returns a caption and a confidence score, which the newer Image Analysis 4.0 caption feature does not. To obtain confidence scores for captions, you must use the older API version. Other options either use features that lack confidence scores or rely on invalid parameters or workarounds that cannot produce a true confidence value.

Exam trap

The trap here is assuming that the newest Image Analysis caption feature includes a confidence score, when only the legacy v3.2 description feature returns one.

37
MCQeasy

Your team is building a mobile app that uses Azure Custom Vision to classify plant species. The app must work offline and sync labeled images when connectivity is restored. Which SDK feature should you use?

A.Azure IoT Edge runtime on the phone
B.Azure API Management with caching
C.Export the model as a TensorFlow or CoreML model for on-device inference
D.Continuous deployment integration
AnswerC

Exported model runs offline.

Why this answer

Azure Custom Vision allows exporting trained models to formats like TensorFlow, CoreML, ONNX, or Docker for on-device inference. This enables the mobile app to run classification locally without network connectivity, and the Custom Vision SDK includes a method to upload labeled images for offline training sync when connectivity is restored.

Exam trap

The trap here is that candidates confuse offline inference with edge computing (IoT Edge) or API caching, not realizing that Custom Vision's export feature is the only option that provides a local model for on-device classification without requiring a network connection.

How to eliminate wrong answers

Option A is wrong because Azure IoT Edge runtime is designed for edge devices like gateways or industrial controllers, not for mobile phones, and it does not provide offline inference or image sync capabilities for Custom Vision. Option B is wrong because Azure API Management with caching only caches API responses to reduce latency, but it does not enable offline model execution or local image storage and sync. Option D is wrong because continuous deployment integration automates model deployment pipelines but does not address offline inference or offline image labeling and sync on a mobile device.

38
MCQmedium

Refer to the exhibit. You have trained an object detection model in Azure Custom Vision. The model is published as 'defect-model'. You need to deploy this model to a Docker container for on-premises inference using the Azure IoT Edge runtime. What should you do first?

A.Create an Azure Container Registry and push the Custom Vision base image.
B.Export the model as a Docker container (e.g., TensorFlow) using the Custom Vision portal.
C.Use the Custom Vision prediction API to call the published endpoint from the edge device.
D.Retrain the model with more images to improve mAP.
AnswerB

IoT Edge requires a containerised model image, so the model must first be exported from Custom Vision as a Docker container for a supported framework such as TensorFlow, producing the artefact later deployed to the edge device.

Why this answer

To deploy a Custom Vision model to an Azure IoT Edge device, you must first export the model as a Docker container (e.g., TensorFlow, ONNX, or DockerFile) from the Custom Vision portal. This export creates a container image that can be deployed to Azure Container Registry and then used as a module in an IoT Edge deployment. Without this export step, you cannot create the containerized module required for on-premises inference.

Exam trap

The trap here is that candidates may think they can directly use the cloud prediction endpoint on an edge device, but Azure IoT Edge requires a containerized module for local execution, making the export step mandatory before any deployment.

How to eliminate wrong answers

Option A is wrong because you do not push the Custom Vision base image; instead, you export the trained model as a container from the portal, which generates a Docker image that you then push to Azure Container Registry. Option C is wrong because calling the prediction API from the edge device would require internet connectivity and defeats the purpose of on-premises inference; IoT Edge runs modules locally without constant cloud access. Option D is wrong because retraining the model to improve mAP is a separate optimization step and does not address the immediate deployment requirement to create a container for IoT Edge.

39
MCQeasy

You need to extract handwritten text from scanned forms. Which Azure Computer Vision feature should you use?

A.OCR API (optical character recognition)
B.Tag API
C.Read API
D.Describe API
AnswerC

The Read API performs optical character recognition on both printed and handwritten text, directly satisfying the stem's handwriting requirement. Unlike the older OCR API, which handles printed text only, Read supports handwriting through Microsoft Entra ID-authenticated Computer Vision resources, returning extracted lines and words for scanned forms.

Why this answer

The Read API is specifically designed for extracting printed and handwritten text from images and documents, including scanned forms. It uses advanced deep-learning models optimized for text recognition and is the correct service for this task in Azure Computer Vision.

Exam trap

The trap here is that candidates confuse the legacy OCR API (which only handles printed text) with the Read API (which handles both printed and handwritten text), leading them to select Option A incorrectly.

How to eliminate wrong answers

Option A is wrong because the OCR API is a legacy service that only extracts printed text and does not support handwritten text recognition. Option B is wrong because the Tag API returns a list of content tags (objects, concepts) based on the image, not text extraction. Option D is wrong because the Describe API generates human-readable captions describing the image content, not text extraction.

40
MCQeasy

A developer is building an Azure AI Vision solution that must detect and locate multiple objects in an image, such as cars, people, and traffic lights. The solution must return bounding boxes for each detected object. Which Azure AI Vision capability should the developer use?

A.Image classification
B.Object detection
C.Optical character recognition (OCR)
D.Face detection
AnswerB

Object detection in Azure AI Vision identifies and localizes multiple objects within an image, returning bounding boxes and labels for each detected object. It is designed for scenarios like detecting cars, people, and traffic lights simultaneously. This capability directly provides the required bounding box coordinates and object labels, making it the correct choice for the scenario.

Why this answer

Object detection is the correct capability because it identifies and localizes multiple objects in an image, returning bounding boxes and labels for each. Image classification only labels the whole image, OCR extracts text, and face detection is limited to faces. For detecting cars, people, and traffic lights with bounding boxes, object detection is the appropriate Azure AI Vision feature.

Exam trap

The trap here is confusing object detection with image classification, where classification only labels the overall image and does not provide bounding boxes for individual objects.

41
MCQeasy

A media company needs to automatically generate descriptive captions for thousands of archived photographs stored in Azure Blob Storage. The solution must be fully managed and require no model training. Which Azure AI Vision capability should they use?

A.Custom Vision classification model
B.Face API person identification
C.Azure AI Document Intelligence prebuilt receipt model
D.Image captioning in Azure AI Vision
AnswerD

Azure AI Vision's image captioning feature generates human-readable descriptions for images without any training. It is a prebuilt capability available through the Image Analysis API, ideal for captioning large archives. The media company can call this API on each blob and store the returned caption, meeting the no-training requirement.

Why this answer

Image captioning in Azure AI Vision is a prebuilt feature that generates natural language descriptions for images without any training. The media company can process each stored photograph and receive a caption, directly meeting the scenario's need for automated, managed captioning. Other options either require training, target documents, or focus only on faces, so they do not provide the required broad image description.

Exam trap

The trap here is assuming that any Azure AI service that processes images can generate captions, when only Azure AI Vision's image captioning feature provides that specific prebuilt capability.

42
MCQmedium

You are developing an application that uses Azure AI Vision to analyze images of products on an assembly line. The application must identify the presence of specific objects, such as screws, bolts, and washers, and return their bounding boxes. You have a limited set of labeled images for each object type. Which Azure service should you use to train a model that meets these requirements?

A.Azure AI Vision Face API
B.Azure AI Vision Image Analysis with the 'objects' feature
C.Azure Video Indexer
D.Azure Custom Vision object detection
AnswerD

Azure Custom Vision object detection allows you to train a model on your own labeled images to detect specific objects and return bounding boxes. It supports custom object types like screws, bolts, and washers, and can work with a limited set of images. This service is designed exactly for this scenario.

Why this answer

Azure Custom Vision object detection is the correct service because it enables you to train a custom model using your labeled images to detect specific objects and return bounding boxes. It is designed for scenarios where you need to identify custom objects with a limited dataset.

Exam trap

The trap here is assuming that the pre-built 'objects' feature in Image Analysis can be customized or that it can detect specific industrial parts without training.

43
MCQmedium

A company uses Azure Computer Vision to moderate user-generated content. The solution must detect adult content and flag it. Which API should you call?

A.Read API
B.Analyze API with visualFeatures set to 'Adult'
C.Detect API
D.Describe API
AnswerB

The Analyze API accepts visualFeatures, and setting it to 'Adult' returns adult and racy classifications with confidence scores, directly satisfying the requirement to detect and flag adult content. Other features such as Tags or Description cannot return moderation ratings.

Why this answer

The Analyze API with the visualFeatures parameter set to 'Adult' is the correct choice because Azure Computer Vision's Analyze Image operation includes an 'Adult' category that specifically detects adult, racy, and gory content in images. This API returns a confidence score (0 to 1) for each category, allowing the solution to flag content based on a threshold. The other APIs do not provide adult content moderation capabilities.

Exam trap

The trap here is that candidates may confuse the Analyze API's 'Adult' feature with the 'Description' or 'Tags' features, assuming that general image analysis can detect adult content, but only the explicit 'Adult' visualFeature parameter provides the specialized moderation scores.

How to eliminate wrong answers

Option A is wrong because the Read API is designed for optical character recognition (OCR) to extract printed and handwritten text from images, not for detecting adult content. Option C is wrong because the Detect API is used for object detection (identifying and locating objects within an image), not for content moderation. Option D is wrong because the Describe API generates human-readable captions describing the content of an image, but it does not include adult content classification or scoring.

44
MCQeasy

A developer needs to build a mobile app that identifies dog breeds from photos. They have a small dataset of labeled images and want to train a custom model with minimal machine learning expertise. Which Azure service should they use?

A.Azure AI Vision Image Analysis
B.Azure AI Face API
C.Azure Machine Learning designer
D.Azure AI Custom Vision
AnswerD

Custom Vision is a no-code/low-code service that allows you to upload labeled images and train a classification model. It provides a simple interface and SDKs, making it ideal for developers with limited ML expertise. It supports multi-class classification, which is suitable for identifying dog breeds from photos. The service handles training and deployment, so the developer can focus on the app.

Why this answer

Custom Vision is designed for custom image classification and object detection with minimal ML expertise. It supports training on small datasets and provides a user-friendly interface. The other services either require more ML knowledge, offer only prebuilt models, or are for a different domain entirely.

Custom Vision directly meets the need to identify dog breeds from photos.

Exam trap

The trap here is assuming that Azure AI Vision prebuilt models can classify specific breeds, when they only provide generic tags and cannot be customized without Custom Vision.

45
MCQmedium

A media company wants to automatically generate alt text for images on its news site using Azure AI Vision. The images are stored in Azure Blob Storage, and the solution must run serverless and respond within seconds. Which approach should you use?

A.Deploy an Azure Function that calls the Image Analysis caption feature with the image URL and returns the generated text.
B.Use the Face API to describe each image based on detected people.
C.Run a batch transcription job over the image files to extract descriptive text.
D.Create a Custom Vision classification project and train it on the site's images to produce captions.
AnswerA

Azure Functions provides a serverless compute model, and Image Analysis 4.0 can accept a publicly reachable image URL, so the function can request the caption feature and return alt text quickly. This combination meets the serverless and low-latency requirements without managing infrastructure or moving image bytes unnecessarily.

Why this answer

Image Analysis captioning produces a one-sentence description suitable for alt text, and it accepts either image bytes or a URL. Hosting the call in an Azure Function keeps the solution serverless and event-driven, so responses come back in seconds. Custom Vision, Face, and speech batch transcription do not generate general-purpose image descriptions and therefore cannot meet the requirement.

Exam trap

The trap here is confusing image classification or face analysis with captioning, which is the only Azure AI Vision feature that produces descriptive natural-language text.

46
MCQeasy

A company needs to identify and tag products on store shelves using a custom model. They have a large dataset of labeled images with bounding boxes around each product. They want to train a model that can detect multiple products in new images and return their locations. Which Azure service should they use?

A.Azure AI Face, using the face detection feature
B.Azure AI Vision, using the prebuilt object detection model
C.Azure AI Custom Vision, using the Object Detection project type
D.Azure AI Custom Vision, using the Classification project type
AnswerC

Custom Vision supports two project types: Classification and Object Detection. Object Detection is designed to identify multiple objects within an image and return bounding box coordinates for each. Since the dataset includes bounding boxes and the goal is to detect and locate products, this project type is the correct choice.

Why this answer

Custom Vision Object Detection is designed for scenarios where you need to detect and locate multiple instances of objects within an image. It uses labeled bounding boxes to train a model that predicts both class labels and bounding box coordinates. Classification, prebuilt Vision models, and Face detection do not provide the required custom object detection with location.

Exam trap

The trap here is assuming that Custom Vision Classification can be used for object detection if you have bounding boxes, but classification ignores bounding box information.

47
MCQhard

A manufacturing company uses Azure AI Custom Vision to detect defects on a production line. The model was trained with 500 images per class and achieves 95% accuracy. After deployment, the model's accuracy drops to 80% due to changes in lighting conditions. What is the most effective first step to improve the model's robustness?

A.Reduce the probability threshold to increase recall.
B.Capture additional images under the new lighting and retrain the model.
C.Use Azure AutoML to automatically find the best algorithm.
D.Add more images from the original lighting conditions to the training set.
AnswerB

Domain shift from changed lighting causes the accuracy drop, so capturing images under the new lighting conditions and retraining lets the model learn those visual features, directly addressing the distribution mismatch rather than tuning unrelated parameters.

Why this answer

The drop in accuracy is caused by a domain shift—specifically, new lighting conditions that were not represented in the original training set. The most effective first step is to capture additional images under the new lighting and retrain the model, as Custom Vision relies on diverse, representative training data to generalize to real-world variations. This directly addresses the root cause by expanding the training distribution to include the new lighting scenario, which is a fundamental principle of supervised learning in computer vision.

Exam trap

The trap here is that candidates may confuse a performance tuning action (like adjusting the probability threshold) with a data quality fix, or assume AutoML can magically fix any accuracy drop, when in fact the root cause is a classic domain shift that requires representative retraining data.

How to eliminate wrong answers

Option A is wrong because reducing the probability threshold increases recall but also increases false positives, which does not improve robustness to lighting changes—it only trades precision for recall without addressing the underlying distribution shift. Option C is wrong because Azure AutoML is designed for automated model selection and hyperparameter tuning, but the problem here is a data distribution mismatch, not a need for a different algorithm; AutoML cannot compensate for missing lighting variations in the training data. Option D is wrong because adding more images from the original lighting conditions does not help the model learn to handle the new lighting; it only reinforces the existing bias toward the old lighting, leaving the domain shift unaddressed.

48
Multi-Selectmedium

A company uses Azure Custom Vision to build a classifier for defect detection on a manufacturing line. They have labeled images of products with and without defects. Which TWO actions should they take to improve model performance?

Select 2 answers
A.Train for more iterations without validation.
B.Use images with balanced numbers of defect and non-defect samples.
C.Set the learning rate manually using the Custom Vision API.
D.Increase the number of images per tag, including variations in lighting and angle.
E.Reduce the number of images per tag to avoid overfitting.
AnswersB, D

Balanced datasets prevent bias toward majority class.

Why this answer

Balanced datasets prevent the model from becoming biased toward the majority class (e.g., non-defect images), which is critical for defect detection where defects are rare. Azure Custom Vision uses a weighted loss function during training, and class imbalance can cause the model to predict the majority class for most inputs, reducing recall for defects. Balanced samples ensure the model learns discriminative features for both classes equally.

Exam trap

The trap here is that candidates may think reducing images prevents overfitting (Option E) or that manual learning rate tuning (Option C) is possible in Custom Vision, but the service abstracts hyperparameter tuning and requires sufficient, varied data for robust defect detection.

49
MCQmedium

You are building a web application that allows users to upload images of restaurant receipts and extract the total amount and merchant name. The receipts may be crumpled, rotated, or have handwritten notes. Which Azure AI service should you use to reliably extract this information?

A.Azure AI Vision Spatial Analysis
B.Azure AI Document Intelligence prebuilt receipt model
C.Azure AI Face API
D.Azure AI Vision Read API
AnswerB

The prebuilt receipt model in Azure AI Document Intelligence is specifically trained to extract key fields such as total, merchant name, transaction date, and line items from receipts. It handles variations in receipt formats, including rotation and crumpling, and returns structured JSON with confidence scores, making it ideal for this scenario without custom parsing.

Why this answer

The prebuilt receipt model in Azure AI Document Intelligence is purpose-built to extract structured fields like total and merchant name from receipts, even when they are crumpled or rotated. It uses a combination of OCR and machine learning to understand receipt layouts and output key-value pairs, eliminating the need for custom parsing or additional services.

Exam trap

The trap here is assuming that the Read API's text extraction is sufficient, but it lacks the semantic understanding to identify specific fields like total amount without extra processing.

50
MCQeasy

Your company wants to moderate user-uploaded images for adult content. Which Azure AI service should you use?

A.Azure AI Content Safety
B.Azure AI Face
C.Azure AI Document Intelligence
D.Azure AI Vision Image Analysis
AnswerA

Azure AI Content Safety provides dedicated image moderation with severity-scored adult, racy, and violent classifications, directly satisfying the requirement to filter user-uploaded images. Unlike Computer Vision, which offers only basic adult/gory flags, Content Safety returns graded severity levels, enabling precise threshold-based blocking aligned with your moderation policy.

Why this answer

Azure AI Content Safety is the correct service because it is specifically designed to detect and moderate inappropriate content, including adult content, in images and text. It provides severity-based classifications (safe, low, medium, high) for categories such as hate, self-harm, sexual, and violence, making it ideal for user-uploaded image moderation.

Exam trap

The trap here is that candidates often confuse Azure AI Vision Image Analysis (which can detect adult content via the 'adult' flag in its Analyze Image API) with the dedicated Azure AI Content Safety service, but the exam expects the service purpose-built for content moderation with granular severity levels and broader category support.

How to eliminate wrong answers

Option B is wrong because Azure AI Face is focused on detecting, analyzing, and recognizing human faces, not on moderating adult content. Option C is wrong because Azure AI Document Intelligence is designed to extract text, key-value pairs, and tables from documents, not to analyze images for adult content. Option D is wrong because Azure AI Vision Image Analysis provides general image descriptions, object detection, and optical character recognition, but lacks the specific content moderation categories and severity scoring needed for adult content detection.

51
MCQmedium

You are developing an app that analyzes images of restaurant receipts. The app must extract the merchant name, transaction date, and total amount from each receipt. You want to minimize development effort and use a prebuilt Azure AI service. Which service should you use?

A.Azure AI Document Intelligence with the prebuilt-receipt model
B.Azure AI Language with custom named entity recognition
C.Azure AI Vision with the Read API
D.Azure AI Custom Vision with a classification model
AnswerA

The prebuilt-receipt model in Azure AI Document Intelligence is specifically trained to extract key fields from receipts, including merchant name, transaction date, and total. It returns structured JSON with these fields, so minimal development effort is needed. It handles common receipt variations and is the recommended prebuilt solution for receipt processing.

Why this answer

Azure AI Document Intelligence provides a prebuilt receipt model that is purpose-built to extract merchant name, transaction date, and total from receipt images. It returns structured data without requiring custom model training or text parsing, which aligns with the goal of minimizing development effort. Other services either return unstructured text or require custom training.

Exam trap

The trap here is assuming that the Read API's OCR output is sufficient for structured field extraction, when it only provides raw text.

52
MCQeasy

A company uses Azure Face API to verify employee identities for building access. They need to ensure that only live faces are used, not photos or videos. Which feature should they enable?

A.Set a high confidence threshold for face matching.
B.Face identification with a large person group.
C.Enable liveness detection using session-based verification.
D.Face detection with attributes such as age and emotion.
AnswerC

Session-based liveness detection challenges the subject to perform a randomised action, then analyses the response to confirm a live person rather than a static photo or replayed video. This satisfies the requirement to reject spoofing attempts during identity verification.

Why this answer

Azure Face API's liveness detection with session-based verification is specifically designed to prevent spoofing attacks using photos, videos, or masks. It analyzes subtle cues such as micro-movements, texture, and depth to confirm the presence of a live person, ensuring that only live faces are accepted for identity verification.

Exam trap

The trap here is that candidates may confuse confidence thresholds or face attributes with liveness detection, not realizing that only session-based verification actively checks for spoofing through motion and depth analysis.

How to eliminate wrong answers

Option A is wrong because setting a high confidence threshold only increases the strictness of face matching scores, but does not differentiate between a live face and a spoofed image or video. Option B is wrong because face identification with a large person group is used to match a detected face against a database of enrolled persons, but it does not verify liveness or detect presentation attacks. Option D is wrong because face detection with attributes like age and emotion extracts demographic and emotional information from a face, but it cannot determine whether the face is live or a reproduction.

53
MCQmedium

A company wants to build a solution that automatically generates alt text for images on their website to improve accessibility. The alt text must be a concise, human-readable description of the image content. Which Azure AI Vision feature should they use?

A.Custom Vision with a classification model trained on descriptive text
B.Image Analysis with the Detect Objects feature
C.Image Analysis with the Caption feature
D.Image Analysis with the Tags feature
AnswerC

The Caption feature generates a concise, human-readable description of an image, which is exactly what is needed for alt text. It is a prebuilt model that requires no training and returns a single sentence describing the image. This directly meets the accessibility requirement.

Why this answer

The Caption feature in Image Analysis is designed to generate a one-sentence description of an image, which is ideal for alt text. It is prebuilt and requires no training. Other features like Tags or Detect Objects provide keywords or bounding boxes but not a coherent description, making them less suitable for accessibility purposes.

Exam trap

The trap here is confusing Tags with Caption; Tags give keywords, while Caption gives a full sentence suitable for alt text.

54
MCQmedium

You are building a web application that lets users upload photos of restaurant menus and receive the extracted text. The menus are often photographed at an angle, with uneven lighting and background clutter. You need an Azure AI Vision capability that returns text lines and words with bounding boxes, and you want to minimize development effort. Which Azure AI Vision feature should you use?

A.Azure AI Custom Vision object detection model
B.Azure AI Vision Image Analysis with the Tags feature
C.Azure AI Face API with the OCR attribute enabled
D.Azure AI Vision OCR (Read) API
AnswerD

The Read API is designed for extracting printed and handwritten text from images, including photos taken at angles or with background noise. It returns text lines, words, and bounding boxes in a structured JSON response, so you can parse and display results with minimal custom code. It directly matches the need for menu text extraction without training a custom model.

Why this answer

The Read API in Azure AI Vision is purpose-built for OCR and returns text lines and words with bounding boxes, making it ideal for extracting menu content from photos. It handles real-world conditions like skew and clutter. The other services either return high-level labels, detect objects, or analyze faces, none of which extract the literal text required by the application.

Exam trap

The trap here is assuming that Image Analysis tags or Custom Vision can read text, when only the dedicated Read OCR capability returns the actual characters and their positions.

55
MCQmedium

A company is building a solution to analyze customer reviews images using Azure AI Vision. They need to extract text from images that may contain both printed and handwritten text. Which feature should they use?

A.Custom Vision
B.OCR API (optical character recognition)
C.Read API
D.Azure AI Document Intelligence
AnswerC

The Read API uses the OCR engine that handles both printed and handwritten text within the same image, satisfying the mixed-content constraint. Legacy OCR and handwriting-only endpoints cannot process both simultaneously, so Read is the only feature meeting this requirement.

Why this answer

The Read API is the correct choice because it is specifically designed to extract text from images containing both printed and handwritten text, using advanced OCR capabilities that support mixed content. Unlike the OCR API, which is optimized for printed text only, the Read API leverages deep learning models to handle varied handwriting styles and complex layouts, making it ideal for analyzing customer review images.

Exam trap

The trap here is that candidates confuse the OCR API with the Read API, assuming both handle handwritten text equally, but the OCR API is limited to printed text while the Read API is the only one that natively supports mixed printed and handwritten content.

How to eliminate wrong answers

Option A is wrong because Custom Vision is a service for training custom image classification and object detection models, not for text extraction. Option B is wrong because the OCR API (optical character recognition) is optimized for printed text and does not reliably extract handwritten text, which is a key requirement. Option D is wrong because Azure AI Document Intelligence (formerly Form Recognizer) is designed for structured document processing (e.g., forms, invoices) and is not the primary service for general text extraction from images with mixed printed and handwritten content.

56
MCQhard

You have a real-time video processing pipeline using Azure AI Video Indexer. You need to detect when a specific person appears in archived video footage. Which approach minimizes latency and cost?

A.Use Video Indexer's face detection and indexing, then search
B.Extract keyframes and use Custom Vision to detect the person
C.Run face detection on every frame using Azure AI Face and store results
D.Use Azure AI Vision to detect faces in video frames and compare against a database
AnswerA

Video Indexer indexes faces once during ingestion, storing face identifiers alongside timestamps, so later searches match against that index rather than re-processing footage. This avoids repeated full-video analysis, minimising both compute cost and query latency for archived footage.

Why this answer

Video Indexer's built-in face detection and indexing automatically identifies and tracks faces during the indexing process, storing the results in a searchable metadata index. To detect when a specific person appears, you can then search the indexed metadata for that person's face ID or name, which avoids re-processing the video and minimizes both latency and cost. This approach leverages the one-time indexing cost and optimized search capabilities rather than running additional AI services on every frame.

Exam trap

The trap here is that candidates often assume Custom Vision or Azure AI Face are needed for custom person detection, overlooking that Video Indexer already provides built-in face detection and search capabilities that are optimized for archived video analysis.

How to eliminate wrong answers

Option B is wrong because extracting keyframes and using Custom Vision requires training a custom model and processing only keyframes, which may miss the person if they appear between keyframes, and the custom training adds overhead and cost without leveraging Video Indexer's built-in face indexing. Option C is wrong because running face detection on every frame using Azure AI Face would incur high compute and API costs per frame, and storing all results creates unnecessary data volume, making it far more expensive and slower than using Video Indexer's pre-indexed search. Option D is wrong because using Azure AI Vision to detect faces in video frames and comparing against a database requires frame-by-frame processing and external database lookups, which introduces latency and cost that Video Indexer's integrated indexing and search avoids.

57
MCQmedium

You are building a mobile app that allows users to take a photo of a product and get detailed information. The app uses Azure AI Custom Vision to classify products. You need to ensure low latency for inference. What should you do?

A.Increase the number of training iterations
B.Use the Azure AI Vision API directly
C.Use Azure Front Door to cache results
D.Export the Custom Vision model as a TensorFlow model and run on-device
AnswerD

Exporting the Custom Vision model as TensorFlow and running inference on-device removes the network round trip to the Azure endpoint entirely, which is the dominant latency source for a mobile app. Local execution satisfies the low-latency constraint.

Why this answer

Exporting the Custom Vision model as a TensorFlow model and running it on-device eliminates network latency entirely. Inference happens locally on the mobile device, which provides the lowest possible latency for real-time classification, especially when network connectivity is poor or inconsistent.

Exam trap

The trap here is that candidates assume cloud-based solutions (like Azure Front Door or Vision API) are always faster, but Microsoft explicitly tests the understanding that on-device inference eliminates network latency and is the optimal choice for low-latency mobile scenarios.

How to eliminate wrong answers

Option A is wrong because increasing the number of training iterations improves model accuracy, not inference latency; latency is determined by model architecture and runtime environment, not training steps. Option B is wrong because using the Azure AI Vision API directly requires a network round-trip to Azure, which introduces higher latency compared to on-device inference, and it does not leverage the custom classification model you built. Option C is wrong because Azure Front Door caches HTTP responses at edge locations, but inference results are dynamic and user-specific (each photo is unique), so caching would rarely hit and cannot reduce the latency of the actual inference call.

58
MCQmedium

You have a computer vision solution that analyzes security camera feeds to detect people and vehicles. The solution uses Azure AI Vision Spatial Analysis. You need to ensure compliance with privacy regulations by blurring detected faces. Which feature should you enable?

A.Use Azure AI Content Safety to filter faces
B.Post-process frames with Azure AI Face client SDK
C.Enable face detection and redact faces using Azure AI Video Indexer
D.Enable face blurring in the Spatial Analysis configuration
AnswerD

Face blurring in the Spatial Analysis configuration redacts detected faces in the video stream before storage or transmission, satisfying the stem's privacy compliance requirement. It operates within the spatial analysis pipeline itself, unlike separate Face service redaction applied after processing.

Why this answer

Azure AI Vision Spatial Analysis includes a built-in face blurring feature that can be enabled directly in the Spatial Analysis configuration. This allows you to automatically blur detected faces in the video feed at the edge or in the cloud, ensuring compliance with privacy regulations without requiring additional services or post-processing steps.

Exam trap

The trap here is that candidates may confuse Azure AI Video Indexer's face redaction capabilities with Spatial Analysis's real-time face blurring, or assume that a separate SDK or service is required for face blurring when it is actually a built-in configuration option in Spatial Analysis.

How to eliminate wrong answers

Option A is wrong because Azure AI Content Safety is designed to detect and filter harmful content (e.g., violence, hate speech) in text, images, and video, not to blur faces. Option B is wrong because post-processing frames with the Azure AI Face client SDK would require additional development effort and latency, and it is not a native feature of Spatial Analysis; the face blurring is already integrated into the Spatial Analysis pipeline. Option C is wrong because Azure AI Video Indexer is a separate service for extracting insights from video files (e.g., transcripts, faces, emotions) and does not provide real-time face blurring for live security camera feeds; it is not part of the Spatial Analysis solution.

59
MCQhard

A company uses the Face API to detect and identify employees for building access. They need to ensure that the system complies with GDPR requirements for biometric data. Which action should they take?

A.Store faces in a secure database and delete after 30 days.
B.Anonymize the face data by blurring key features.
C.Obtain explicit consent from each employee before enrollment.
D.Use encryption for stored face templates.
AnswerC

Biometric templates derived from facial images are special-category personal data under GDPR, requiring a lawful basis beyond legitimate interest. Explicit, freely given consent from each employee before enrolment satisfies Article 9, and employees must be able to withdraw it without detriment.

Why this answer

Under GDPR, biometric data (such as facial recognition templates) is classified as special category data requiring explicit consent for processing. The Face API itself does not manage consent; the responsibility lies with the application layer. Option C is correct because obtaining explicit consent from each employee before enrollment is a fundamental GDPR requirement for lawful processing of biometric data.

Exam trap

The trap here is that candidates often focus on technical security measures (encryption, deletion, anonymization) as sufficient for GDPR compliance, overlooking the foundational legal requirement for explicit consent when processing special category biometric data.

How to eliminate wrong answers

Option A is wrong because merely storing faces in a secure database and deleting after 30 days does not address the GDPR requirement for a lawful basis (e.g., explicit consent) before processing biometric data; retention limits are a separate compliance aspect. Option B is wrong because anonymizing face data by blurring key features would render the Face API ineffective for identification, as the API requires clear facial features to generate a unique face template; this approach would break the system's core functionality. Option D is wrong because encryption of stored face templates protects data at rest but does not provide a lawful basis for processing; GDPR requires a valid legal ground (such as explicit consent) regardless of encryption.

60
Multi-Selectmedium

You are designing an Azure AI Vision solution that must detect and extract text from identity documents such as passports and driver's licenses. The solution must also identify the document type and extract key fields like name and date of birth. You need to choose the appropriate Azure AI service and features. (Choose two.)

Select 2 answers
A.Azure AI Vision Read API
B.Azure AI Document Intelligence with the prebuilt-idDocument model
C.Azure AI Vision Image Analysis with the read feature
D.Azure AI Document Intelligence with the prebuilt-receipt model
E.Azure AI Document Intelligence with a custom extraction model
AnswersB, E

The prebuilt-idDocument model in Azure AI Document Intelligence is specifically trained to analyze identity documents such as passports and driver's licenses. It extracts key fields like name, date of birth, document number, and expiration date, and it identifies the document type. This directly meets the requirement for detecting and extracting text and fields from identity documents.

Why this answer

The prebuilt-idDocument model in Azure AI Document Intelligence is specifically designed to analyze identity documents and extract fields like name and date of birth. A custom extraction model in the same service can also be trained to extract custom fields and classify document types when prebuilt models are insufficient. The Read API and Image Analysis read feature only extract text without structured field extraction, and the prebuilt-receipt model is for receipts, not identity documents.

Exam trap

The trap here is assuming that any OCR service can extract structured fields from identity documents, when only specialized Document Intelligence models provide that capability.

61
MCQmedium

A company uses Azure AI Vision to extract text from scanned invoices. They need to preserve the layout information, such as tables and key-value pairs, to automate data entry. Which Azure service should they use?

A.Azure AI Face API
B.Azure AI Custom Vision
C.Azure AI Vision Read API
D.Azure AI Document Intelligence (formerly Form Recognizer)
AnswerD

Document Intelligence is designed to extract text, tables, and key-value pairs from documents such as invoices. It provides prebuilt models for invoices that understand layout and can return structured data like invoice ID, date, and line items. This directly meets the need to preserve layout information and automate data entry, making it the correct choice for this scenario.

Why this answer

Document Intelligence provides prebuilt invoice models that extract text, tables, and key-value pairs while preserving layout. The Read API only returns text lines, and Custom Vision and Face API are for different domains. For automating data entry from invoices with tables, Document Intelligence is the appropriate service.

Exam trap

The trap here is assuming the Read API is sufficient for invoices, but it lacks layout understanding and structured field extraction that Document Intelligence provides.

62
MCQhard

You are creating a new Custom Vision project with the above JSON. The domainId corresponds to the 'Logo' domain. Which type of model will this project train?

A.An object detection model for logo detection
B.An optical character recognition model
C.A multilabel image classification model for logo detection
D.A general image classification model
AnswerC

The Logo domain is a multilabel classification domain, meaning each image can be assigned multiple logo tags simultaneously rather than one exclusive label. Training with that domainId therefore produces a multilabel image classification model for logo detection.

Why this answer

The 'Logo' domain in Custom Vision is specifically designed for image classification tasks, not object detection. When you create a project with the 'Logo' domain, it trains a multilabel image classification model, meaning each image can be assigned multiple labels (e.g., multiple logos in one image). This domain is optimized for identifying and classifying logos within images, making it distinct from object detection or general classification.

Exam trap

The trap here is that candidates often confuse the 'Logo' domain with object detection, assuming it draws bounding boxes around logos, when in fact it performs multilabel classification without localization.

How to eliminate wrong answers

Option A is wrong because the 'Logo' domain does not correspond to object detection; object detection requires a domain like 'General (Object Detection)' or 'Logo (Object Detection)' if available, and the JSON specifies the 'Logo' domain which is for classification. Option B is wrong because optical character recognition (OCR) is not a Custom Vision domain; OCR is handled by Azure Cognitive Services like Computer Vision's Read API, not Custom Vision. Option D is wrong because while the 'Logo' domain is a type of image classification, it is specifically a multilabel classification model, not a general image classification model (which typically uses single-label classification).

63
MCQmedium

You are building a computer vision solution to detect defects on a manufacturing assembly line. The solution must process images in real-time with low latency, and you need to choose an Azure service. Which service should you use?

A.Azure Computer Vision API
B.Azure Video Indexer
C.Azure Form Recognizer
D.Azure Custom Vision
AnswerD

Azure Custom Vision trains and hosts a purpose-built image classification or object detection model, letting you deploy a compact domain-specific model to a prediction endpoint for low-latency inline defect scoring, satisfying the real-time assembly-line constraint.

Why this answer

Azure Custom Vision is the correct choice because it allows you to train a custom image classification or object detection model tailored to detect specific manufacturing defects. It supports real-time, low-latency inference via a Docker container deployed to edge devices or directly through the prediction API, meeting the assembly line's performance requirements.

Exam trap

The trap here is that candidates often choose Azure Computer Vision API (Option A) because it sounds like a general-purpose vision service, but they overlook the requirement for custom defect detection, which necessitates a trainable model like Custom Vision.

How to eliminate wrong answers

Option A is wrong because Azure Computer Vision API provides pre-trained models for general image analysis (e.g., OCR, tagging) and cannot be customized to detect specific manufacturing defects without retraining. Option B is wrong because Azure Video Indexer is designed for analyzing video content (e.g., speech, faces, scenes) and is not optimized for real-time, low-latency image processing on a per-frame basis. Option C is wrong because Azure Form Recognizer is specialized for extracting text and structure from documents (e.g., invoices, forms), not for detecting visual defects in manufacturing images.

64
Multi-Selecthard

Which THREE factors should you consider when selecting a pricing tier for Azure Computer Vision in a production environment?

Select 3 answers
A.Availability of free tier
B.Type of storage account for images
C.Data residency requirements
D.Latency requirements
E.Transactions per second limit
AnswersC, D, E

May require specific region and tier.

Why this answer

Data residency requirements (Option C) are critical when selecting a pricing tier for Azure Computer Vision because the service processes images in specific regional data centers, and some tiers (e.g., Standard S0) support multi-region processing while others may be restricted. Compliance with regulations like GDPR or HIPAA may require that image data never leaves a particular geography, directly influencing which tier and region you can choose.

Exam trap

The trap here is that candidates often confuse the free tier's availability as a valid production option, or mistakenly think storage account type influences pricing tier selection, when in reality the key factors are operational constraints like TPS, latency, and data residency compliance.

65
MCQeasy

A company wants to extract key-value pairs from scanned invoices using Azure AI. Which service should they use?

A.Read API
B.Custom Vision
C.OCR API
D.Azure AI Document Intelligence
AnswerD

Azure AI Document Intelligence's prebuilt invoice model extracts key-value pairs such as invoice number, date, and total directly from scanned documents, satisfying the requirement for structured field extraction. Its OCR and layout analysis handle scanned images, unlike generic vision or language services that lack invoice-specific field mapping.

Why this answer

Azure AI Document Intelligence (formerly Form Recognizer) is the correct choice because it is specifically designed to extract key-value pairs, tables, and structured data from scanned documents like invoices. Unlike the Read API or OCR API, which only return raw text or OCR output, Document Intelligence uses prebuilt models (e.g., 'prebuilt-invoice') that understand the semantic layout of invoices, enabling direct extraction of fields such as invoice number, date, and total amount.

Exam trap

The trap here is that candidates often confuse the Read API or OCR API with Document Intelligence because all three involve text extraction, but only Document Intelligence provides key-value pair extraction and document understanding capabilities.

How to eliminate wrong answers

Option A is wrong because the Read API extracts printed and handwritten text as lines and words, but it does not parse key-value pairs or understand document structure. Option B is wrong because Custom Vision is an image classification and object detection service, not designed for text extraction or document understanding. Option C is wrong because the OCR API (part of Computer Vision) performs optical character recognition to return raw text and bounding boxes, but it lacks the ability to identify and extract key-value pairs or structured fields from invoices.

66
MCQhard

A manufacturing company uses Azure Custom Vision to detect defects on an assembly line. The model is deployed to a container on a local edge server. Recently, the model's accuracy dropped. You suspect data drift. What should you do to monitor and retrain the model?

A.Use Azure Machine Learning data drift monitoring on the Custom Vision endpoint.
B.Periodically collect new images with labels, retrain the model in Custom Vision, and redeploy the updated container.
C.Configure Custom Vision to send alerts when drift is detected.
D.Enable active learning in Custom Vision to automatically retrain the model.
AnswerB

Data drift is countered by capturing fresh labelled images from the line, retraining in Custom Vision so the model learns current defect patterns, then exporting and redeploying the updated container to the edge server. This restores accuracy while keeping inference local.

Why this answer

Custom Vision models deployed to containers on edge devices do not expose a REST endpoint that Azure Machine Learning's data drift monitoring can directly access. The only way to detect drift and retrain is to periodically collect new labeled images from the production line, retrain the model in Custom Vision, and redeploy the updated container to the edge server.

Exam trap

The trap here is that candidates assume Azure Machine Learning's data drift monitoring works with any deployed model, but it specifically requires an Azure-hosted endpoint, not a local container, and Custom Vision lacks native drift detection or auto-retraining features.

How to eliminate wrong answers

Option A is wrong because Azure Machine Learning data drift monitoring requires an Azure-hosted endpoint (e.g., AKS or ACI) with a scoring URI; Custom Vision containers on local edge servers do not provide such an endpoint, so drift monitoring cannot be configured. Option C is wrong because Custom Vision does not have built-in drift detection or alerting capabilities; it only provides training and prediction APIs, not monitoring. Option D is wrong because active learning in Custom Vision is a feature for image classification that suggests images for labeling to improve the model, but it does not automatically retrain the model or handle drift detection on edge deployments.

67
MCQhard

You are training an Azure Custom Vision object detection model to locate pallets in warehouse photos. Your training set contains 500 images, but only 40 images include pallets, while the rest are empty aisles. The model performs poorly, often missing pallets. You need to improve detection while keeping training time reasonable. What should you do first?

A.Enable the 'General' domain and retrain with the same dataset without changes.
B.Switch the project domain to a compact domain to speed up training.
C.Increase the probability threshold so fewer false positives are returned.
D.Add more labeled images that contain pallets in varied conditions and balance the dataset.
AnswerD

The model misses pallets because positive examples are scarce relative to empty aisles, so it learns the background class too strongly. Adding and labeling more pallet images across lighting, angles, and stacking arrangements gives the model representative positive features, improving recall. Balancing the dataset directly targets the root cause rather than masking it with threshold changes.

Why this answer

Object detection models learn from the distribution of labeled examples. When pallet images are only 8 percent of the set, the model is biased toward the dominant empty-aisle class and misses pallets. Adding more varied, labeled pallet images and balancing classes improves the positive signal.

Threshold tuning, domain changes, or simple retraining do not correct the underlying data imbalance.

Exam trap

The trap here is reaching for a threshold or domain tweak, when the real cause is too few labeled positive examples relative to empty background images.

68
MCQmedium

You need to build a solution that reads text from images in multiple languages, including Arabic and English, and translates the text into English. The solution must preserve the original layout as much as possible. Which combination of Azure AI services should you use?

A.Azure AI Document Intelligence Read and Azure AI Translator
B.Azure AI Document Intelligence Read and Azure AI Language
C.Azure AI Vision OCR and Azure AI Translator
D.Azure AI Speech and Azure AI Translator
AnswerA

Azure AI Document Intelligence's Read model extracts printed and handwritten text with layout preserved as lines and words, supporting Arabic and English. Azure AI Translator then converts the extracted text into English. This combination satisfies both the multilingual OCR requirement and the layout-preservation constraint, which standalone Translator or Vision OCR cannot fully meet.

Why this answer

Azure AI Document Intelligence Read (formerly Form Recognizer Read) is optimized for extracting text from images and documents while preserving the original layout, including bounding box coordinates for each text element. Azure AI Translator then translates the extracted text into English. This combination meets the requirement for multi-language OCR (including Arabic and English) and layout preservation.

Exam trap

The trap here is that candidates often confuse Azure AI Vision OCR (legacy) with Azure AI Document Intelligence Read, assuming both provide equivalent layout preservation, but only Document Intelligence Read is designed for structured layout-aware extraction.

How to eliminate wrong answers

Option B is wrong because Azure AI Language provides text analytics (e.g., sentiment, key phrases) but does not include OCR capabilities; it cannot read text from images. Option C is wrong because Azure AI Vision OCR (legacy OCR API) does not preserve layout information as effectively as Document Intelligence Read, which is specifically designed for layout-aware extraction. Option D is wrong because Azure AI Speech is for speech-to-text and text-to-speech, not for reading text from images.

69
MCQeasy

You are building a mobile app that uses Azure AI Vision to generate captions for photos taken by users. The app must work offline in areas with no internet connectivity. Which Azure AI Vision feature should you use?

A.Azure AI Vision Docker container for Image Analysis
B.Azure AI Custom Vision
C.Azure AI Vision Read API
D.Azure AI Vision Image Analysis API
AnswerA

Azure AI Vision provides Docker containers that can run the Image Analysis capabilities locally, including caption generation. These containers can be deployed on-premises or on edge devices, enabling offline operation. By using the container, the mobile app can communicate with a local endpoint without internet, satisfying the offline requirement.

Why this answer

The Azure AI Vision Docker container for Image Analysis allows you to run the captioning feature locally, enabling offline operation. This is the only option that provides the required functionality without internet connectivity, making it the correct choice for a mobile app that must work in areas with no network.

Exam trap

The trap here is overlooking that Docker containers can run AI services locally, while the cloud APIs require connectivity.

70
MCQeasy

You need to analyze a video stream from a security camera to count the number of people entering a building. Which Azure AI service is most suitable?

A.Azure AI Spatial Analysis
B.Azure AI Custom Vision
C.Azure AI Computer Vision
D.Azure AI Video Indexer
AnswerA

Azure AI Spatial Analysis ingests live video and applies computer vision operations, including people counting and zone-based entry/exit tracking, directly on the stream. It satisfies the real-time counting constraint that image-only services cannot, since Azure AI Vision handles still images rather than continuous camera feeds.

Why this answer

Azure AI Spatial Analysis is the most suitable service because it is purpose-built for real-time spatial analysis of video streams, including counting people entering a building. It can process live camera feeds and detect when people cross a line or enter a zone, providing real-time counts. In contrast, Azure AI Video Indexer is designed for analyzing recorded video files, not live streams, making it less suitable for real-time security camera analysis.

Other options like Custom Vision and Computer Vision are for image classification and image analysis, respectively, not optimized for video streams.

Exam trap

The key trap is that candidates often choose Azure AI Video Indexer (Option D) because it is associated with video analysis, but it is designed for indexing and analyzing video files, not real-time streaming. Azure AI Spatial Analysis (Option A) is the correct choice for live video streams and people counting scenarios.

How to eliminate wrong answers

Option A is wrong because Azure AI Spatial Analysis is a feature within Azure AI Computer Vision that focuses on real-time spatial relationships and occupancy analysis (e.g., social distancing, people counting in a space) but is not optimized for video stream analysis with event-based counting like entering a building; it is more suited for static spatial monitoring. Option B is wrong because Azure AI Custom Vision requires training a custom model with labeled images to detect specific objects, which is overkill and less efficient for a standard people-counting task that can be handled by pre-built video analytics. Option C is wrong because Azure AI Computer Vision provides image analysis APIs (e.g., OCR, object detection) but lacks native video stream processing capabilities and event-based counting for people entering a building; it is designed for single-image analysis, not continuous video streams.

71
MCQeasy

You need to detect if a photo contains adult or racy content. Which Azure AI Computer Vision feature should you use?

A.Describe Image API
B.OCR API
C.Analyze Image API with the 'adult' parameter
D.Tag Image API
AnswerC

The Analyze Image API accepts an adult parameter that returns adult and racy classification scores with confidence values. This directly satisfies the requirement to detect adult or racy content in a photo, unlike OCR or tagging features.

Why this answer

The Analyze Image API with the 'adult' parameter is the correct feature because it specifically detects adult, racy, and gory content in images. When you call the Analyze Image API and include the 'adult' visual feature, Azure AI Computer Vision returns a boolean flag and a confidence score for adult and racy content classification, enabling content moderation.

Exam trap

The trap here is that candidates often confuse the Tag Image API's generic object tagging with the specialized adult content detection feature, assuming tags like 'swimsuit' or 'underwear' would suffice, but only the Analyze Image API with the 'adult' parameter provides the explicit moderation scores required by the question.

How to eliminate wrong answers

Option A is wrong because the Describe Image API generates human-readable captions summarizing the image content, but it does not provide explicit adult/racy content detection or confidence scores. Option B is wrong because the OCR API extracts printed or handwritten text from images, and has no capability to analyze visual content for adult or racy themes. Option D is wrong because the Tag Image API returns a list of content tags (e.g., 'person', 'tree') based on objects and actions, but it does not include a dedicated adult/racy content moderation feature.

72
MCQhard

You are deploying a Custom Vision model to a production environment. The model must handle 100 predictions per second with low latency. Which deployment option should you choose?

A.Use the Free tier prediction endpoint.
B.Export the model as a Docker container and run it on Azure Container Instances.
C.Use the Training API to make predictions.
D.Use a paid tier prediction endpoint with sufficient capacity.
AnswerD

A paid tier prediction endpoint supplies dedicated throughput and lower latency than the free tier, whose strict rate limits cannot sustain 100 predictions per second. Scaling the endpoint's capacity directly satisfies the stem's high-volume, low-latency constraint, making it the appropriate deployment choice for production Custom Vision inference.

Why this answer

A paid tier prediction endpoint in Azure Custom Vision is designed to handle production-scale workloads with dedicated compute resources, supporting up to 100 predictions per second with low latency. The Free tier is rate-limited and cannot sustain this throughput, while exporting as a Docker container introduces additional overhead and scaling complexity that may not guarantee the required latency or throughput without manual orchestration.

Exam trap

The trap here is that candidates may assume exporting a model as a Docker container (Option B) is always the best for performance, but they overlook the operational overhead and lack of built-in scaling for high-throughput cloud predictions, whereas the paid endpoint is optimized for exactly this scenario.

How to eliminate wrong answers

Option A is wrong because the Free tier prediction endpoint is rate-limited to 20 predictions per minute and cannot handle 100 predictions per second. Option B is wrong because exporting the model as a Docker container and running it on Azure Container Instances requires manual scaling and does not provide built-in load balancing or guaranteed low latency for high-throughput production workloads; it is better suited for offline or edge scenarios. Option C is wrong because the Training API is used for training and managing models, not for making real-time predictions; using it for predictions would be inefficient and unsupported.

73
Multi-Selecthard

You are building a document processing solution that extracts information from invoices. The invoices come in various formats and languages. You need to extract line items, totals, and supplier names. Which THREE services should you combine?

Select 3 answers
A.Azure AI Custom Vision
B.Azure AI Translator
C.Azure AI Content Safety
D.Azure AI Document Intelligence
E.Azure AI Vision OCR
AnswersB, D, E

Translates text if invoices are in multiple languages.

Why this answer

Azure AI Translator is correct because invoices arrive in various languages, and translating extracted text to a common language (e.g., English) is necessary for downstream processing like entity extraction and validation. Without translation, multilingual invoice data would be inconsistent or unprocessable by language-specific models.

Exam trap

The trap here is that candidates may mistakenly choose Azure AI Custom Vision for 'extracting' invoice data, confusing its image classification capabilities with the structured document extraction provided by Document Intelligence.

74
MCQeasy

Refer to the exhibit. An Azure Cognitive Services Computer Vision API call for image captioning is returning only one caption. The developer wants to get three possible captions ranked by confidence. Which parameter should be modified in the request?

A.Use a different API version, such as 2023-04-01.
B.Modify the URL to point to a different image.
C.Change the language parameter to 'multi'.
D.Set the maxCandidates value to 3.
AnswerD

The maxCandidates parameter controls how many alternative captions the Image Captioning service returns, each with a confidence score. Setting it to 3 satisfies the requirement for three ranked captions; leaving it at the default of 1 explains why only one caption currently appears.

Why this answer

The `maxCandidates` parameter in the Computer Vision Image Analysis API controls the maximum number of captions returned in the response. By default, this value is 1, so only the top-ranked caption is returned. Setting `maxCandidates=3` instructs the API to return up to three captions, each with its own confidence score, ranked from highest to lowest confidence.

Exam trap

The trap here is that candidates may confuse the `maxCandidates` parameter with other parameters like `language` or `details`, or assume that changing the API version or image source would increase the number of captions, when in fact the default behavior is to return only one caption unless explicitly overridden.

How to eliminate wrong answers

Option A is wrong because changing the API version (e.g., to 2023-04-01) does not affect the number of captions returned; the `maxCandidates` parameter is available across supported versions. Option B is wrong because pointing to a different image changes the input but does not alter the request parameter that controls the number of captions; the API would still return only one caption per image unless `maxCandidates` is set. Option C is wrong because the `language` parameter specifies the language of the returned text (e.g., 'en' for English), not the count of captions; 'multi' is not a valid language value for this API.

75
MCQmedium

A retail company uses Azure Computer Vision to analyze customer traffic in stores. They deploy a custom object detection model to count customers and detect occupancy. After deployment, the model consistently underestimates the number of customers during peak hours. The company has retrained the model with more data but the issue persists. What is the most likely cause?

A.The model is not being batch-processed for inference.
B.The training data does not adequately represent peak-hour scenarios.
C.The model is overfitting to the training data.
D.The Computer Vision API version is outdated.
AnswerB

Persistent underestimation despite retraining indicates the training data under-represents peak-hour conditions such as crowding and occlusion, so the model never learns those patterns. The cause is a data representation gap, not model architecture or inference configuration.

Why this answer

The model consistently underestimates customer counts during peak hours, which indicates a distribution shift between the training data and the inference environment. Even after retraining with more data, the issue persists because the additional data likely still lacks sufficient representation of peak-hour scenarios (e.g., high density, occlusion, rapid movement). In Azure Custom Vision, object detection models learn from labeled examples; if the training set does not include diverse peak-hour images with varied lighting, crowd densities, and angles, the model will fail to generalize to those conditions.

Exam trap

The trap here is that candidates may assume retraining with 'more data' automatically fixes the issue, but the key is that the additional data must be representative of the specific failure scenario (peak hours), not just any data.

How to eliminate wrong answers

Option A is wrong because batch processing affects throughput and latency, not the accuracy of individual inference results; the model's underestimation is a precision/recall issue, not a processing mode issue. Option C is wrong because overfitting would cause the model to perform well on training data but poorly on new data in general, not specifically during peak hours; the consistent underestimation only in peak hours points to a data distribution mismatch, not overfitting. Option D is wrong because the Computer Vision API version affects available features and endpoints, not the learned weights of a custom object detection model; the model's behavior is determined by its training data and architecture, not the API version used for deployment.

Page 1 of 2 · 90 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Implement computer vision solutions questions.