AI-102 Implement computer vision solutions Practice Question
You are designing a solution that uses Azure AI Vision to analyze images uploaded by users. You need to extract text and also generate a descriptive caption for each image. You want to use the Image Analysis API. Which two capabilities should you enable? (Choose two.)
⚠ Common exam trap
Test-takers frequently confuse Tags with Caption; Tags provide keywords, whereas Caption produces a full descriptive sentence, which is what the scenario requires.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Read
The Read capability extracts text, and the Caption capability generates a descriptive sentence for the image. Together they satisfy both requirements. Tags, Objects, and Brands provide different types of analysis that are not needed here. Enabling Read and Caption allows a single API call to return both OCR results and a caption, streamlining the solution.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Read
Why this is correct
The Read capability in Image Analysis extracts printed and handwritten text from images and returns it as structured lines and words. This directly satisfies the requirement to extract text from user-uploaded images. It is the correct choice because it provides OCR functionality within the same API call, avoiding the need to call a separate OCR endpoint.
- ✗
Tags
Why it's wrong here
Tags return a list of content labels such as 'outdoor', 'tree', or 'person'. While useful for categorization, tags do not provide a full descriptive caption or extract text. The scenario specifically asks for a descriptive caption and text extraction, so tags alone would not fulfill either requirement. They could be used in addition, but they are not the correct selections here.
- ✗
Brands
Why it's wrong here
Brands detection identifies commercial brands or logos in an image. This is a specialized feature that does not extract text or produce a general caption. The scenario does not mention brand recognition, so this capability is irrelevant. Including it would not meet the requirements and could increase cost without providing value.
- ✓
Caption
Why this is correct
The Caption capability generates a human-readable description of the image content, such as 'a person riding a bike'. This meets the requirement to generate a descriptive caption for each image. It is a correct choice because it provides a concise natural language summary that can be displayed to users or stored for later use.
- ✗
Objects
Why it's wrong here
Objects detection identifies and localizes objects within an image, returning bounding boxes and labels. This is different from generating a descriptive caption or extracting text. The scenario does not require object localization; it requires text extraction and a caption. Therefore, enabling Objects would not satisfy the stated needs and would add unnecessary processing.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AI-102 question from scratch — 761 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.