AI-102 Implement computer vision solutions Practice Question
A developer needs to generate a descriptive caption for an image using Azure AI Vision. The image contains a dog catching a frisbee in a park. Which feature of Image Analysis should they use?
⚠ Common exam trap
Watch out — candidates often confuse tags with captions, as both provide descriptive information, but only captions produce a full sentence.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Caption
The caption feature of Azure AI Vision Image Analysis generates a natural language description of an image's content. It is the only feature that produces a sentence like 'a dog catching a frisbee in a park'. Tags, objects, and read serve different purposes: tags provide keywords, objects detect and locate items, and read extracts text. Therefore, caption is the correct choice.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Caption
Why this is correct
The caption feature in Image Analysis generates a human-readable sentence that describes the image content, such as 'a dog catching a frisbee in a park'. It is specifically designed to produce descriptive captions. This directly meets the developer's requirement for a descriptive caption.
- ✗
Read
Why it's wrong here
The read feature extracts text from images, which is irrelevant for an image of a dog catching a frisbee unless there is text present. It does not generate captions. Using read would not produce a description of the scene. Thus, it is not suitable for this scenario.
- ✗
Objects
Why it's wrong here
The objects feature detects physical objects in an image and returns their locations and labels, but it does not generate a descriptive sentence. It would identify 'dog' and 'frisbee' with bounding boxes, but not describe the action. This does not provide the required caption.
- ✗
Tags
Why it's wrong here
Tags provide a list of relevant keywords for an image, such as 'dog', 'frisbee', and 'park', but they do not generate a grammatical sentence describing the scene. The requirement is for a descriptive caption, which is a natural language sentence, not just tags. Therefore, tags do not fulfill the need.
Go deeper
Related to this question
About these practice questions
One of 761 original AI-102 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.