AI-102 Implement computer vision solutions Practice Question
A developer is building a mobile app that uses Azure AI Vision to generate a descriptive caption for user-uploaded photos. The app must return a human-readable sentence describing the main content of each image. Which Image Analysis feature should the developer use?
⚠ Common exam trap
The trap here is assuming that tags or objects can be concatenated into a sentence, but only the Caption feature is designed to output a fluent description.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Caption
The Caption feature of Image Analysis is purpose-built to generate a one-sentence description of an image's main content. Tags, Objects, and Read serve different purposes: tags provide keywords, objects give bounding boxes, and Read extracts text. Only Caption produces the natural language sentence needed for the mobile app.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Tags
Why it's wrong here
Tags return a list of content tags with confidence scores, such as 'outdoor', 'tree', or 'person'. While tags can indicate what is in the image, they do not form a grammatical sentence or describe relationships between objects. The requirement is for a human-readable sentence, so tags alone are insufficient. Tags are better for filtering or categorization, not for generating captions.
- ✗
Objects
Why it's wrong here
Objects detection returns bounding boxes for specific objects like cars, people, or furniture, along with their labels. It identifies where objects are located but does not generate a natural language description. While useful for spatial analysis, it does not provide the human-readable sentence that the app requires. It is a different feature from captioning.
- ✗
Read
Why it's wrong here
The Read feature extracts printed or handwritten text from images, returning words and lines. It is used for OCR tasks, not for describing the overall scene. If the photo contains a sign with text, Read would extract that text, but it would not describe the main content or produce a caption. Therefore, it does not satisfy the requirement.
- ✓
Caption
Why this is correct
The Caption feature in Image Analysis generates a single, human-readable sentence that describes the main content of an image, such as 'a person riding a bike on a beach'. It is specifically designed for this purpose and returns a confidence score for the caption. This directly meets the requirement of producing a descriptive sentence for each photo.
Go deeper
Related to this question
About these practice questions
One of 761 original AI-102 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.