AI-102 Implement computer vision solutions Practice Question
A company wants to build a solution that automatically generates alt text for images on their website to improve accessibility. The alt text must be a concise, human-readable description of the image content. Which Azure AI Vision feature should they use?
⚠ Common exam trap
Many candidates confuse Tags with Caption; Tags give keywords, while Caption gives a full sentence suitable for alt text.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Image Analysis with the Caption feature
The Caption feature in Image Analysis is designed to generate a one-sentence description of an image, which is ideal for alt text. It is prebuilt and requires no training. Other features like Tags or Detect Objects provide keywords or bounding boxes but not a coherent description, making them less suitable for accessibility purposes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Custom Vision with a classification model trained on descriptive text
Why it's wrong here
Custom Vision classification models output class labels, not descriptive sentences. Training a model to generate alt text would require a complex dataset and would not produce the fluent descriptions needed. This approach is overkill and does not directly solve the problem.
- ✗
Image Analysis with the Detect Objects feature
Why it's wrong here
Detect Objects returns bounding boxes and labels for objects in the image, but not a descriptive sentence. It is useful for identifying items but does not generate alt text. Using it would require significant post-processing to create a readable description, which is not ideal for the scenario.
- ✓
Image Analysis with the Caption feature
Why this is correct
The Caption feature generates a concise, human-readable description of an image, which is exactly what is needed for alt text. It is a prebuilt model that requires no training and returns a single sentence describing the image. This directly meets the accessibility requirement.
- ✗
Image Analysis with the Tags feature
Why it's wrong here
The Tags feature returns a list of keywords associated with the image, not a coherent sentence. While tags could be used to construct alt text, they are not human-readable descriptions and would require additional processing. This does not provide the concise description required for accessibility.
Go deeper
Related to this question
About these practice questions
This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.