Courseiva

AI-102 Implement computer vision solutions Practice Question

A company wants to build a solution that automatically generates alt text for images on their website to improve accessibility. The alt text must be a concise, human-readable description of the image content. Which Azure AI Vision feature should they use?

⚠ Common exam trap

Many candidates confuse Tags with Caption; Tags give keywords, while Caption gives a full sentence suitable for alt text.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Image Analysis with the Caption feature

The Caption feature in Image Analysis is designed to generate a one-sentence description of an image, which is ideal for alt text. It is prebuilt and requires no training. Other features like Tags or Detect Objects provide keywords or bounding boxes but not a coherent description, making them less suitable for accessibility purposes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Custom Vision with a classification model trained on descriptive text

    Why it's wrong here

    Custom Vision classification models output class labels, not descriptive sentences. Training a model to generate alt text would require a complex dataset and would not produce the fluent descriptions needed. This approach is overkill and does not directly solve the problem.

  • ✗

    Image Analysis with the Detect Objects feature

    Why it's wrong here

    Detect Objects returns bounding boxes and labels for objects in the image, but not a descriptive sentence. It is useful for identifying items but does not generate alt text. Using it would require significant post-processing to create a readable description, which is not ideal for the scenario.

  • ✓

    Image Analysis with the Caption feature

    Why this is correct

    The Caption feature generates a concise, human-readable description of an image, which is exactly what is needed for alt text. It is a prebuilt model that requires no training and returns a single sentence describing the image. This directly meets the accessibility requirement.

  • ✗

    Image Analysis with the Tags feature

    Why it's wrong here

    The Tags feature returns a list of keywords associated with the image, not a coherent sentence. While tags could be used to construct alt text, they are not human-readable descriptions and would require additional processing. This does not provide the concise description required for accessibility.

About these practice questions

This AI-102 question is part of Courseiva's 761-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.