Courseiva

AI-102 Implement computer vision solutions Practice Question

A developer needs to generate a descriptive caption for an image using Azure AI Vision. The image contains a dog catching a frisbee in a park. Which feature of Image Analysis should they use?

⚠ Common exam trap

Watch out — candidates often confuse tags with captions, as both provide descriptive information, but only captions produce a full sentence.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Caption

The caption feature of Azure AI Vision Image Analysis generates a natural language description of an image's content. It is the only feature that produces a sentence like 'a dog catching a frisbee in a park'. Tags, objects, and read serve different purposes: tags provide keywords, objects detect and locate items, and read extracts text. Therefore, caption is the correct choice.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Caption

    Why this is correct

    The caption feature in Image Analysis generates a human-readable sentence that describes the image content, such as 'a dog catching a frisbee in a park'. It is specifically designed to produce descriptive captions. This directly meets the developer's requirement for a descriptive caption.

  • ✗

    Read

    Why it's wrong here

    The read feature extracts text from images, which is irrelevant for an image of a dog catching a frisbee unless there is text present. It does not generate captions. Using read would not produce a description of the scene. Thus, it is not suitable for this scenario.

  • ✗

    Objects

    Why it's wrong here

    The objects feature detects physical objects in an image and returns their locations and labels, but it does not generate a descriptive sentence. It would identify 'dog' and 'frisbee' with bounding boxes, but not describe the action. This does not provide the required caption.

  • ✗

    Tags

    Why it's wrong here

    Tags provide a list of relevant keywords for an image, such as 'dog', 'frisbee', and 'park', but they do not generate a grammatical sentence describing the scene. The requirement is for a descriptive caption, which is a natural language sentence, not just tags. Therefore, tags do not fulfill the need.

About these practice questions

One of 761 original AI-102 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.