Courseiva

AI-102 Implement computer vision solutions Practice Question

You are designing a solution that uses Azure AI Vision to analyze images uploaded by users. You need to extract text and also generate a descriptive caption for each image. You want to use the Image Analysis API. Which two capabilities should you enable? (Choose two.)

⚠ Common exam trap

Test-takers frequently confuse Tags with Caption; Tags provide keywords, whereas Caption produces a full descriptive sentence, which is what the scenario requires.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Read

The Read capability extracts text, and the Caption capability generates a descriptive sentence for the image. Together they satisfy both requirements. Tags, Objects, and Brands provide different types of analysis that are not needed here. Enabling Read and Caption allows a single API call to return both OCR results and a caption, streamlining the solution.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Read

    Why this is correct

    The Read capability in Image Analysis extracts printed and handwritten text from images and returns it as structured lines and words. This directly satisfies the requirement to extract text from user-uploaded images. It is the correct choice because it provides OCR functionality within the same API call, avoiding the need to call a separate OCR endpoint.

  • ✗

    Tags

    Why it's wrong here

    Tags return a list of content labels such as 'outdoor', 'tree', or 'person'. While useful for categorization, tags do not provide a full descriptive caption or extract text. The scenario specifically asks for a descriptive caption and text extraction, so tags alone would not fulfill either requirement. They could be used in addition, but they are not the correct selections here.

  • ✗

    Brands

    Why it's wrong here

    Brands detection identifies commercial brands or logos in an image. This is a specialized feature that does not extract text or produce a general caption. The scenario does not mention brand recognition, so this capability is irrelevant. Including it would not meet the requirements and could increase cost without providing value.

  • ✓

    Caption

    Why this is correct

    The Caption capability generates a human-readable description of the image content, such as 'a person riding a bike'. This meets the requirement to generate a descriptive caption for each image. It is a correct choice because it provides a concise natural language summary that can be displayed to users or stored for later use.

  • ✗

    Objects

    Why it's wrong here

    Objects detection identifies and localizes objects within an image, returning bounding boxes and labels. This is different from generating a descriptive caption or extracting text. The scenario does not require object localization; it requires text extraction and a caption. Therefore, enabling Objects would not satisfy the stated needs and would add unnecessary processing.

About these practice questions

Courseiva writes every AI-102 question from scratch — 761 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Microsoft exam blueprint

This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.