AI-102 Implement computer vision solutions Practice Question
A developer is using the Azure AI Vision Image Analysis API to extract text from a photo of a street sign. The sign contains text in both English and Japanese arranged in multiple columns. The developer needs the response to include the detected language and the bounding box for each text line. Which feature should they use?
⚠ Common exam trap
Many candidates confuse general image analysis features like captioning or tagging with OCR, when only the read feature extracts text and provides bounding boxes and language detection.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The 'read' feature in Image Analysis
The read feature in Azure AI Vision Image Analysis is specifically designed for OCR. It returns extracted text with bounding boxes for lines and words, and it detects the language of the text. This matches the need to capture English and Japanese text in multiple columns and provide bounding boxes for each line. Other features like caption, detectObjects, and tags do not perform OCR.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The 'read' feature in Image Analysis
Why this is correct
The read feature in Image Analysis performs OCR and returns extracted text along with bounding boxes for lines and words. It also detects the language of the text. This directly provides the language and bounding box information required, and it handles mixed-language and multi-column text in a single call.
- ✗
The 'detectObjects' feature in Image Analysis
Why it's wrong here
The detectObjects feature identifies common objects and returns their bounding boxes, but it does not perform OCR or extract text. It might detect the sign as an object, yet it would not read the text or provide language information. This makes it unsuitable for extracting text from the sign.
- ✗
The 'caption' feature in Image Analysis
Why it's wrong here
The caption feature generates a descriptive sentence about the image content but does not extract text or return bounding boxes. It would describe the sign generally, not read the text. Therefore it cannot provide the language or bounding box details needed for the street sign.
- ✗
The 'tags' feature in Image Analysis
Why it's wrong here
The tags feature returns content tags for the image, such as 'sign' or 'outdoor', but it does not extract text or bounding boxes. It is useful for categorization, not for OCR. Using it would not yield the text, language, or bounding box data required.
Go deeper
Related to this question
About these practice questions
One of 761 original AI-102 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Microsoft exam blueprint
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.