AI-102 Implement computer vision solutions Practice Question
Which THREE are valid uses of Azure AI Vision Image Analysis 4.0? (Select three.)
⚠ Common exam trap
Test-takers frequently confuse the capabilities of Azure AI Vision with those of Azure AI Speech or Azure AI Translator, assuming Image Analysis can handle audio transcription or text translation when it strictly processes visual content only.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Extract printed text from an image using OCR
Option A is correct because Image Analysis 4.0 includes a Read OCR feature that extracts printed and handwritten text from images and documents, returning lines and words. Option C is correct because the object detection capability identifies objects in an image and returns their coordinates as bounding boxes along with confidence scores. Option D is correct because Image Analysis 4.0 can generate captions (and dense captions) describing the content of an image in natural language. Option B is not valid because audio transcription is a speech service capability (e.g., Azure AI Speech), not Image Analysis. Option E is not valid because translating text is performed by Azure AI Translator; Image Analysis only extracts the text via OCR, it does not translate it.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Extract printed text from an image using OCR
Why this is correct
Image Analysis 4.0 includes a Read OCR feature that extracts printed and handwritten text from images via the Read API, satisfying the OCR use case. It returns text lines and words with bounding polygons, so printed text extraction is a supported capability of this service.
- ✗
Transcribe spoken audio from a video file
Why it's wrong here
Image Analysis 4.0 processes still images for captions, tags, objects and OCR; it does not transcribe audio, which needs Azure AI Speech. It tempts because video frames can be sampled as images, but speech-to-text from the audio track is a Speech service capability.
- ✓
Detect objects in an image and return bounding boxes
Why this is correct
Image Analysis 4.0 provides object detection, returning bounding boxes with labels and confidence scores for recognised objects. This directly satisfies the requirement to detect objects and return their coordinates, distinguishing it from pure classification or captioning features.
- ✓
Generate a human-readable caption for an image
Why this is correct
Image Analysis 4.0 includes a captioning feature that returns a human-readable description of an image's content, satisfying the stem's requirement for a valid capability. It uses the same multimodal model as the Florence-based vision services, producing natural-language sentences rather than tags alone.
- ✗
Translate text found in an image to another language
Why it's wrong here
Image Analysis 4.0 returns OCR text via the Read feature but performs no translation; that requires Azure AI Translator. It tempts because OCR output is commonly piped into translation, yet the translation step sits in a separate Azure AI service, not Image Analysis.
Go deeper
Related to this question
About these practice questions
One of 761 original AI-102 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.