AI-102 Plan and manage an Azure AI solution Practice Question
You are developing an Azure AI solution that uses pre-built models from Azure AI Vision to analyze images. The solution must be able to detect objects and read printed text. Which TWO capabilities should you use?
⚠ Common exam trap
Watch out — candidates often confuse Image Tagging (which only provides labels) with Object Detection (which provides both labels and spatial localization), and may mistakenly choose the legacy OCR API instead of the modern Read API for text extraction.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Object detection
Object detection (C) is correct because it is the Azure AI Vision pre-built capability that locates and classifies multiple objects within an image, returning bounding boxes and labels, which directly satisfies the requirement to detect objects. Read (OCR) (E) is correct because it is the modern Azure AI Vision OCR engine that extracts printed and handwritten text from images and documents, satisfying the requirement to read printed text. The legacy OCR (A) option is an older, deprecated recognition model with weaker accuracy and limited language support, so it is not the recommended choice for new solutions. Facial detection (B) only returns face locations and attributes and does not detect general objects or read text, and image tagging (D) produces descriptive labels for the overall image rather than object locations or extracted text, so neither meets the stated requirements.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
OCR (legacy)
Why it's wrong here
Legacy OCR supports only synchronous recognition of small, upright printed text and cannot detect objects. It is tempting because it does read printed text, but the Read API supersedes it for printed text, and object detection is a separate capability entirely.
- ✗
Facial detection
Why it's wrong here
Facial detection locates human faces and estimates attributes such as age or emotion; it neither detects general objects nor reads printed text. It is tempting because it returns bounding boxes, but those boxes cover faces only, so it cannot satisfy either stated requirement.
- ✓
Object detection
Why this is correct
Object detection returns bounding boxes and labels for multiple objects within an image, directly satisfying the requirement to detect objects. Azure AI Vision's pre-built model provides this without training, meeting the scenario's use of standard capabilities.
- ✗
Image tagging
Why it's wrong here
Image tagging returns descriptive labels about overall scene content, not bounding boxes for individual objects nor text transcription. It is tempting because tags do describe image contents, but object detection supplies the coordinates and Read OCR supplies printed text, which the scenario requires.
- ✓
Read (OCR)
Why this is correct
The Read capability performs optical character recognition on printed and handwritten text in images, satisfying the requirement to read printed text. It is a pre-built Azure AI Vision model, so no custom training is needed.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AI-102 question from scratch — 761 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.