AI-102 Azure AI Vision OCR Read API Practice Question
A logistics company uses Azure AI Vision to analyze images of packages on conveyor belts. They need to detect damaged packages and read tracking numbers. The solution must process high throughput (1000 images per minute) with low latency (<500ms per image). The images are captured by fixed cameras. Which approach should you recommend?
⚠ Common exam trap
The trap is believing that Custom Vision can perform OCR natively. Custom Vision only handles image classification and object detection; text reading requires a dedicated OCR service. Candidates may think a single model is simpler, but it's not technically possible.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Azure AI Vision OCR Read API for tracking numbers and a separate Custom Vision model for damage detection
It combines Azure AI Vision OCR Read API (for text extraction) with a separate Custom Vision model (for damage detection). This leverages the specialized strengths of each service: the OCR Read API is optimized for text reading with high accuracy, while Custom Vision excels at object detection. The latency requirement (<500ms per image) can be met through parallel processing or edge deployment, and throughput can be achieved by scaling API calls. Option B is incorrect because Custom Vision does not include built-in OCR capability; it cannot read tracking numbers. Options A and C are less suitable: Document Intelligence is designed for structured documents, not real-time package analysis, and Video Indexer is for video streams, not still images.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Azure AI Document Intelligence to process package labels
Why it's wrong here
Document Intelligence extracts text and structure from documents such as labels and forms; it does not detect physical damage in package imagery. It would be the correct choice for parsing scanned shipping documents, not for camera-based damage detection.
- ✗
Train a single Custom Vision model that detects damage and reads tracking numbers using OCR
Why it's wrong here
Custom Vision performs classification or object detection on images and cannot read tracking numbers; OCR is a separate Azure AI Vision capability, so one model cannot do both. Custom Vision would be correct for damage classification alone, without the text-reading requirement.
- ✗
Use Azure AI Video Indexer to analyze the video stream from cameras
Why it's wrong here
Video Indexer analyses stored or streamed video for insights such as transcription and scene detection, not per-frame package inspection at 1000 images per minute under 500ms. It would be the right service for indexing recorded footage, not real-time conveyor inspection.
- ✓
Use Azure AI Vision OCR Read API for tracking numbers and a separate Custom Vision model for damage detection
Why this is correct
This approach splits the tasks: the OCR Read API reads tracking numbers, and a Custom Vision model detects damage. Both can run in parallel or be deployed at the edge to meet latency and throughput. This is the only technically feasible option.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AI-102 question from scratch — 761 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-102 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-102 exam.