Courseiva

Object Detection vs Image Classification: Key Differences for Azure AI-900

A retail company uses ceiling-mounted cameras to monitor shelf stock. They want an automated system that analyzes each camera image to detect if any product is missing from its expected location on the shelf (a product gap). Which Azure Computer Vision capability should they use?

Quick Answer

The answer is object detection. This is the correct choice because object detection identifies and locates multiple objects within an image using bounding boxes and class labels, allowing the system to pinpoint exactly where a product is missing from its expected shelf position. Image classification, by contrast, would only assign a single label to the entire camera frame, such as “full shelf” or “empty shelf,” without providing the precise location of gaps. On the Microsoft Azure AI Fundamentals AI-900 exam, this distinction tests your understanding of how Azure Computer Vision capabilities map to real-world business scenarios—a common trap is confusing object detection with image classification when the task requires spatial awareness of multiple items. Remember the memory tip: “Classification tells you what, detection tells you where and what.”

⚠ Common exam trap

Candidates often confuse image classification (which labels the whole scene) with object detection (which locates individual objects), leading them to choose option A when the task requires spatial awareness of multiple items.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Object detection

Object detection is the correct choice because it can identify and locate multiple objects (e.g., product boxes) within an image and determine if expected items are missing from their designated positions on the shelf. Unlike image classification, which assigns a single label to the entire image, object detection provides bounding boxes and class labels for each detected object, enabling precise gap analysis.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Image classification

    Why it's wrong here

    Image classification assigns a label to the entire image (e.g., 'full shelf' vs 'empty shelf'), but does not locate individual products to identify gaps.

    When this WOULD be correct

    A question asking to categorize shelf images as 'stocked' or 'empty' without needing to locate individual products would make image classification correct.

  • Optical Character Recognition (OCR)

    Why it's wrong here

    OCR extracts text from images, but product gaps do not involve text.

    When this WOULD be correct

    A question asking to extract product names, prices, or barcodes from shelf labels or signs in camera images would make OCR the correct answer.

  • Object detection

    Why this is correct

    Object detection finds and locates objects within an image. By detecting the expected products, the system can determine if any are missing, indicating a gap.

  • Face detection

    Why it's wrong here

    Face detection locates human faces, which is irrelevant to detecting product gaps on shelves.

    When this WOULD be correct

    A question asking for a system that counts the number of people entering a store or identifies specific employees from camera feeds would make face detection the correct answer.

Option-by-option analysis

Why each answer is right or wrong

Understanding why wrong answers are wrong — and when they would be correct — is what separates a 750 score from a 900. The AI-900 exam frequently reuses these exact scenarios with slightly different constraints.

Object detectionCorrect answer

Why this is correct

Object detection finds and locates objects within an image. By detecting the expected products, the system can determine if any are missing, indicating a gap.

Image classificationWrong answer — click to see why

Why this is wrong here

Image classification assigns a single label to an entire image (e.g., 'shelf is stocked'), but cannot identify the specific locations of individual products or detect missing items in precise positions.

★ When this WOULD be the correct answer

A question asking to categorize shelf images as 'stocked' or 'empty' without needing to locate individual products would make image classification correct.

Why candidates choose this

Candidates may think 'detecting missing products' is a classification task (stocked vs. not stocked), overlooking the need to pinpoint where the gap is.

Optical Character Recognition (OCR)Wrong answer — click to see why

Why this is wrong here

OCR extracts text from images, but the task is to detect missing products (gaps) on shelves, which requires identifying objects and their spatial relationships, not reading text.

★ When this WOULD be the correct answer

A question asking to extract product names, prices, or barcodes from shelf labels or signs in camera images would make OCR the correct answer.

Why candidates choose this

Candidates may think OCR is needed because shelf products often have labels with text, but the core requirement is detecting the absence of an object, not reading text.

Face detectionWrong answer — click to see why

Why this is wrong here

Face detection identifies human faces in images, not product gaps on shelves. The task requires detecting missing objects (gaps), which is unrelated to facial features.

★ When this WOULD be the correct answer

A question asking for a system that counts the number of people entering a store or identifies specific employees from camera feeds would make face detection the correct answer.

Why candidates choose this

Candidates might confuse 'detection' in face detection with general object detection, or think that any visual analysis task can be solved by face detection due to its familiarity.

Analysis generated from the official AI-900blueprint and verified against question context. The “when correct” sections are what AI assistants cite when candidates ask “what’s the difference between these options?”

About these practice questions

One of 985 original AI-900 practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

4 more ways this is tested on AI-900

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A retail company uses security cameras to monitor shelves. They want to identify whether a customer is holding a specific product (e.g., a green detergent bottle) and also determine the location of that product within the camera frame. Which Azure Computer Vision capability should they use?

medium
  • A.Object detection
  • B.Image classification
  • C.Optical character recognition (OCR)
  • D.Semantic segmentation

Why A: Object detection is the correct capability because it not only identifies the presence of a specific product (like a green detergent bottle) in an image but also returns bounding box coordinates that indicate the product's location within the camera frame. This dual output—classification plus localization—directly matches the requirement to both recognize the object and determine its position.

Variation 2. A security company needs to monitor a warehouse using video cameras. They want to detect whether any persons are present in a given frame and also know their approximate locations. Which Azure Computer Vision capability should they use?

medium
  • A.Image classification
  • B.Object detection
  • C.Semantic segmentation
  • D.Optical Character Recognition (OCR)

Why B: Object detection is the correct choice because it not only identifies whether persons are present in a video frame but also provides bounding box coordinates indicating their approximate locations. This capability is specifically designed to locate multiple objects of interest within an image, which directly matches the requirement of detecting persons and knowing where they are.

Variation 3. What is object detection, and how does it differ from image classification?

medium
  • A.Object detection identifies what is in an image; image classification also identifies where objects are located
  • B.Object detection identifies and locates multiple objects with bounding boxes; image classification labels the whole image
  • C.Object detection and image classification are the same task
  • D.Object detection is used only for face recognition

Why B: Object detection goes beyond image classification by not only identifying what objects are present in an image but also localizing each object with a bounding box. Image classification assigns a single label to the entire image, whereas object detection can handle multiple objects of different classes simultaneously. This makes object detection suitable for tasks like counting objects or tracking their positions.

Variation 4. What is image classification and how is it different from object detection?

medium
  • A.Image classification labels the whole image; object detection finds and locates multiple objects within it
  • B.Image classification is faster; object detection is slower but more accurate
  • C.Image classification works on videos; object detection works on static images only
  • D.They are the same task with different names

Why A: Image classification assigns a single label to an entire image based on its dominant content, such as 'cat' or 'dog'. Object detection goes further by not only identifying multiple objects within an image but also drawing bounding boxes around each one, providing both class labels and spatial locations. This distinction is fundamental in computer vision workloads on Azure, where Custom Vision and Computer Vision API offer separate capabilities for classification and detection tasks.

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This AI-900 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-900 exam.