AI-900 Practice Question: Describe features of computer vision workloads on Azure
What is 'Azure AI Vision's image analysis v4.0' and what new capability does it add?
⚠ Common exam trap
Candidates often confuse 'version 4.0' with a simple incremental update (like resolution or performance tweaks) rather than recognizing it as a paradigm shift powered by the Florence foundation model, which is the core new capability tested.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Florence-powered advanced capabilities including dense captioning, embeddings, and improved background removal
Azure AI Vision's image analysis v4.0 is a major update that leverages the Florence foundation model to deliver advanced capabilities such as dense captioning (generating detailed descriptions for multiple regions in an image), image embeddings (vector representations for similarity search), and improved background removal. This version significantly enhances the depth and accuracy of image understanding compared to previous versions.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A version supporting 4K resolution images for the first time
Why it's wrong here
Pixel resolution is a low-level image-processing constraint, and previous Azure AI Vision versions already accepted high-resolution images for tasks like OCR and tagging. v4.0's headline improvements are semantic rather than pixel-based: Florence introduces dense captioning that describes individual image regions, image embeddings for vector search, and more precise background removal. Claiming '4K for the first time' confuses camera capture hardware specifications with cloud API functional enhancements, so it does not explain what the API version actually delivers.
- ✓
Florence-powered advanced capabilities including dense captioning, embeddings, and improved background removal
Why this is correct
v4.0 is powered by Microsoft's Florence model, a large vision-language transformer pretrained on billions of image-text pairs to learn joint visual-linguistic representations. Its dense captioning produces natural language descriptions for multiple salient regions in an image rather than a single whole-image caption, while embeddings map images and text to a common vector space for semantic search, and improved background removal uses fine-grained segmentation to isolate foreground objects. These Florence-powered capabilities are the defining advances that distinguish v4.0 from older Azure AI Vision models.
- ✗
A version requiring 4x more compute than the previous version
Why it's wrong here
No Microsoft specification states that v4.0 requires four times the compute of its predecessor, and compute consumption is an operational/deployment metric rather than a customer-facing capability. Florence is a transformer that can be fine-tuned and served efficiently, and the API abstracts away GPU allocation so users cannot (and need not) pick an API version based on a compute multiplier. This option misidentifies the version's value proposition by focusing on presumed resource requirements instead of the functional AI improvements introduced in v4.0.
- ✗
The fourth iteration of Microsoft's Kinect 3D depth sensor SDK
Why it's wrong here
The Kinect SDK was Microsoft's runtime and toolchain for the Kinect depth-sensing camera, designed for skeletal tracking, gesture recognition, and 3D scene capture on Xbox and Windows. Azure AI Vision v4.0, by contrast, is a cloud-based REST API built on the Florence vision-language foundation model, with no lineage to consumer gaming peripherals. The 'v4.0' version number refers to the Azure AI Vision service API iteration, not a successor to Kinect SDK releases.
Go deeper
Related to this question
Learn chapter
Azure Machine Learning Studio
Key term
Azure AI Vision
Azure AI Vision is a cloud-based service from Microsoft that uses pre-built machine learning models to extract information from images and videos, such as objects, text, faces, and scene descriptions.
Key term
Model
In IT and AI, a model is a trained mathematical representation that learns patterns from data to make predictions or decisions.
About these practice questions
This AI-900 question is part of Courseiva's 985-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-900 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-900 exam.