AI-900 Practice Question: Describe features of generative AI workloads on Azure
What is 'Azure AI Foundry's model benchmarks' and how do they help you choose a model?
⚠ Common exam trap
Many candidates confuse operational metrics (SLA, pricing) or UI performance with the actual AI task performance benchmarks, which are specifically designed to compare model capabilities on reasoning, code, and math tasks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Standardised AI task performance comparisons (reasoning, code, math) across models in the catalogue
Azure AI Foundry's model benchmarks provide standardized performance comparisons across models in the catalog, evaluating key AI tasks such as reasoning, code generation, and math. These benchmarks allow you to objectively compare models based on their performance on specific tasks, helping you select the most suitable model for your workload.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Performance tests for Azure AI Foundry's web interface loading speed
Why it's wrong here
Performance tests for the Azure AI Foundry web interface measure front-end responsiveness, rendering time, and network latency—operational UX concerns—not model intelligence. The catalogue benchmarks use static, well-known evaluation datasets and compute metrics like accuracy or pass rate on model outputs, independent of browser or portal behaviour. Evaluating UI speed would tell you nothing about a model's reasoning or coding ability.
- ✓
Standardised AI task performance comparisons (reasoning, code, math) across models in the catalogue
Why this is correct
These benchmarks are standardised evaluation scores—such as MMLU for broad reasoning, HumanEval for code generation, and GSM8K for math problem solving—precomputed on public datasets and shown in the catalogue to compare models objectively. They let you select a model based on demonstrated task competence without running your own costly evaluation harness from scratch, directly reflecting the model's underlying capability.
- ✗
Azure's SLA guarantees for model availability and API response time
Why it's wrong here
Azure SLA commitments define uptime percentages, standard API status codes, and first-byte latency thresholds—infrastructure reliability guarantees—while benchmarks score a model's output quality on cognitive tasks. An SLA might state 99.9% monthly availability but says nothing about whether a model has high accuracy on a benchmark like MMLU. These are separate service dimensions; uptime cannot substitute for capability measurement.
- ✗
Pricing benchmarks comparing Azure OpenAI costs against competitor services
Why it's wrong here
Pricing benchmarks focus on Azure OpenAI per-token cost, provisioned throughput, and reserved capacity—commercial factors used in procurement—whereas model benchmarks are standardised ML evaluation scores that measure task-specific capability such as reasoning or code generation. Comparing costs against competitors may inform budget decisions but does not indicate which model performs better on a given workload, so it is not what the catalogue's benchmark feature displays.
Go deeper
Related to this question
About these practice questions
Courseiva writes every AI-900 question from scratch — 985 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This AI-900 practice question is part of Courseiva's free Microsoft certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI-900 exam.