AI0-001 AI Infrastructure and Technologies Practice Question
A media company wants to generate short video summaries from long recordings using a generative AI model. The model is hosted in the cloud, and the company needs to minimize cost while handling unpredictable traffic spikes. Which cloud service model is most appropriate?
⚠ Common exam trap
The trap here is equating reserved instances with cost savings for all workloads, when they only benefit steady, predictable usage.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Serverless inference endpoint with automatic scaling
A serverless inference endpoint with automatic scaling aligns cost with actual usage and handles unpredictable traffic without manual intervention. Dedicated VMs, on-premises clusters, and reserved instances all involve either idle costs or commitment risks that are suboptimal for variable demand.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Dedicated GPU virtual machine with a fixed hourly rate
Why it's wrong here
A dedicated GPU VM incurs costs even when idle, which is inefficient for unpredictable traffic. While it provides consistent performance, the fixed hourly rate leads to overspending during low-demand periods. It does not automatically scale, so spikes may cause queuing or require manual intervention.
- ✗
Reserved instances with a one-year commitment
Why it's wrong here
Reserved instances offer discounts for steady, predictable workloads, but they require a long-term commitment and do not scale automatically. For unpredictable spikes, the company may over-provision or under-provision. This option is cost-effective only when usage is stable and known in advance.
- ✓
Serverless inference endpoint with automatic scaling
Why this is correct
Serverless inference endpoints scale automatically with traffic and charge only for the compute used during requests, which minimizes cost for unpredictable spikes. This model eliminates idle capacity costs and matches the variable demand of video summarization workloads.
- ✗
On-premises GPU cluster with a load balancer
Why it's wrong here
On-premises infrastructure requires upfront capital and ongoing maintenance, and it cannot scale elastically to unpredictable cloud-level spikes. The scenario specifies a cloud-hosted model, so on-premises is not aligned. This option increases cost and operational complexity without addressing the elasticity requirement.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.