Courseiva

AI0-001 AI Infrastructure and Technologies Practice Question

A media company wants to generate short video summaries from long recordings using a generative AI model. The model is hosted in the cloud, and the company needs to minimize cost while handling unpredictable traffic spikes. Which cloud service model is most appropriate?

⚠ Common exam trap

The trap here is equating reserved instances with cost savings for all workloads, when they only benefit steady, predictable usage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Serverless inference endpoint with automatic scaling

A serverless inference endpoint with automatic scaling aligns cost with actual usage and handles unpredictable traffic without manual intervention. Dedicated VMs, on-premises clusters, and reserved instances all involve either idle costs or commitment risks that are suboptimal for variable demand.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Dedicated GPU virtual machine with a fixed hourly rate

    Why it's wrong here

    A dedicated GPU VM incurs costs even when idle, which is inefficient for unpredictable traffic. While it provides consistent performance, the fixed hourly rate leads to overspending during low-demand periods. It does not automatically scale, so spikes may cause queuing or require manual intervention.

  • ✗

    Reserved instances with a one-year commitment

    Why it's wrong here

    Reserved instances offer discounts for steady, predictable workloads, but they require a long-term commitment and do not scale automatically. For unpredictable spikes, the company may over-provision or under-provision. This option is cost-effective only when usage is stable and known in advance.

  • ✓

    Serverless inference endpoint with automatic scaling

    Why this is correct

    Serverless inference endpoints scale automatically with traffic and charge only for the compute used during requests, which minimizes cost for unpredictable spikes. This model eliminates idle capacity costs and matches the variable demand of video summarization workloads.

  • ✗

    On-premises GPU cluster with a load balancer

    Why it's wrong here

    On-premises infrastructure requires upfront capital and ongoing maintenance, and it cannot scale elastically to unpredictable cloud-level spikes. The scenario specifies a cloud-hosted model, so on-premises is not aligned. This option increases cost and operational complexity without addressing the elasticity requirement.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.