Courseiva
Implementing AI Solutions →mediumMultiple Choice

AI0-001 Implementing AI Solutions Practice Question

A developer is integrating an AI microservice that accepts image uploads and returns classification labels. The service must handle spikes of up to 1,000 requests per minute but average 100 requests per minute. Which deployment architecture BEST meets these requirements with cost efficiency?

⚠ Common exam trap

AI0-001 often tests whether candidates default to 'serverless = always cheapest' — the trap is missing that synchronous serverless hits concurrency limits under bursty load, while async queue + autoscaling workers is the cost-efficient pattern for spiky workloads.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use an async processing queue (e.g., RabbitMQ) with a pool of worker instances that auto-scale based on queue depth

An async queue with auto-scaling workers decouples ingestion from processing, absorbs bursts by buffering requests, and scales worker count based on queue depth — so you only pay for capacity during actual load. This matches the 10x peak-to-average ratio cost-effectively. Synchronous designs either over-provision for peak or drop requests during spikes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Expose the model via a serverless function (e.g., AWS Lambda) with synchronous invocation

    Why it's wrong here

    Synchronous invocation caps concurrency at the account limit and bills per request duration, so 1,000 requests per minute would throttle rather than scale. It suits steady low-volume event triggers, not bursty image classification where queued asynchronous processing with a backlog absorbs spikes.

  • ✓

    Use an async processing queue (e.g., RabbitMQ) with a pool of worker instances that auto-scale based on queue depth

    Why this is correct

    Queue-depth-based autoscaling lets worker instances expand only during the 1,000-request spikes and contract back to baseline for the 100-request average, so you pay for capacity actually consumed. RabbitMQ decouples ingestion from classification, absorbing bursts without dropping uploads — satisfying both the throughput ceiling and the cost-efficiency constraint.

  • ✗

    Deploy the service as a synchronous REST API on a single always-on VM sized for peak load

    Why it's wrong here

    Sizing a single VM for 1,000 requests per minute leaves it idle at the 100-request average, paying peak capacity continuously. It fits predictable, constant workloads needing persistent local state, not spiky classification traffic where horizontal autoscaling matches cost to demand.

  • ✗

    Stream results directly from the model to the client using WebSockets

    Why it's wrong here

    WebSockets stream continuous results over a persistent connection, which image classification request-response does not need; holding 1,000 concurrent sockets exhausts connection limits. It is correct for real-time token streaming from generative models, not discrete label returns.

About these practice questions

Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official CompTIA exam blueprint

This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.