AI0-001 AI Implementation and Operations Practice Question
A retail bank operates a real-time AI service that approves or declines card transactions in under 100 ms. During a marketing campaign, transaction volume triples and the inference service's p99 latency rises to 1.4 seconds, causing checkout timeouts. The model is unchanged and CPU utilization on the inference nodes is only 35%. Which action BEST addresses the latency increase while preserving the sub-100 ms requirement?
⚠ Common exam trap
The trap here is assuming that low CPU utilization means the service is healthy and that latency must be a model-quality problem rather than a queueing capacity problem.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable autoscaling of the inference replicas based on request concurrency or queue depth.
High p99 latency with low CPU utilization is the classic signature of request queueing, not compute saturation. Scaling the number of inference replicas based on concurrency or queue depth increases parallel capacity so requests are served promptly, restoring the sub-100 ms SLA. Neither retraining, bigger instances, nor batching addresses the queueing root cause, and batching would actually increase per-request latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Move the model to a larger instance type with more vCPUs and memory per replica.
Why it's wrong here
Scaling up helps when a single request is CPU-bound, but CPU utilization is only 35%, showing the compute is not saturated. Adding vCPUs to each replica increases per-instance cost without reducing queueing caused by insufficient parallel capacity. Horizontal scaling across more replicas addresses the concurrency bottleneck far more cost-effectively than vertical scaling of underutilized nodes.
- ✗
Increase the inference batch size so the service processes more transactions per request.
Why it's wrong here
Batching improves throughput but adds queueing delay because each request must wait for the batch to fill, which directly worsens the per-transaction latency the bank must keep under 100 ms. With CPU already at only 35%, the bottleneck is not raw compute throughput, so larger batches would increase waiting time without solving the tail-latency problem observed at p99.
- ✓
Enable autoscaling of the inference replicas based on request concurrency or queue depth.
Why this is correct
Low CPU utilization combined with high p99 latency indicates requests are queueing rather than computing, so scaling out replicas reduces the queue and restores the sub-100 ms target. Scaling on concurrency or queue depth reacts to the actual saturation signal, unlike CPU-based scaling that would remain idle at 35% and never trigger during the campaign traffic spike.
- ✗
Retrain the fraud model on the campaign-period transaction data and redeploy it.
Why it's wrong here
Retraining changes model quality, not the serving path's capacity to handle three times the request volume. The latency spike is caused by queueing under load, not by model complexity or accuracy, so a retrained model of similar size would exhibit the same p99 degradation. Retraining also takes days and requires validation, which does not meet an immediate production latency incident.
About these practice questions
This AI0-001 question is part of Courseiva's 962-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.