AI0-001 AI Implementation and Operations Practice Question
A media company serves personalized article recommendations through a model hosted on a cloud inference service. During a major news event, request volume spikes tenfold and p95 latency rises from 120 ms to over 2 seconds, causing timeouts on the web front end. The model itself is unchanged and the endpoint is healthy. The team wants to keep serving personalized results during spikes without degrading the user experience. Which action should the team take first?
⚠ Common exam trap
The trap here is treating a capacity and serving-path problem as a model-quality problem, which leads to retraining or threshold changes that cannot reduce inference latency.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cache or precompute recommendations for high-traffic article and user segments and serve those cached results during the spike.
The latency spike is caused by a tenfold surge in request volume against a fixed-capacity inference path, not by a model defect. Caching or precomputing recommendations for high-traffic segments removes redundant inference work from the request path, cutting p95 latency immediately while preserving personalization. Retraining, threshold tuning, and static fallbacks either do not reduce per-request compute or sacrifice the personalization the team must retain.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Retrain the recommendation model on a larger dataset so it produces results faster under load.
Why it's wrong here
Training data volume does not determine inference latency; the model architecture and serving configuration do. Retraining on more data would take time, require redeployment, and leave the latency spike during the news event unresolved. The scenario states the model is unchanged and healthy, so the bottleneck is capacity and serving behavior, not model quality.
- ✗
Switch the front end to a static, non-personalized article list whenever request volume exceeds the autoscaling limit.
Why it's wrong here
Falling back to non-personalized content abandons the personalization the business depends on and is a blunt degradation rather than a fix. The team explicitly wants to keep serving personalized results during spikes, so disabling personalization contradicts the stated requirement even though it might reduce load.
- ✗
Lower the model's prediction confidence threshold so fewer candidates are scored per request.
Why it's wrong here
A confidence threshold filters or ranks outputs after scoring; it does not reduce the number of candidate items the model must evaluate, so per-request compute and latency remain essentially unchanged. Adjusting it would also change recommendation quality, which is an unintended side effect that does not address the capacity-driven latency spike.
- ✓
Cache or precompute recommendations for high-traffic article and user segments and serve those cached results during the spike.
Why this is correct
Serving precomputed or cached recommendations for the hottest segments removes most inference calls from the critical path during the spike, directly reducing p95 latency and preventing front-end timeouts. This is the fastest operational lever because it requires no model change and can be enabled immediately, preserving personalization quality for the segments that matter most during the event.
About these practice questions
Courseiva writes every AI0-001 question from scratch — 962 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official CompTIA exam blueprint
This AI0-001 practice question is part of Courseiva's free CompTIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the AI0-001 exam.