hardMultiple Choice
PMLE Practice Question: A recommendation system model is updated daily…
A recommendation system model is updated daily via a retraining pipeline. After each update, the online prediction latency increases significantly for about 30 minutes before returning to normal. What is the most likely cause and solution?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The new model version causes cold start in the serving infrastructure; pre-warm the model by sending a dummy request after deployment.
The most likely cause is a cold start when the new model version is deployed. The serving infrastructure needs to load the model into memory, which takes time and causes increased latency for the first requests. Pre-warming the model by sending dummy requests after deployment can mitigate this. Option A is incorrect because autoscaling policy affecting scale-down would not cause a temporary spike after each update; it would be more persistent. Option B is incorrect because GKE cluster sharing resources would cause consistent latency issues, not just for 30 minutes after update. Option C is incorrect because CPU-to-GPU switching is not a typical cause for temporary latency spikes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The Vertex AI endpoint autoscaling policy is too aggressive, causing scale-down during retraining.
Why it's wrong here
Autoscaling reacts to traffic and utilisation, not to retraining; scale-down would not produce a fixed 30-minute latency spike after each update. It tempts because aggressive scaling genuinely causes latency under fluctuating load, and tuning minReplicaCount would be correct if the stem described traffic-driven capacity shortfalls.
- ✗
The retraining pipeline runs on a GKE cluster that shares resources with the serving endpoint.
Why it's wrong here
Shared GKE resources would cause variable contention-driven latency, not a consistent 30-minute spike ending on its own after each retrain. It tempts because co-located workloads genuinely degrade serving, and isolating the pipeline would be correct if the stem described ongoing resource starvation rather than a transient warm-up pattern.
- ✗
The model is being switched from CPU to GPU at deployment.
Why it's wrong here
Switching CPU to GPU at deployment would change latency persistently, not for a fixed 30-minute window after each retrain. It tempts because GPU/CPU mismatches do cause inference problems, and matching accelerator hardware would be correct if the stem described sustained latency or throughput regression rather than a transient spike.
- ✓
The new model version causes cold start in the serving infrastructure; pre-warm the model by sending a dummy request after deployment.
Why this is correct
Cold start occurs because the newly deployed model version's serving containers must load weights and initialise before handling traffic, inflating latency until caches warm. Sending a dummy request immediately after deployment pre-warms the infrastructure, eliminating the 30-minute degradation window.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.