hardMultiple Choice
PMLE Practice Question: A company has deployed a machine learning model…
A company has deployed a machine learning model that uses a large input tensor. They notice that the prediction latency varies significantly between requests of the same size. Cloud Monitoring shows that the serving endpoint's CPU utilization is consistently below 50%, but memory utilization fluctuates between 70% and 95%. What is the most likely cause?
⚠ Common exam trap
Google Cloud often tests the misconception that high memory utilization always indicates a memory leak, but the key differentiator is the pattern of fluctuation versus monotonic increase, and the fact that GC pauses cause latency spikes without high CPU usage.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The model is performing garbage collection cycles
The described symptoms—low CPU utilization (below 50%) and high, fluctuating memory utilization (70%–95%) with variable latency—are classic indicators of garbage collection (GC) pauses in a managed runtime like Python or Java. When the model processes large input tensors, it allocates significant memory; as memory pressure builds, the garbage collector runs more frequently, causing stop-the-world pauses that increase latency unpredictably, even though CPU is not fully utilized.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The model is performing garbage collection cycles
Why this is correct
With CPU below 50% but memory swinging between 70% and 95%, the variability stems from memory reclamation rather than compute saturation. Garbage collection cycles pause request handling while reclaiming the large input tensor's allocations, producing latency spikes that correlate with memory pressure, not CPU load.
- ✗
The model is using excessive memory due to a memory leak
Why it's wrong here
A leak grows memory monotonically until the process is killed, producing a crash or restart rather than latency varying between identical requests. Leak diagnosis fits steadily climbing memory across hours; here memory oscillates between 70% and 95% within normal serving.
- ✗
The prediction latency is being affected by CPU throttling
Why it's wrong here
CPU throttling requires sustained CPU saturation, yet utilisation stays below 50%, so the CPU quota is never exhausted. Throttling is the right diagnosis when CPU sits at its limit and latency spikes track quota periods, which the monitoring data here contradicts.
- ✗
The model is hitting a cold start due to autoscaling
Why it's wrong here
Cold starts occur when an instance initialises after scaling, adding startup delay to the first request, not to every same-sized request. Autoscaling with cold-start mitigation suits bursty traffic where new replicas spin up; here memory swings during steady serving point elsewhere.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.