easyMultiple Choice
PMLE Practice Question: You have an online prediction model that is…
You have an online prediction model that is showing increasing prediction latency. You have already verified that the request rate and input data size are unchanged. Which of the following should you investigate next?
⚠ Common exam trap
Google Cloud often tests the distinction between network-level latency (e.g., geographic location) and compute-level latency (e.g., model size), tempting candidates to pick the geographic option when the root cause is model-related.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Check if the model was recently updated to a larger version
If request rate and input data size are unchanged, increased prediction latency often points to a change in the model itself. A larger model (e.g., deeper neural network, more parameters) requires more computation per inference, directly increasing latency. This is a common root cause when monitoring ML pipelines, as model version updates can silently alter performance characteristics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Check if the model was recently updated to a larger version
Why this is correct
Latency rising while request rate and payload size stay constant points to increased per-inference compute. A larger model version, with more parameters or layers, raises inference time. Checking recent deployment history isolates this cause before investigating infrastructure or network factors.
- ✗
Check the monitoring dashboard configuration
Why it's wrong here
Dashboard configuration governs how metrics are displayed and aggregated, not the model's actual inference path. It would be worth checking when latency figures look implausible or missing, but unchanged request rate and payload size point to a genuine runtime bottleneck rather than a reporting artefact.
- ✗
Check if the feature engineering logic was changed
Why it's wrong here
Feature engineering logic runs before inference and alters the values passed to the model, not the model's compute cost. It would be the right suspect if input schema or payload size had changed, but those are confirmed unchanged, so latency growth must originate in the serving stack itself.
- ✗
Check the geographic location of the endpoint
Why it's wrong here
Endpoint geography affects network round-trip time, which would appear as a step change if the deployment region moved. It would be the correct line of enquiry after a migration or failover, but a gradual latency increase with stable traffic indicates resource contention within the serving infrastructure.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.