hardMultiple Choice
PMLE Practice Question: Your team has deployed a text classification…
Your team has deployed a text classification model on Vertex AI Endpoints. You notice that the model's latency has increased significantly over the last week, but the request rate has remained stable. Which of the following is the most likely cause?
⚠ Common exam trap
Test-takers frequently confuse 'model latency' with 'request rate' and assume any latency increase must be due to scaling issues, ignoring that preprocessing logic changes can dramatically affect per-request performance without altering throughput.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A change in the preprocessing logic that now includes a computationally expensive step
A computationally expensive preprocessing step directly increases per-request latency on the inference path, even when request rate is stable. Vertex AI Endpoints execute user-provided preprocessing code before model inference, so adding a heavy operation (e.g., large regex, image resizing, or external API call) will linearly increase response time for every prediction.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A sudden increase in the number of prediction requests
Why it's wrong here
Request rate is explicitly stated as stable, so added traffic cannot explain the latency rise; Vertex AI Endpoints would also autoscale replicas to absorb load. Increased request volume is the genuine cause when QPS climbs beyond provisioned capacity, making this the right diagnosis only when monitoring shows a matching spike in prediction requests.
- ✗
The model was replaced with a larger version without updating the endpoint
Why it's wrong here
Replacing a model without redeploying leaves the endpoint serving the old artefact, so latency would not change. It is tempting because a larger model does increase inference time, making it the correct cause only if the endpoint had actually been redeployed with the new version.
- ✓
A change in the preprocessing logic that now includes a computationally expensive step
Why this is correct
A change in preprocessing logic adds a computationally expensive step, directly increasing per-request processing time while request volume stays constant. This satisfies the stem's constraint: stable request rate with rising latency points to per-request cost, not load. Vertex AI Endpoints latency reflects the full prediction pipeline, including client-side preprocessing.
- ✗
A misconfiguration in the autoscaling policy
Why it's wrong here
Autoscaling adjusts replica count to match demand; with request rate flat, it neither triggers nor explains rising latency. It is tempting because autoscaling genuinely addresses latency caused by traffic spikes or queueing under load, so it would be the right answer if the stem showed request rate climbing.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.