Courseiva
hardMultiple Choice

PMLE Practice Question: Your team has deployed a text classification…

Your team has deployed a text classification model on Vertex AI Endpoints. You notice that the model's latency has increased significantly over the last week, but the request rate has remained stable. Which of the following is the most likely cause?

⚠ Common exam trap

Test-takers frequently confuse 'model latency' with 'request rate' and assume any latency increase must be due to scaling issues, ignoring that preprocessing logic changes can dramatically affect per-request performance without altering throughput.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

A change in the preprocessing logic that now includes a computationally expensive step

A computationally expensive preprocessing step directly increases per-request latency on the inference path, even when request rate is stable. Vertex AI Endpoints execute user-provided preprocessing code before model inference, so adding a heavy operation (e.g., large regex, image resizing, or external API call) will linearly increase response time for every prediction.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    A sudden increase in the number of prediction requests

    Why it's wrong here

    Request rate is explicitly stated as stable, so added traffic cannot explain the latency rise; Vertex AI Endpoints would also autoscale replicas to absorb load. Increased request volume is the genuine cause when QPS climbs beyond provisioned capacity, making this the right diagnosis only when monitoring shows a matching spike in prediction requests.

  • ✗

    The model was replaced with a larger version without updating the endpoint

    Why it's wrong here

    Replacing a model without redeploying leaves the endpoint serving the old artefact, so latency would not change. It is tempting because a larger model does increase inference time, making it the correct cause only if the endpoint had actually been redeployed with the new version.

  • ✓

    A change in the preprocessing logic that now includes a computationally expensive step

    Why this is correct

    A change in preprocessing logic adds a computationally expensive step, directly increasing per-request processing time while request volume stays constant. This satisfies the stem's constraint: stable request rate with rising latency points to per-request cost, not load. Vertex AI Endpoints latency reflects the full prediction pipeline, including client-side preprocessing.

  • ✗

    A misconfiguration in the autoscaling policy

    Why it's wrong here

    Autoscaling adjusts replica count to match demand; with request rate flat, it neither triggers nor explains rising latency. It is tempting because autoscaling genuinely addresses latency caused by traffic spikes or queueing under load, so it would be the right answer if the stem showed request rate climbing.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.