Courseiva
easyMultiple Choice

PMLE Practice Question: You have an online prediction model that is…

You have an online prediction model that is showing increasing prediction latency. You have already verified that the request rate and input data size are unchanged. Which of the following should you investigate next?

⚠ Common exam trap

Google Cloud often tests the distinction between network-level latency (e.g., geographic location) and compute-level latency (e.g., model size), tempting candidates to pick the geographic option when the root cause is model-related.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Check if the model was recently updated to a larger version

If request rate and input data size are unchanged, increased prediction latency often points to a change in the model itself. A larger model (e.g., deeper neural network, more parameters) requires more computation per inference, directly increasing latency. This is a common root cause when monitoring ML pipelines, as model version updates can silently alter performance characteristics.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Check if the model was recently updated to a larger version

    Why this is correct

    Latency rising while request rate and payload size stay constant points to increased per-inference compute. A larger model version, with more parameters or layers, raises inference time. Checking recent deployment history isolates this cause before investigating infrastructure or network factors.

  • ✗

    Check the monitoring dashboard configuration

    Why it's wrong here

    Dashboard configuration governs how metrics are displayed and aggregated, not the model's actual inference path. It would be worth checking when latency figures look implausible or missing, but unchanged request rate and payload size point to a genuine runtime bottleneck rather than a reporting artefact.

  • ✗

    Check if the feature engineering logic was changed

    Why it's wrong here

    Feature engineering logic runs before inference and alters the values passed to the model, not the model's compute cost. It would be the right suspect if input schema or payload size had changed, but those are confirmed unchanged, so latency growth must originate in the serving stack itself.

  • ✗

    Check the geographic location of the endpoint

    Why it's wrong here

    Endpoint geography affects network round-trip time, which would appear as a step change if the deployment region moved. It would be the correct line of enquiry after a migration or failover, but a gradual latency increase with stable traffic indicates resource contention within the serving infrastructure.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.