Courseiva
mediumMultiple Choice

PMLE Practice Question: A machine learning engineer notices that the…

A machine learning engineer notices that the Vertex AI Prediction endpoint's error rate has increased over the past week. The model was retrained with new data and redeployed. Which step should the engineer take first to diagnose the issue?

⚠ Common exam trap

PMLE often tests whether candidates jump to remediation (rollback, scaling) instead of the observability-first principle — the trap is treating 'roll back' as the safe default when the exam expects evidence gathering first.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Check the Cloud Monitoring dashboard for latency and error codes, and review the model's prediction logs.

The correct first step in any incident diagnosis is to gather observability data — Cloud Monitoring metrics (error rate, latency, error codes) and prediction logs — to characterize the failure before changing anything. This reveals whether errors are 4xx (bad input), 5xx (model/server), timeouts, or a specific code path, and whether the issue correlates with the redeployment. Only after this evidence is collected should the engineer consider rollback or data-distribution analysis.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of replicas to reduce error rate.

    Why it's wrong here

    Adding replicas scales capacity to absorb load; it cannot fix errors caused by a bad model artefact or skewed inputs, and masks the fault. Replica increases suit genuine throughput saturation, where latency rises under volume while prediction quality stays unchanged.

  • ✗

    Compare the input data distribution of recent requests to the training data distribution using Explainable AI.

    Why it's wrong here

    Explainable AI attributes feature contributions to individual predictions; it does not compare distributions, so drift between serving and training inputs stays invisible. Distribution comparison belongs to Vertex AI Model Monitoring, which computes skew and drift against a training baseline.

  • ✗

    Roll back to the previous model version immediately.

    Why it's wrong here

    Rolling back restores availability but destroys the evidence needed to diagnose why the retrained model regressed, and the stem asks for the first diagnostic step. Rollback suits confirmed production incidents where stopping user impact outweighs root-cause investigation.

  • ✓

    Check the Cloud Monitoring dashboard for latency and error codes, and review the model's prediction logs.

    Why this is correct

    Cloud Monitoring exposes endpoint-level latency percentiles and HTTP error codes, while prediction logs capture request payloads and model responses. Together they isolate whether the regression stems from infrastructure, malformed inputs, or the retrained model itself — the fastest evidence-gathering step before altering the deployment.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.