Courseiva
mediumMultiple Choice

PMLE Practice Question: A model deployed on Vertex AI Prediction is…

A model deployed on Vertex AI Prediction is returning high latency for real-time requests. The model is a small TensorFlow model. Which troubleshooting step should the team take first?

⚠ Common exam trap

Google Cloud often tests the principle of 'start with the simplest infrastructure fix before optimizing the model or container,' so candidates mistakenly jump to retraining or custom containers without first checking if the instance type and scaling settings are appropriate.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Check if the machine type is too small and enable autoscaling

High latency for real-time predictions from a small TensorFlow model often indicates that the serving infrastructure is under-provisioned. Checking the machine type and enabling autoscaling directly addresses whether the instance is too small to handle the request volume, which is the most common first step in diagnosing latency issues on Vertex AI Prediction.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Retrain the model with a larger batch size

    Why it's wrong here

    Batch size affects throughput during training or batch prediction, not per-request latency for a deployed model. Retraining also does not address serving infrastructure. The first step is inspecting the endpoint's machine type and autoscaling. Larger batch sizes are correct for improving training throughput or batch inference efficiency.

  • ✓

    Check if the machine type is too small and enable autoscaling

    Why this is correct

    A small TensorFlow model should not be inherently slow, so insufficient machine resources or lack of scaling is the likely bottleneck. Checking the machine type and enabling autoscaling addresses CPU or memory saturation before deeper model-level investigation.

  • ✗

    Use a custom container with optimized runtime

    Why it's wrong here

    A custom container addresses dependency or framework mismatches, not latency for a small TensorFlow model already served by a prebuilt container. The first step is checking whether the endpoint uses a GPU or CPU machine type and whether traffic is hitting a cold instance. Custom containers suit non-standard runtimes or libraries.

  • ✗

    Enable Cloud Armor to reduce traffic

    Why it's wrong here

    Cloud Armor filters and blocks malicious or unwanted HTTP(S) traffic at the edge; it does not reduce legitimate request latency on Vertex AI Prediction. High latency for a small model usually stems from machine type, scaling, or cold starts. Cloud Armor is correct when mitigating DDoS or applying WAF rules to a load balancer.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.