Courseiva
easyMultiple Choice

PMLE Practice Question: A company deploys a TensorFlow model on Vertex AI…

A company deploys a TensorFlow model on Vertex AI Prediction with a single node. During peak hours, inference latency increases. What should they do first to reduce latency?

⚠ Common exam trap

Candidates often confuse improving throughput (batching or bigger machines) with reducing latency under load, but the first action should always be to add more replicas via autoscaling to handle concurrent requests, not to optimize a single node's performance.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable autoscaling for the deployment

Enabling autoscaling for the deployment is the correct first step because it allows Vertex AI Prediction to dynamically adjust the number of replicas based on incoming traffic. During peak hours, autoscaling can add more nodes to distribute the inference load, directly reducing latency without requiring manual intervention or over-provisioning.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable autoscaling for the deployment

    Why this is correct

    Autoscaling adds replicas when traffic rises, spreading inference load across nodes so each handles fewer requests. This directly addresses the single-node bottleneck causing peak-hour latency, and it is the least disruptive first step before considering larger machines or GPUs.

  • ✗

    Increase the machine type of the node

    Why it's wrong here

    A larger machine type raises per-node throughput but leaves the single node as the sole replica, so concurrent peak requests still queue behind one another; horizontal scaling adds serving capacity. It tempts because CPU saturation on one node does look like a sizing problem, and resizing is the right fix when a single replica is genuinely underpowered.

  • ✗

    Decrease the min replicas to 0

    Why it's wrong here

    Setting min replicas to zero lets Vertex AI scale the deployment down to no running replicas, so peak traffic hits cold starts and latency worsens rather than improves. It tempts because reducing idle replicas cuts cost, which is the correct goal when traffic is predictably low and cost, not latency, is the constraint.

  • ✗

    Enable automatic batching of requests

    Why it's wrong here

    Automatic batching groups incoming requests into larger inference calls, which raises throughput but adds queueing delay before each request is served, so tail latency grows under peak load. It tempts because batching genuinely improves hardware utilisation, and it is the right choice when throughput per node, not per-request latency, is the bottleneck.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.