Courseiva
easyMultiple Choice

PMLE Practice Question: A company deploys a model on Vertex AI Endpoints…

A company deploys a model on Vertex AI Endpoints for real-time inference. They notice latency spikes during peak hours. Which action is most effective to reduce latency without sacrificing accuracy?

⚠ Common exam trap

PMLE often tests whether candidates pick a static fix (bigger machine, pruning) when the scenario describes a dynamic load problem — the trap is missing that 'peak hours' implies autoscaling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Enable autoscaling based on CPU utilization

Latency spikes during peak hours indicate the endpoint is under-provisioned for concurrent load. Enabling autoscaling based on CPU utilization (or a custom metric like request count per replica) lets Vertex AI add replicas as demand rises, absorbing the spike without changing the model or sacrificing accuracy. This is the most direct and effective action for peak-hour latency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Enable autoscaling based on CPU utilization

    Why this is correct

    Autoscaling adds replica capacity when CPU utilisation rises, spreading inference requests across more nodes during peak load. This reduces per-request queueing latency while the same model and precision are served, so accuracy is unchanged. It directly addresses the peak-hour latency spikes.

  • ✗

    Use a larger machine type

    Why it's wrong here

    A larger machine type raises per-instance throughput but does not reduce the queueing delay that causes peak-hour latency; scaling the endpoint's replica count (autoscaling) distributes concurrent requests instead. Larger machines suit sustained high CPU or memory demand from a single model, not bursty traffic spikes.

  • ✗

    Reduce model size by pruning

    Why it's wrong here

    Pruning alters the model's weights and structure, which can degrade accuracy, so it does not satisfy the no-sacrifice constraint. It tempts as a latency reduction technique, but the correct approach adds replicas or autoscaling to Vertex AI Endpoints to absorb peak traffic without touching the model.

  • ✗

    Implement client-side caching

    Why it's wrong here

    Client-side caching only helps repeated identical requests and does nothing for the endpoint's own compute latency during peak load. It tempts because caching reduces perceived response time, but the stem requires reducing server-side inference latency, which endpoint autoscaling or additional replicas address.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.