Courseiva
hardMultiple Choice

PMLE Practice Question: Your team is serving a large language model on…

Your team is serving a large language model on Vertex AI using a custom container. The endpoint experiences intermittent 502 errors during traffic spikes. The autoscaling configuration uses a CPU utilization target of 60% and the model is deployed on n1-standard-4 instances. The model requires significant memory. Which combination of changes is most likely to resolve the issue?

⚠ Common exam trap

The trap here is assuming that autoscaling alone can resolve performance issues, when the root cause is often insufficient per-instance resources; candidates may focus on scaling parameters rather than instance sizing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Switch to a machine type with more memory, e.g., n1-highmem-8, and increase min_replica_count.

The 502 errors during traffic spikes indicate that instances are becoming unhealthy or crashing, most likely due to memory exhaustion. The model requires significant memory, and n1-standard-4 instances have only 15 GB of RAM, which is insufficient for a large language model. Switching to n1-highmem-8 (which provides 52 GB of RAM) directly addresses the memory bottleneck, and increasing min_replica_count ensures that enough instances are available to handle baseline traffic and absorb spikes without overloading any single instance. This combination resolves the root cause: insufficient memory leading to instance failures under load.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the target CPU utilization to 90% to allow more requests per instance.

    Why it's wrong here

    Raising the CPU target to 90% delays scale-out, so instances stay saturated longer and requests queue until the container's health probe times out, producing 502s. The memory-heavy model needs a larger machine type and a lower CPU target so replicas are added before exhaustion. A high CPU target suits lightweight, latency-tolerant services.

  • ✓

    Switch to a machine type with more memory, e.g., n1-highmem-8, and increase min_replica_count.

    Why this is correct

    Increasing memory headroom on n1-highmem-8 prevents the container from being killed or stalled when the model's working set exceeds the 15 GB available on n1-standard-4 during spikes, which is what produces the 502s. Raising min_replica_count keeps warm capacity ready so autoscaling lag cannot drop requests.

  • ✗

    Enable canary traffic splitting to reduce load on the main endpoint.

    Why it's wrong here

    Canary splitting routes a percentage of traffic to a second model version; it does not add capacity, so both revisions still saturate during spikes and 502s persist. Traffic splitting is for safely validating a new model version before full rollout, not for relieving load on a single undersized deployment.

  • ✗

    Reduce the model batch size from 32 to 1 to lower memory per request.

    Why it's wrong here

    Dropping batch size to 1 cuts throughput per replica, so each instance serves fewer concurrent requests and queues grow faster during spikes, worsening 502s. Smaller batches suit latency-sensitive single-request inference, not sustained high-throughput serving where memory pressure is better addressed by a larger machine type.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.