Courseiva
mediumMultiple Choice

PMLE Practice Question: A team deploys a PyTorch model on Vertex AI for…

A team deploys a PyTorch model on Vertex AI for online predictions. They notice that after deployment, the latency increases over time, especially during peak hours. The model is served using a custom container. What is the most likely cause?

⚠ Common exam trap

Google Cloud often tests the misconception that latency increases are always due to resource exhaustion (memory/CPU) rather than concurrency or request handling inefficiencies, leading candidates to pick Option C.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The prediction requests are not being batched, and the model inference code is not optimized for concurrency.

The latency increase over time, especially during peak hours, indicates that the model inference code is not handling concurrent requests efficiently. Without batching or optimized concurrency, each request is processed sequentially, causing a queue buildup under load. This is a common issue with custom containers on Vertex AI when the prediction handler is single-threaded or lacks async processing.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The custom container does not have a health check, causing instances to be prematurely terminated.

    Why it's wrong here

    Absent health checks cause failed instances to keep serving or be replaced, producing errors and restarts rather than steadily climbing latency at peak. It is tempting because health checks matter for availability, and adding one is right when instances are being terminated prematurely.

  • ✗

    The model is not using GPU even though a GPU machine is selected.

    Why it's wrong here

    Would affect throughput but not latency growth over time.

  • ✗

    The model is too large for the machine's memory, causing swapping.

    Why it's wrong here

    Memory exhaustion would cause out-of-memory crashes or consistent slowness from the first request, not a gradual rise tied to peak-hour traffic. It is tempting because oversized models do degrade serving, and resizing the machine would be the fix if memory were genuinely the bottleneck.

  • ✓

    The prediction requests are not being batched, and the model inference code is not optimized for concurrency.

    Why this is correct

    Without request batching, each call triggers a separate forward pass, and unoptimised inference code serialises concurrent requests. Under peak load this queues requests, so latency climbs as the container saturates rather than scaling with traffic.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.