Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

A machine learning team deploys a PyTorch model for online prediction on Vertex AI using a custom container. They notice that the first few requests after scaling up experience high latency. What is the most likely cause and how should they mitigate it?

⚠ Common exam trap

Test-takers frequently confuse slow initialization with autoscaling misconfiguration, assuming that scale-to-zero or smaller machines would fix the latency, when in fact the root cause is the readiness probe timing.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The container has a slow initialization; set initialDelaySeconds in the health probe to give more time before considering the pod ready.

The high latency on the first few requests after scaling up is a classic symptom of a slow container initialization. By setting `initialDelaySeconds` in the health probe, you allow the container more time to start up and become ready before it receives traffic, preventing premature routing that causes timeouts or retries. This is a common tuning parameter for custom containers on Vertex AI, where model loading or dependency initialization can take several seconds.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The endpoint is not configured for autoscaling; enable min_replica=0 to allow scale-to-zero.

    Why it's wrong here

    Setting min_replica=0 causes cold starts, worsening the latency described; the real cause is container image pull and model load during scale-up, mitigated by min_replica≥1 or a warm pool. It is tempting because scale-to-zero cuts idle cost, which suits sporadic, latency-tolerant batch traffic.

  • ✗

    The model file is corrupted; re-upload to Vertex AI Model Registry.

    Why it's wrong here

    A corrupted model file would fail predictions consistently, not only during the first requests after scaling up; the latency stems from cold-start initialisation. It is tempting because re-uploading is a quick fix, but it would be correct if every request returned errors regardless of replica state.

  • ✓

    The container has a slow initialization; set initialDelaySeconds in the health probe to give more time before considering the pod ready.

    Why this is correct

    Slow container initialisation delays readiness, so the pod receives traffic before PyTorch and model weights finish loading, producing cold-start latency after scaling. Raising initialDelaySeconds postpones the readiness probe's first check, preventing premature traffic routing until initialisation completes. This directly addresses the scaling-induced latency constraint.

  • ✗

    Use a smaller machine type (n1-standard-2) to reduce startup overhead.

    Why it's wrong here

    A smaller machine type reduces CPU and memory, slowing model loading and inference rather than cutting startup overhead; the cause is cold-start container and weight initialisation. It is tempting because smaller instances cost less, but it would be correct for steady low-throughput workloads where latency is not critical.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.