Courseiva
hardMultiple Choice

PMLE Practice Question: A team deploys a real-time model using a custom…

A team deploys a real-time model using a custom container on Vertex AI Prediction. The container is large (5 GB) and cold starts are causing latency spikes. The endpoint is configured with `min_replica_count=0` to reduce cost. The team wants to keep the cost low while reducing cold starts. What is the best approach?

⚠ Common exam trap

PMLE often tests the trade-off between cost and latency, tempting candidates to pick image-size or infrastructure tweaks when the direct, supported fix is simply raising min_replica_count to keep a warm replica.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set `min_replica_count=1` to keep at least one replica always warm.

Setting min_replica_count=1 keeps at least one replica always provisioned and warm, eliminating cold starts for the first request after idle while still allowing autoscaling to add replicas under load. This directly addresses the latency spikes caused by the 5 GB container's slow initialization, and the cost of one always-on replica is typically far lower than the business impact of cold-start latency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Set `min_replica_count=1` to keep at least one replica always warm.

    Why this is correct

    Setting min_replica_count to 1 keeps one replica permanently provisioned, eliminating cold-start latency from the 5 GB image while retaining autoscaling for spikes. It balances the stem's dual constraints: reducing cold starts without abandoning cost control entirely.

  • ✗

    Use a prebuilt container for the model framework to reduce image size.

    Why it's wrong here

    Swapping to a prebuilt container changes the runtime image and may not support the model's custom dependencies or inference code, so it cannot be assumed to shrink the 5 GB artefact. It is tempting because prebuilt containers are smaller, which is correct when the model uses a standard framework without bespoke serving logic.

  • ✗

    Enable container memory optimization to reduce startup time.

    Why it's wrong here

    Vertex AI Prediction exposes no container memory optimisation setting that alters image pull or process start time; the flag does not exist as described. It is tempting because memory tuning sounds like a startup lever, and it would be relevant when the bottleneck is runtime memory pressure during inference rather than image download.

  • ✗

    Provision a Persistent Disk (SSD) for the container image to speed up download.

    Why it's wrong here

    Attaching a Persistent Disk does not cache the container image layers; Vertex AI still pulls the 5 GB image to the replica on each cold start. It is tempting because SSD-backed storage accelerates I/O, which is the right fix when the workload reads large datasets from disk at runtime, not when the delay is image retrieval.

About these practice questions

Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.