Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You are deploying a model to a Vertex AI Endpoint that must serve predictions with a strict 50 ms latency SLA. The model is a custom container that loads a large model file from Cloud Storage at startup. You notice that the first few predictions after a new deployment are slow, and sometimes the endpoint scales up and the new replicas also have slow first predictions. What should you do to reduce this cold-start latency?

⚠ Common exam trap

The trap here is focusing on autoscaling or machine size when the real cold-start bottleneck is the model file download at container startup.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Bake the model file into the custom container image so that it does not need to be downloaded from Cloud Storage at startup.

The most direct way to reduce cold-start latency is to eliminate the need to download the large model file at startup by baking it into the container image. This removes a major source of delay for new replicas. The other options either do not address the startup download, add cost without solving the root cause, or are not applicable to custom models.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Bake the model file into the custom container image so that it does not need to be downloaded from Cloud Storage at startup.

    Why this is correct

    Embedding the model file in the container image removes the Cloud Storage download from the startup path, which is often the dominant cause of cold-start latency. The container image is cached on the host, so new replicas can start faster and serve predictions within the SLA. This is a common technique for latency-sensitive custom containers on Vertex AI.

  • ✗

    Increase the minimum number of replicas and enable autoscaling with a lower target utilization so that replicas are always warm.

    Why it's wrong here

    Increasing minimum replicas keeps more instances running, which can reduce the frequency of cold starts, but it does not address the slow startup of each new replica when it loads the large model file. New replicas will still experience the same cold-start latency, so this does not fully solve the problem.

  • ✗

    Configure the endpoint to use a pre-built container instead of a custom container so that the model is loaded more efficiently.

    Why it's wrong here

    Pre-built containers are optimized for specific frameworks, but they do not automatically solve cold-start latency for a custom model. If the model is not compatible with a pre-built container, switching is not possible; even if it is, the startup time depends on model loading, which the pre-built container does not magically optimize.

  • ✗

    Use a larger machine type for the endpoint so that the model file is downloaded faster and the model loads more quickly.

    Why it's wrong here

    A larger machine type may increase network bandwidth and CPU, but the bottleneck is often the download and initialization of the large model file, not raw compute. Increasing machine size does not eliminate the download step and may not meaningfully reduce cold-start time, while also increasing cost.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.