PMLE Serving and Scaling Models Practice Question
An ML platform team deploys the same custom container to two Vertex AI endpoints: one for interactive scoring and one for nightly bulk scoring. The interactive endpoint must return predictions in under 200 ms and receives small single-record requests. The bulk endpoint sends large batched requests and tolerates seconds of latency. Both endpoints currently use the same machine type and the same container image, and the interactive endpoint frequently misses its latency target. Which change best resolves the interactive latency problem?
⚠ Common exam trap
The trap here is treating a latency SLO breach as a pure capacity problem and adding replicas or vCPUs, when the fixed per-request overhead of batch-oriented model loading is the actual cause.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy a container variant tuned for the interactive endpoint: load the model once at startup, keep it resident, and avoid per-request batch assembly, then serve it on a latency-optimized machine type.
Interactive single-record scoring and bulk batched scoring have opposite performance profiles, so they should not share one serving configuration. The interactive path is dominated by fixed per-request overhead such as model loading and batch assembly, so a container that loads the model once, keeps it resident, and skips per-request batching removes that cost. Pairing it with a machine type selected for low-latency single requests, while leaving the bulk endpoint tuned for throughput, resolves the SLO breach without over-sizing the bulk tier.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the request timeout on the interactive endpoint and add more replicas.
Why it's wrong here
Raising a timeout does not make predictions faster; it only changes when the client gives up, so the SLO breach persists from the caller's perspective. Adding replicas helps when the endpoint is saturated by concurrent requests, but the described symptom occurs even on small single-record calls, which points to per-request processing cost rather than queueing. Neither change removes that fixed overhead.
- ✗
Enable autoscaling on the bulk endpoint so it absorbs the interactive traffic as well.
Why it's wrong here
The two endpoints serve distinct workloads and are not interchangeable; routing interactive traffic to the bulk endpoint would mix small low-latency requests with large batched ones. Autoscaling the bulk endpoint responds to aggregate load but cannot guarantee per-request latency, and the interactive endpoint would still be running a container configuration optimized for throughput rather than single-record response time.
- ✗
Move both endpoints to a machine type with more vCPUs and more memory.
Why it's wrong here
Scaling up both endpoints raises cost and does not isolate the interactive workload from the bulk workload's resource profile. If the container performs per-record preprocessing that dominates small-request latency, extra vCPUs give little benefit because the request path is serial. The bulk endpoint gains capacity it does not need, while the interactive endpoint may still miss the target.
- ✓
Deploy a container variant tuned for the interactive endpoint: load the model once at startup, keep it resident, and avoid per-request batch assembly, then serve it on a latency-optimized machine type.
Why this is correct
Latency on small requests is dominated by fixed per-request overhead, so removing repeated model loading and batch assembly from the request path gives the largest gain. A container that loads the model once and keeps it warm in memory, paired with a machine type chosen for single-request speed rather than bulk throughput, directly targets the 200 ms SLO while leaving the bulk endpoint's throughput-optimized configuration untouched.
Visual reference
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.