hardMultiple Select
PMLE Practice Question: Which THREE factors are critical when designing a…
Which THREE factors are critical when designing a model serving architecture for a global user base with strict latency SLAs? (Choose 3.)
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable autoscaling with request-based metrics to handle traffic spikes.
Options C, D, and E are correct. Option A is wrong because batch prediction is designed for offline processing of large volumes, not for real-time serving with strict latency SLAs. Option B is wrong because a single-region deployment cannot provide low latency to a global user base and does not address data sovereignty issues appropriately; multi-region deployment with fine-grained controls is needed. Option C is correct: autoscaling with request-based metrics ensures that resources scale dynamically to handle traffic spikes while maintaining latency. Option D is correct: caching idempotent predictions reduces latency for repeated queries and offloads the model server. Option E is correct: multi-region deployment with Vertex AI Endpoints in multiple locations minimizes network distance and latency for users worldwide.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use batch prediction to process requests in bulk for efficiency.
Why it's wrong here
Batch prediction returns results only after a job completes, so it cannot satisfy per-request latency SLAs requiring immediate responses. It is tempting because batching amortises compute cost for throughput-oriented workloads, such as overnight scoring, where latency is irrelevant.
- ✗
Deploy the model in a single region to avoid data sovereignty issues.
Why it's wrong here
A single-region deployment forces distant users onto long-haul network paths, adding round-trip latency that breaches strict SLAs. It is tempting because data sovereignty genuinely constrains region choice, but that is a compliance driver, not a latency design.
- ✓
Enable autoscaling with request-based metrics to handle traffic spikes.
Why this is correct
Request-based autoscaling adjusts serving replicas dynamically as traffic spikes, preventing queueing latency that would breach strict SLAs during demand surges. Scaling on CPU alone reacts too slowly for request-driven load, so request metrics directly protect the latency target.
- ✓
Implement request caching for idempotent predictions when appropriate.
Why this is correct
Caching idempotent predictions returns stored results for repeated identical requests, bypassing model inference and its associated compute latency. This reduces response time and load, directly helping meet strict latency SLAs for the global user base.
- ✓
Use multi-region deployment with Vertex AI Endpoints in multiple locations.
Why this is correct
Deploying Vertex AI Endpoints across multiple regions places model replicas near users, cutting network round-trip time that dominates latency for a global audience. This geographic distribution is what satisfies the strict latency SLAs unachievable from a single region.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.