PMLE Scaling Prototypes into ML Models Practice Question
You are deploying a scikit-learn model to Vertex AI for online prediction. The model expects a JSON payload with a single feature vector. You need to ensure the endpoint can handle bursts of traffic up to 1000 requests per second while maintaining low latency. What should you do?
⚠ Common exam trap
The trap here is thinking that provisioning for peak load or using GPUs will solve latency issues, when the key is elastic scaling based on actual demand.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy the model to a Vertex AI Endpoint with a single machine type and enable autoscaling based on CPU utilization.
Vertex AI Endpoints provide autoscaling to dynamically adjust replicas based on traffic. By enabling autoscaling with CPU utilization as the metric, the endpoint can scale out during bursts and scale in during low traffic, ensuring low latency and cost efficiency. This is the correct way to handle variable online prediction traffic.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Deploy the model to a Vertex AI Endpoint with a single machine type and enable autoscaling based on CPU utilization.
Why this is correct
Vertex AI Endpoint supports autoscaling, which automatically adjusts the number of replicas based on traffic. Setting a minimum and maximum replica count with CPU utilization as the metric allows the endpoint to handle bursts while keeping latency low. This is the standard and recommended approach for scaling online predictions.
- ✗
Deploy the model to a Vertex AI Endpoint with a fixed number of replicas equal to the expected peak traffic.
Why it's wrong here
A fixed number of replicas sized for peak traffic wastes resources during low traffic and may still be insufficient if traffic exceeds the estimate. It does not adapt to bursts and lacks the elasticity of autoscaling. This approach is not cost-effective and does not guarantee low latency under variable load.
- ✗
Use a batch prediction job to process incoming requests every minute.
Why it's wrong here
Batch prediction is designed for asynchronous, offline processing of large datasets, not for real-time online prediction with low latency. It cannot handle per-request bursts and introduces significant delay. This option does not meet the requirement for online prediction with low latency.
- ✗
Deploy the model to a Vertex AI Endpoint with a GPU machine type to accelerate inference.
Why it's wrong here
For a scikit-learn model, GPU acceleration is typically unnecessary and may not be supported efficiently. The bottleneck for online prediction is often CPU and network, not GPU. Adding GPUs increases cost without addressing the need for autoscaling to handle traffic bursts.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.