mediumMultiple Choice
PMLE An ML engineer at a fintech company Practice Question
You are an ML engineer at a fintech company. You have a prototype credit risk model built using XGBoost that achieves high accuracy on historical data. The model is trained on a dataset with 500,000 rows and 50 features. The company wants to deploy this model to production to score loan applications in real-time. The production environment must handle a peak load of 100 requests per second with a latency under 200ms. You have decided to use Vertex AI for deployment. After deploying the model as a Vertex AI endpoint with a single n1-standard-4 machine, you notice that latency exceeds 500ms at peak load and some requests time out. You have verified that the model prediction itself (excluding network overhead) takes about 50ms on average. What should you do to meet the latency and throughput requirements?
⚠ Common exam trap
Candidates often assume latency issues are always due to model inference speed (leading them to choose GPU or model pruning), when in fact the bottleneck is often the lack of horizontal scaling to handle concurrent requests under load.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable autoscaling with a minimum of 2 replicas and use a larger machine type (e.g., n1-standard-8) to handle more concurrent requests.
The latency bottleneck is not the model inference time (50ms) but the inability of a single n1-standard-4 machine to handle 100 concurrent requests per second without queuing. By enabling autoscaling with a minimum of 2 replicas and upgrading to n1-standard-8, you increase both the number of concurrent requests the endpoint can process and the CPU/memory resources per replica, reducing queue wait times and keeping total latency under 200ms. This directly addresses the throughput and latency requirements without changing the model or switching to batch processing.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Change the machine type to a GPU-accelerated machine like n1-standard-4 with a T4 GPU.
Why it's wrong here
Changing to a GPU-accelerated machine is ineffective for this XGBoost model's inference, as its 50ms CPU prediction time indicates the bottleneck is not individual computation speed but concurrent request handling. GPUs primarily accelerate highly parallelisable deep learning or large model inference, where their many cores significantly reduce per-prediction latency. For this scenario, GPU overhead could even increase latency, failing to address the throughput requirement by providing more concurrent processing capacity.
- ✗
Prune the model to reduce size and improve prediction speed.
Why it's wrong here
Pruning shrinks model size but the stem shows prediction already takes only 50ms; the 500ms latency comes from a single n1-standard-4 saturating under 100 requests per second, so pruning leaves the throughput bottleneck. It is tempting because pruning genuinely speeds inference on resource-constrained or edge deployments.
- ✓
Enable autoscaling with a minimum of 2 replicas and use a larger machine type (e.g., n1-standard-8) to handle more concurrent requests.
Why this is correct
Autoscaling to at least two replicas distributes the 100 requests per second across instances, while the larger n1-standard-8 machine provides more CPU for concurrent inference. Together these cut the 500ms latency below the 200ms target.
- ✗
Switch from online prediction to batch prediction using Vertex AI Batch Prediction.
Why it's wrong here
Batch prediction returns results asynchronously to a storage destination, so it cannot answer individual loan applications within 200ms. It is tempting because it removes endpoint latency and cost for bulk scoring, which suits nightly scoring of accumulated applications rather than real-time requests.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.