easyMultiple Choice
PDE Practice Question: Your company has a machine learning model that…
Your company has a machine learning model that predicts customer churn. The model is deployed on Vertex AI Endpoints with autoscaling. After a marketing campaign, traffic to the endpoint increases by 10x. Some predictions start failing with 'HTTP 503 Service Unavailable' errors. What is the most likely cause?
⚠ Common exam trap
Google Cloud often tests the distinction between model-level errors (e.g., data drift, accuracy degradation) and infrastructure-level errors (e.g., 503, 429, timeout), so the trap here is that candidates confuse a model performance issue with a scaling/availability issue.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The autoscaling configuration has insufficient maximum nodes to handle the traffic.
A 503 Service Unavailable error from Vertex AI Endpoints indicates that the endpoint is overwhelmed and cannot handle the incoming request volume. With a 10x traffic spike and autoscaling configured, the most likely cause is that the autoscaling configuration has insufficient maximum nodes, so the endpoint cannot scale out enough to handle the load, causing requests to be rejected.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The model container has a memory leak.
Why it's wrong here
A memory leak would eventually crash containers, but 503 errors appearing immediately after a 10x traffic surge point to capacity, not gradual memory exhaustion. Leak diagnosis is tempting for intermittent failures, and would be correct if errors accumulated over time independent of traffic volume.
- ✗
The model's accuracy has degraded due to data drift.
Why it's wrong here
Data drift degrades prediction quality, not availability; 503 errors indicate the endpoint cannot serve requests, typically because autoscaling lags the 10x traffic surge. Accuracy monitoring is tempting when traffic patterns change, and would be correct if predictions returned successfully but with degraded quality.
- ✓
The autoscaling configuration has insufficient maximum nodes to handle the traffic.
Why this is correct
Autoscaling adds replicas only up to the configured maximum node count. A 10x traffic surge can exhaust that ceiling, leaving requests unserved and returning HTTP 503. Raising the maximum node limit lets autoscaling provision enough capacity to absorb the campaign load.
- ✗
The model is using an older version that is not supported.
Why it's wrong here
An older model version still serves predictions; Vertex AI does not return 503 because a version is outdated. Version mismatch is tempting when deployments change, and would be correct if requests failed with a version-specific error such as a missing or incompatible model artefact rather than capacity exhaustion.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.