mediumMultiple Choice
PDE Practice Question: Refer to the exhibit
Network Topology
Refer to the exhibit. A data scientist deploys a model using this configuration. Users report that after a few hours of inactivity, the first prediction request takes over 30 seconds. What is the most likely cause?
⚠ Common exam trap
Google Cloud often tests the distinction between cold start latency (caused by scaling to zero) and persistent performance issues like network latency or resource exhaustion, so candidates must recognize that a delay only after inactivity points to replica provisioning, not a constant problem.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The automatic scaling configuration allows scaling down to zero replicas, causing a cold start on the first request.
The automatic scaling configuration that allows scaling down to zero replicas means that after a period of inactivity, all model replicas are terminated. When a new prediction request arrives, the endpoint must provision a new replica from scratch, which involves loading the model artifacts, initializing the inference container, and performing health checks. This cold start process typically takes 30 seconds or more, matching the reported behavior.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The automatic scaling configuration allows scaling down to zero replicas, causing a cold start on the first request.
Why this is correct
Scaling to zero removes all serving containers, so the next request must provision a replica and load the TensorFlow model, producing a cold-start delay of tens of seconds. Keeping a minimum replica count avoids this.
- ✗
The network latency between the client and the endpoint is high due to regional distance.
Why it's wrong here
Regional distance adds a constant per-request delay, so every call would be slow, not only the first after inactivity. It is tempting because latency genuinely degrades endpoint performance, and it would be correct if response times were uniformly elevated regardless of traffic pattern.
- ✗
The endpoint is misconfigured with the wrong regional endpoint.
Why it's wrong here
A wrong regional endpoint would fail outright or route incorrectly, producing consistent errors rather than a one-off cold-start delay. It is tempting because endpoint configuration errors do break deployments, and it would be correct if requests failed or hit an unintended region entirely.
- ✗
The model is too large and exceeds the instance memory.
Why it's wrong here
Memory exhaustion would cause out-of-memory failures or sustained slow inference, not a delay confined to the first request after idle hours. It is tempting because oversized models do cause deployment problems, and it would be correct if every request, not just the first, ran slowly.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.