easyMultiple Select
PMLE Practice Question: Which THREE factors should be considered when…
Which THREE factors should be considered when choosing a compute option for serving a deep learning model in production on Google Cloud? (Choose three.)
⚠ Common exam trap
The trap here is that candidates might think the training language (D) matters for serving, but Google Cloud serving infrastructure is language-agnostic as long as the model is exported in a supported format, making this a common distractor.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Integration with Vertex AI for model monitoring
Option A is correct because integration with Vertex AI enables managed model monitoring, which is essential for detecting training-serving skew, drift, and performance degradation in production deep learning deployments. Option B is correct because autoscaling capabilities allow the serving infrastructure to dynamically adjust compute resources based on traffic patterns, ensuring cost efficiency and consistent latency under variable load. Option C is correct because deep learning inference often requires GPU or TPU acceleration to meet latency and throughput requirements, and the chosen compute option must support the necessary hardware accelerators. Option D is not a relevant factor because the programming language used during training does not constrain the production serving compute option, as models are typically exported to framework-agnostic formats or served via standard runtimes. Option E is clearly irrelevant because the team's logo color has no bearing on compute selection for model serving.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Integration with Vertex AI for model monitoring
Why this is correct
Integration with Vertex AI satisfies the requirement for production serving, since Vertex AI provides managed endpoints with built-in model monitoring for drift and skew detection. This removes the need to build custom observability around the compute layer, directly addressing the operational constraint of maintaining model quality once deployed.
- ✓
Autoscaling capabilities to handle variable traffic
Why this is correct
Autoscaling directly addresses variable inference traffic, a core production constraint when serving deep learning models. GPU and TPU accelerators are costly, so scaling replicas to match demand controls spend while preserving latency during peaks. Static provisioning either wastes accelerator hours or drops requests, making autoscaling essential for the stem's production serving scenario.
- ✓
GPU or TPU requirements for model inference
Why this is correct
Inference hardware dictates throughput and cost: GPUs suit parallel tensor maths, while TPUs accelerate large matrix operations for TensorFlow models. Selecting a machine type without matching accelerator support would leave the trained model unable to serve predictions at the required latency.
- ✗
The programming language used for training
Why it's wrong here
Training language is irrelevant to serving inference; the serving framework, latency, throughput and accelerator requirements govern the compute choice. It is tempting because the training stack does constrain reproducibility, and would be correct when selecting a training platform rather than a production inference option.
- ✗
The color of the team's logo
Why it's wrong here
Logo colour has no bearing on compute selection; the factors are workload characteristics such as latency, throughput, GPU/TPU requirements and cost. It is tempting because team branding matters for dashboards and documentation, but that is a presentation concern, not a serving-infrastructure decision.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.