PMLE Serving and Scaling Models Practice Question
Which Vertex AI feature allows you to reduce the size of a trained model to improve inference speed on edge devices without significant accuracy loss?
⚠ Common exam trap
Google PMLE exams often test the distinction between 'optimization' (size/speed improvements) and 'monitoring' (observability), leading candidates to confuse Model Monitoring with performance tuning because both involve 'model performance' terminology.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Vertex AI Model Optimization
Vertex AI Model Optimization is the correct feature because it provides model quantization, pruning, and distillation techniques specifically designed to reduce model size and improve inference latency on edge devices. This service applies post-training quantization (e.g., FP32 to INT8) and structured weight pruning to shrink the model footprint while maintaining accuracy within acceptable thresholds, directly addressing the need for efficient deployment on resource-constrained hardware.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Vertex AI Model Optimization
Why this is correct
Vertex AI Model Optimization applies quantisation and pruning to shrink trained models, reducing size and inference latency on constrained edge hardware. It meets the stem's requirement of faster edge inference without significant accuracy loss, unlike retraining or distillation approaches.
- ✗
Vertex AI Model Monitoring
Why it's wrong here
Model Monitoring tracks prediction drift and skew on deployed endpoints; it neither quantises nor prunes weights, so model size and edge inference speed are untouched. It is tempting because it is a deployment-time Vertex AI feature, and would be correct when detecting training-serving skew or data drift.
- ✗
Vertex AI Matching Engine
Why it's wrong here
Matching Engine performs approximate nearest-neighbour vector search for embeddings; it does not compress or shrink model artefacts, so edge inference speed is unaffected. It is tempting because it is a Vertex AI serving component, and would be correct when building a similarity search or recommendation retrieval system.
- ✗
Vertex AI Continuous Training
Why it's wrong here
Continuous Training automates retraining pipelines when new data or drift is detected; it produces new model versions rather than reducing the size of an existing one. It is tempting because it is a Vertex AI MLOps feature, and would be correct when models must be refreshed regularly on fresh data.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.