easyMultiple Select
Best Practices for Vertex AI Model Deployment
Which TWO are best practices for deploying models to Vertex AI Prediction? (Choose 2.)
Quick Answer
The answer is to use a dedicated service account with minimal permissions for the endpoint and to leverage version aliases for easy rollback. These are best practices for deploying models to Vertex AI because a dedicated, least-privilege service account enforces the principle of least privilege, limiting the blast radius if credentials are compromised, while version aliases allow you to point traffic to a specific model version without changing the endpoint configuration, enabling seamless rollbacks and canary deployments. On the Google Professional Machine Learning Engineer exam, this question tests your understanding of secure and resilient deployment patterns, often appearing as a trap where you must distinguish between security hygiene and operational convenience—common distractors include logging all inputs (which risks PII exposure) or insisting on identical environments (which is impractical). A helpful memory tip: think “lock it down, then alias it out”—secure the endpoint with a minimal service account, then manage versions with aliases for safe updates.
⚠ Common exam trap
PMLE often tests whether candidates over-apply 'log everything' as a best practice — the trap is picking option B, which sounds like good auditing but actually violates privacy and cost best practices.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Monitor prediction latency and error rates with Cloud Monitoring alerts.
Option A is correct because Vertex AI Prediction exposes built-in metrics such as prediction latency, request count, and error rates to Cloud Monitoring, and configuring alerting policies on these metrics is a recommended operational best practice to detect model degradation or endpoint failures early. Option C is correct because the endpoint and model deployment should run under a dedicated service account scoped with least-privilege IAM roles (e.g., only the permissions needed to read model artifacts from Cloud Storage and write logs), which limits blast radius if the endpoint is compromised. Option B is not a best practice because logging every raw prediction input and output can expose sensitive data (PII), incur high logging costs, and is not required for auditing; sampling or redaction is preferred. Option D is incorrect because training and serving environments are typically decoupled, and Vertex AI supports deploying models across environments as long as the serving container and dependencies are compatible; forcing identical environments is not a deployment best practice. Option E is incorrect because relying solely on the 'default' alias for all deployments removes version control and makes rollback or A/B testing harder; explicit, immutable model versions or dedicated aliases should be used.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Monitor prediction latency and error rates with Cloud Monitoring alerts.
Why this is correct
Cloud Monitoring alerts on prediction latency and error rates surface degradation before users are broadly affected, enabling timely rollback or scaling. This observability practise is essential for maintaining reliable Vertex AI Prediction deployments in production.
- ✗
Log all raw prediction inputs and outputs for every request for auditing.
Why it's wrong here
Logging every raw prediction input and output indiscriminately captures sensitive customer data, inflating storage cost and breaching privacy obligations. Auditing requires sampling or redaction, not blanket capture. Full logging suits debugging a specific failing request, not routine production deployment of a Vertex AI model.
- ✓
Use a dedicated service account with minimal permissions for the endpoint.
Why this is correct
A dedicated service account scoped to only the permissions the endpoint needs enforces least privilege, limiting blast radius if the container is compromised. This directly satisfies the security best-practise requirement for Vertex AI Prediction deployments, rather than reusing broad default credentials.
- ✗
Always deploy the model in the same environment as training to avoid incompatibility.
Why it's wrong here
Deploying in the training environment conflates infrastructure with compatibility; Vertex AI Prediction serves models from a managed container regardless of where training ran. Pinning deployment to the training host prevents independent scaling and reproducible serving. It would suit tightly coupled on-premises inference, not managed Vertex AI endpoints.
- ✗
Use the default model version alias 'default' for all deployments to simplify updates.
Why it's wrong here
Reusing the 'default' alias for every deployment overwrites the pointer, so rollback and traffic splitting between versions become impossible. Aliases exist to label specific versions for staged rollout or canary testing. Correct practise assigns distinct aliases per version and moves traffic deliberately between them.
Quick reference
AAA Protocol Comparison
| Protocol | Port(s) | Encryption | Transport | Primary Use |
|---|---|---|---|---|
| RADIUS | 1812 / 1813 | Password only | UDP | Network access control |
| TACACS+ | 49 | Full packet | TCP | Device administration |
| Diameter | 3868 | Full session | TCP / SCTP | Carrier / mobile networks |
| 802.1X | — | EAP-based | Layer 2 | Port-based access control |
TACACS+ encrypts the entire packet; RADIUS only encrypts the password field — a key exam distinction.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
Same concept, more angles
1 more way this is tested on PMLE
These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.
Variation 1. Which TWO options are best practices for reducing model serving latency on Vertex AI Endpoints? (Choose two.)
easy- A.Use a larger machine type with more memory
- ✓ B.Optimize the model using quantization or pruning
- ✓ C.Deploy the model in the same region as the clients
- D.Use batch prediction instead of online prediction
- E.Enable model caching at the endpoint
Why B: Option B is correct because quantization (e.g., reducing weights from FP32 to INT8) and pruning (removing redundant parameters) shrink the model size and reduce the compute required per inference, directly lowering prediction latency on Vertex AI Endpoints. Option C is correct because deploying the model in the same region as the clients minimizes network round-trip time, which is a significant component of end-to-end serving latency for online predictions. Option A is not a best practice for latency specifically: a larger machine with more memory may help with throughput or memory-bound models, but it does not inherently reduce per-request latency and increases cost. Option D is wrong because batch prediction is an asynchronous, offline mode that does not serve real-time requests and is not a latency optimization for online endpoints. Option E is not a supported Vertex AI Endpoint feature; there is no endpoint-level 'model caching' toggle that reduces serving latency.
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.