Courseiva
easyMultiple Select

Best Practices for Vertex AI Model Deployment

Which TWO are best practices for deploying models to Vertex AI Prediction? (Choose 2.)

Quick Answer

The answer is to use a dedicated service account with minimal permissions for the endpoint and to leverage version aliases for easy rollback. These are best practices for deploying models to Vertex AI because a dedicated, least-privilege service account enforces the principle of least privilege, limiting the blast radius if credentials are compromised, while version aliases allow you to point traffic to a specific model version without changing the endpoint configuration, enabling seamless rollbacks and canary deployments. On the Google Professional Machine Learning Engineer exam, this question tests your understanding of secure and resilient deployment patterns, often appearing as a trap where you must distinguish between security hygiene and operational convenience—common distractors include logging all inputs (which risks PII exposure) or insisting on identical environments (which is impractical). A helpful memory tip: think “lock it down, then alias it out”—secure the endpoint with a minimal service account, then manage versions with aliases for safe updates.

⚠ Common exam trap

PMLE often tests whether candidates over-apply 'log everything' as a best practice — the trap is picking option B, which sounds like good auditing but actually violates privacy and cost best practices.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Monitor prediction latency and error rates with Cloud Monitoring alerts.

Option A is correct because Vertex AI Prediction exposes built-in metrics such as prediction latency, request count, and error rates to Cloud Monitoring, and configuring alerting policies on these metrics is a recommended operational best practice to detect model degradation or endpoint failures early. Option C is correct because the endpoint and model deployment should run under a dedicated service account scoped with least-privilege IAM roles (e.g., only the permissions needed to read model artifacts from Cloud Storage and write logs), which limits blast radius if the endpoint is compromised. Option B is not a best practice because logging every raw prediction input and output can expose sensitive data (PII), incur high logging costs, and is not required for auditing; sampling or redaction is preferred. Option D is incorrect because training and serving environments are typically decoupled, and Vertex AI supports deploying models across environments as long as the serving container and dependencies are compatible; forcing identical environments is not a deployment best practice. Option E is incorrect because relying solely on the 'default' alias for all deployments removes version control and makes rollback or A/B testing harder; explicit, immutable model versions or dedicated aliases should be used.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Monitor prediction latency and error rates with Cloud Monitoring alerts.

    Why this is correct

    Cloud Monitoring alerts on prediction latency and error rates surface degradation before users are broadly affected, enabling timely rollback or scaling. This observability practise is essential for maintaining reliable Vertex AI Prediction deployments in production.

  • ✗

    Log all raw prediction inputs and outputs for every request for auditing.

    Why it's wrong here

    Logging every raw prediction input and output indiscriminately captures sensitive customer data, inflating storage cost and breaching privacy obligations. Auditing requires sampling or redaction, not blanket capture. Full logging suits debugging a specific failing request, not routine production deployment of a Vertex AI model.

  • ✓

    Use a dedicated service account with minimal permissions for the endpoint.

    Why this is correct

    A dedicated service account scoped to only the permissions the endpoint needs enforces least privilege, limiting blast radius if the container is compromised. This directly satisfies the security best-practise requirement for Vertex AI Prediction deployments, rather than reusing broad default credentials.

  • ✗

    Always deploy the model in the same environment as training to avoid incompatibility.

    Why it's wrong here

    Deploying in the training environment conflates infrastructure with compatibility; Vertex AI Prediction serves models from a managed container regardless of where training ran. Pinning deployment to the training host prevents independent scaling and reproducible serving. It would suit tightly coupled on-premises inference, not managed Vertex AI endpoints.

  • ✗

    Use the default model version alias 'default' for all deployments to simplify updates.

    Why it's wrong here

    Reusing the 'default' alias for every deployment overwrites the pointer, so rollback and traffic splitting between versions become impossible. Aliases exist to label specific versions for staged rollout or canary testing. Correct practise assigns distinct aliases per version and moves traffic deliberately between them.

Quick reference

AAA Protocol Comparison

ProtocolPort(s)EncryptionTransportPrimary Use
RADIUS1812 / 1813Password onlyUDPNetwork access control
TACACS+49Full packetTCPDevice administration
Diameter3868Full sessionTCP / SCTPCarrier / mobile networks
802.1X—EAP-basedLayer 2Port-based access control

TACACS+ encrypts the entire packet; RADIUS only encrypts the password field — a key exam distinction.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

1 more way this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. Which TWO options are best practices for reducing model serving latency on Vertex AI Endpoints? (Choose two.)

easy
  • A.Use a larger machine type with more memory
  • ✓ B.Optimize the model using quantization or pruning
  • ✓ C.Deploy the model in the same region as the clients
  • D.Use batch prediction instead of online prediction
  • E.Enable model caching at the endpoint

Why B: Option B is correct because quantization (e.g., reducing weights from FP32 to INT8) and pruning (removing redundant parameters) shrink the model size and reduce the compute required per inference, directly lowering prediction latency on Vertex AI Endpoints. Option C is correct because deploying the model in the same region as the clients minimizes network round-trip time, which is a significant component of end-to-end serving latency for online predictions. Option A is not a best practice for latency specifically: a larger machine with more memory may help with throughput or memory-bound models, but it does not inherently reduce per-request latency and increases cost. Option D is wrong because batch prediction is an asynchronous, offline mode that does not serve real-time requests and is not a latency optimization for online endpoints. Option E is not a supported Vertex AI Endpoint feature; there is no endpoint-level 'model caching' toggle that reduces serving latency.

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.