mediumMultiple Choice
PDE Practice Question: A team uses Vertex AI AutoML Tables to train a…
A team uses Vertex AI AutoML Tables to train a model. They need to deploy the model for real-time predictions with high availability. Which deployment configuration should they use?
⚠ Common exam trap
It's easy for candidates to confuse batch prediction with real-time serving, or assume that a single replica is sufficient for high availability, not realizing that high availability requires redundancy and automatic scaling.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Deploy to a Vertex AI Endpoint with multiple replicas and auto-scaling
For real-time predictions with high availability, you need a deployment that can handle traffic spikes and failover. Deploying to a Vertex AI Endpoint with multiple replicas and auto-scaling ensures that the model is served from multiple instances, providing redundancy and the ability to scale up or down based on demand. This configuration meets the high-availability requirement by distributing load and automatically recovering from instance failures.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Export as a Cloud Function
Why it's wrong here
Cloud Functions cannot host AutoML Tables models; they run event-driven code, not model servers, so no real-time prediction endpoint or replica-based availability exists. It is tempting because serverless functions suit lightweight inference for custom code, but AutoML Tables requires deployment to a Vertex AI Endpoint.
- ✗
Deploy to a Vertex AI Endpoint with 1 replica
Why it's wrong here
A single replica provides no redundancy, so a zone or instance failure takes predictions offline, failing the high-availability requirement. Multiple replicas behind a Vertex AI Endpoint distribute traffic and survive failures. One replica would suffice only for development or non-critical, latency-tolerant workloads where downtime is acceptable.
- ✗
Use a Vertex AI Batch Prediction job
Why it's wrong here
Batch Prediction processes large datasets asynchronously into Cloud Storage or BigQuery, returning results after job completion with no serving endpoint, so real-time requests cannot be answered. It is tempting because it reuses the same trained model cheaply, and would be correct for scheduled bulk scoring where latency is irrelevant.
- ✓
Deploy to a Vertex AI Endpoint with multiple replicas and auto-scaling
Why this is correct
Deploying to a Vertex AI Endpoint with multiple replicas and auto-scaling directly satisfies the high-availability requirement: replicas distribute traffic across instances, so a single node failure does not take the service down, while auto-scaling absorbs real-time prediction load spikes without manual intervention.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.