Courseiva
mediumMultiple Choice

PDE Practice Question: A team uses Vertex AI AutoML Tables to train a…

A team uses Vertex AI AutoML Tables to train a model. They need to deploy the model for real-time predictions with high availability. Which deployment configuration should they use?

⚠ Common exam trap

It's easy for candidates to confuse batch prediction with real-time serving, or assume that a single replica is sufficient for high availability, not realizing that high availability requires redundancy and automatic scaling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Deploy to a Vertex AI Endpoint with multiple replicas and auto-scaling

For real-time predictions with high availability, you need a deployment that can handle traffic spikes and failover. Deploying to a Vertex AI Endpoint with multiple replicas and auto-scaling ensures that the model is served from multiple instances, providing redundancy and the ability to scale up or down based on demand. This configuration meets the high-availability requirement by distributing load and automatically recovering from instance failures.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Export as a Cloud Function

    Why it's wrong here

    Cloud Functions cannot host AutoML Tables models; they run event-driven code, not model servers, so no real-time prediction endpoint or replica-based availability exists. It is tempting because serverless functions suit lightweight inference for custom code, but AutoML Tables requires deployment to a Vertex AI Endpoint.

  • ✗

    Deploy to a Vertex AI Endpoint with 1 replica

    Why it's wrong here

    A single replica provides no redundancy, so a zone or instance failure takes predictions offline, failing the high-availability requirement. Multiple replicas behind a Vertex AI Endpoint distribute traffic and survive failures. One replica would suffice only for development or non-critical, latency-tolerant workloads where downtime is acceptable.

  • ✗

    Use a Vertex AI Batch Prediction job

    Why it's wrong here

    Batch Prediction processes large datasets asynchronously into Cloud Storage or BigQuery, returning results after job completion with no serving endpoint, so real-time requests cannot be answered. It is tempting because it reuses the same trained model cheaply, and would be correct for scheduled bulk scoring where latency is irrelevant.

  • ✓

    Deploy to a Vertex AI Endpoint with multiple replicas and auto-scaling

    Why this is correct

    Deploying to a Vertex AI Endpoint with multiple replicas and auto-scaling directly satisfies the high-availability requirement: replicas distribute traffic across instances, so a single node failure does not take the service down, while auto-scaling absorbs real-time prediction load spikes without manual intervention.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.