Courseiva
hardMultiple Choice

PDE Vertex AI Batch Prediction Practice Question

A company needs to serve predictions for a model that runs an expensive computation on each request. The model is used by a batch job that processes millions of records each night, and also by a real-time API for a few thousand queries per hour. Which prediction strategy minimizes cost and latency for both use cases?

⚠ Common exam trap

Google often tests the misconception that a single Vertex AI endpoint can handle both batch and online workloads efficiently, but the trap is that batch and online have fundamentally different latency and throughput requirements, and using the same infrastructure for both leads to cost or performance penalties.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Vertex AI batch prediction for the nightly job and a separate online endpoint with auto-scaling for the real-time API.

It separates the batch and online workloads to optimize cost and latency. Vertex AI batch prediction is designed for high-throughput, asynchronous processing of large datasets at lower cost, while a separate online endpoint with auto-scaling ensures low-latency responses for real-time API queries by scaling resources based on demand. This avoids over-provisioning for the batch job and prevents the batch workload from interfering with the latency-sensitive API.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy two identical models, one on a Compute Engine VM for batch, one on Vertex AI for online, and synchronize updates.

    Why it's wrong here

    Two independently deployed models duplicate serving infrastructure and require synchronising artefacts, risking version drift between batch and online predictions. Separate deployments suit isolation or differing hardware needs; here a single model with batch prediction plus an online endpoint serves both workloads without duplication.

  • ✓

    Use Vertex AI batch prediction for the nightly job and a separate online endpoint with auto-scaling for the real-time API.

    Why this is correct

    Batch prediction processes the nightly millions of records asynchronously at far lower cost than an always-on endpoint, while a separate auto-scaling online endpoint serves the few thousand real-time queries with low latency. This split matches each workload's cost and latency profile.

  • ✗

    Use Vertex AI batch prediction for both workloads.

    Why it's wrong here

    Batch prediction queues records and returns results to storage, so the real-time API's few thousand queries per hour would wait for job completion rather than return synchronously. Batch prediction is correct for the nightly millions of records, but the interactive workload needs an online endpoint.

  • ✗

    Use a single online Vertex AI endpoint with auto-scaling to handle both workloads.

    Why it's wrong here

    An online endpoint bills for provisioned compute continuously, so millions of nightly batch records incur idle-capacity cost and per-request latency. Online endpoints suit low-volume, low-latency interactive serving; the batch workload belongs in Vertex AI batch prediction, which scales down between jobs.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.