mediumMultiple Choice
PMLE Practice Question: An ML engineer needs to run batch predictions on…
An ML engineer needs to run batch predictions on tens of petabytes of data using a trained model. The data is stored in Cloud Storage. Which service should they choose?
⚠ Common exam trap
Google Cloud often tests the distinction between batch inference and data processing pipelines, so the trap here is that candidates confuse Cloud Dataflow (a data processing tool) with a batch prediction service, not realizing that Vertex AI Batch Prediction is the dedicated service for running models on large static datasets.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Vertex AI Batch Prediction
Vertex AI Batch Prediction is the correct choice because it is a managed service specifically designed for high-throughput, large-scale batch inference on data stored in Cloud Storage. It automatically handles sharding, scaling, and resource management for tens of petabytes, without requiring the engineer to manage infrastructure or write custom distributed processing code.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Dataflow with the model as a side input
Why it's wrong here
Dataflow side inputs are broadcast to every worker and held in memory, so a model loaded this way is replicated per worker rather than shared, and the pipeline still requires custom inference code. It fits enriching streaming records with small reference datasets, not petabyte-scale batch scoring.
- ✗
Cloud Dataproc running Spark ML
Why it's wrong here
Spark ML trains models within Dataproc; it does not natively load and serve an already-trained model for distributed batch inference across tens of petabytes. Dataproc suits custom Spark ETL and model training jobs. The scenario needs a managed service that applies an existing model to Cloud Storage data.
- ✗
Cloud Run with multiple revisions
Why it's wrong here
Cloud Run revisions manage traffic splitting between deployed container versions, and its request-scoped execution model cannot stream tens of petabytes from Cloud Storage. It suits serving low-latency online inference endpoints. Batch scoring at that scale needs a distributed processing engine that shards input across many workers.
- ✓
Vertex AI Batch Prediction
Why this is correct
Vertex AI Batch Prediction handles petabyte-scale batch inference directly from Cloud Storage, satisfying the tens-of-petabytes constraint. It streams data without loading it all into memory, distributing the job across managed compute, unlike online prediction endpoints, which suit low-latency single requests rather than massive offline scoring.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.