Courseiva
mediumMultiple Choice

PMLE Practice Question: An ML engineer needs to run batch predictions on…

An ML engineer needs to run batch predictions on tens of petabytes of data using a trained model. The data is stored in Cloud Storage. Which service should they choose?

⚠ Common exam trap

Google Cloud often tests the distinction between batch inference and data processing pipelines, so the trap here is that candidates confuse Cloud Dataflow (a data processing tool) with a batch prediction service, not realizing that Vertex AI Batch Prediction is the dedicated service for running models on large static datasets.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Vertex AI Batch Prediction

Vertex AI Batch Prediction is the correct choice because it is a managed service specifically designed for high-throughput, large-scale batch inference on data stored in Cloud Storage. It automatically handles sharding, scaling, and resource management for tens of petabytes, without requiring the engineer to manage infrastructure or write custom distributed processing code.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Dataflow with the model as a side input

    Why it's wrong here

    Dataflow side inputs are broadcast to every worker and held in memory, so a model loaded this way is replicated per worker rather than shared, and the pipeline still requires custom inference code. It fits enriching streaming records with small reference datasets, not petabyte-scale batch scoring.

  • ✗

    Cloud Dataproc running Spark ML

    Why it's wrong here

    Spark ML trains models within Dataproc; it does not natively load and serve an already-trained model for distributed batch inference across tens of petabytes. Dataproc suits custom Spark ETL and model training jobs. The scenario needs a managed service that applies an existing model to Cloud Storage data.

  • ✗

    Cloud Run with multiple revisions

    Why it's wrong here

    Cloud Run revisions manage traffic splitting between deployed container versions, and its request-scoped execution model cannot stream tens of petabytes from Cloud Storage. It suits serving low-latency online inference endpoints. Batch scoring at that scale needs a distributed processing engine that shards input across many workers.

  • ✓

    Vertex AI Batch Prediction

    Why this is correct

    Vertex AI Batch Prediction handles petabyte-scale batch inference directly from Cloud Storage, satisfying the tens-of-petabytes constraint. It streams data without loading it all into memory, distributing the job across managed compute, unlike online prediction endpoints, which suit low-latency single requests rather than massive offline scoring.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.