Courseiva

PMLE Serving and Scaling Models Practice Question

A machine learning engineer needs to run batch predictions on 50 TB of data stored in BigQuery using a Vertex AI model. The model is a custom container. What is the most efficient way to set up the batch prediction job?

⚠ Common exam trap

PMLE often tests whether candidates default to exporting data to GCS out of habit — the trap is missing that Vertex AI batch prediction supports BigQuery natively, making export steps unnecessary and inefficient.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Create a Vertex AI batch prediction job with BigQuery source and BigQuery destination.

Vertex AI batch prediction natively supports BigQuery as both input source and output destination, allowing the service to read the 50 TB directly from BigQuery and write predictions back without exporting data. This avoids data movement, leverages BigQuery's scalability, and is the most efficient, fully managed approach for large-scale batch inference with a custom container.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Create a Vertex AI batch prediction job with BigQuery source and BigQuery destination.

    Why this is correct

    Native BigQuery source and destination lets Vertex AI read and write directly without exporting 50 TB to Cloud Storage, avoiding costly intermediate copies. This satisfies the efficiency constraint by keeping data in place and streaming results back to BigQuery.

  • ✗

    Use Dataflow to process the data and call the model via Vertex AI online prediction.

    Why it's wrong here

    Calling online prediction endpoints via Dataflow introduces significant network latency and overhead for each record, failing to leverage the high-throughput distributed reading capabilities of Vertex AI Batch Prediction jobs designed specifically for large-scale BigQuery datasets. This approach is intended for real-time streaming pipelines where data must be processed as it arrives, making it a valid pattern for low-latency transformations rather than massive batch processing tasks.

  • ✗

    Export BigQuery data to CSV in GCS, then create a batch prediction job with GCS source.

    Why it's wrong here

    Exporting 50 TB to CSV in Cloud Storage adds a costly serialisation and transfer stage before prediction, whereas a BigQuery source lets Vertex AI read the table directly. It is tempting because GCS input suits small datasets or models needing files, but here it multiplies cost and latency.

  • ✗

    Create a Cloud Function to iterate over BigQuery rows and call the endpoint.

    Why it's wrong here

    A Cloud Function iterating row by row cannot scale to 50 TB and would time out, whereas a native batch prediction job processes the BigQuery table in parallel. It is tempting because per-row endpoint calls suit small real-time scoring, not bulk offline inference at this volume.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.