Courseiva
Serving and Scaling Models →mediumMultiple Choice

PMLE Serving and Scaling Models Practice Question

You are using Vertex AI batch prediction and your model requires preprocessing that involves joining two BigQuery tables. The preprocessing logic is complex and must be done before inference. How should you design the pipeline?

⚠ Common exam trap

Google often tests the misconception that batch prediction can handle live data transformations within the prediction container, but the correct design is to preprocess data in a separate, scalable data processing service like Dataflow before feeding it to batch prediction.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Dataflow to read from both BigQuery tables, perform the join and preprocessing, write the results to GCS, then run Vertex AI batch prediction with GCS source.

Dataflow (Apache Beam) is designed for complex, stateful data processing like joining two BigQuery tables and performing custom preprocessing. It can read from BigQuery, execute the join logic, and write the preprocessed results to Cloud Storage (GCS). Vertex AI batch prediction then reads the preprocessed data from GCS, which is the recommended pattern for non-trivial transformations before inference, as it decouples preprocessing from prediction and avoids resource contention.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Write a Cloud Composer workflow that runs the preprocessing and then triggers the batch prediction job.

    Why it's wrong here

    Cloud Composer orchestrates the join, but the stem requires the complex preprocessing inside the batch prediction pipeline itself, so the prediction job receives unjoined data unless Composer also writes the joined output. It tempts because Composer is the standard tool for scheduling multi-step BigQuery workflows.

  • ✓

    Use Dataflow to read from both BigQuery tables, perform the join and preprocessing, write the results to GCS, then run Vertex AI batch prediction with GCS source.

    Why this is correct

    Dataflow performs the complex two-table BigQuery join and preprocessing, writing results to Cloud Storage, which Vertex AI batch prediction then reads as its source. This satisfies the constraint that preprocessing must complete before inference, since batch prediction cannot join BigQuery tables itself.

  • ✗

    Use Vertex AI batch prediction with a custom container that includes logic to read and join tables on the fly.

    Why it's wrong here

    A custom container runs preprocessing at inference time, but joining two BigQuery tables from inside the container requires BigQuery client credentials and network access, which batch prediction containers do not provide by default. It tempts because custom containers are correct when preprocessing must ship with the model.

  • ✗

    Use BigQuery to create a materialized view that joins the tables and directly use that as the batch prediction source.

    Why it's wrong here

    A materialised view only stores the join, not the complex preprocessing logic, so the batch prediction job still receives untransformed features. It tempts because materialised views are the correct choice when the required transformation is purely a reusable SQL join with no additional feature engineering.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.