Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You need to preprocess a large dataset (terabytes) for training a TensorFlow model. The preprocessing includes scaling and bucketizing features, and the same transformations must be applied during serving. Which tool should you use?

⚠ Common exam trap

PMLE often tests the distinction between tools that can handle large-scale preprocessing and those that are meant for feature storage or model training, so candidates might incorrectly choose Vertex AI Feature Store or BigQuery ML thinking they handle preprocessing, but the key is the need for consistent transformations during serving with TensorFlow.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dataflow with Apache Beam and tf.Transform

tf.Transform is a library built on Apache Beam that allows you to define preprocessing functions (like scaling and bucketizing) once and then apply them consistently in both training and serving. Dataflow provides a scalable, fully managed runner for Apache Beam pipelines, making it ideal for processing terabyte-scale datasets. This combination ensures that the exact same transformations are used during training and serving, avoiding training-serving skew.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Dataflow with Apache Beam and tf.Transform

    Why this is correct

    Dataflow with Apache Beam and tf.Transform applies identical scaling and bucketising transformations across terabyte-scale preprocessing and serving. tf.Transform exports the transform graph so training and prediction use consistent feature engineering, avoiding training-serving skew that separate pipelines would introduce.

  • ✗

    Dataproc with Spark ML

    Why it's wrong here

    Spark ML transformations are not TensorFlow ops, so the training graph and serving graph would diverge unless logic is reimplemented. Dataproc suits large-scale batch ETL in Spark, and would be correct when serving uses the same Spark pipeline rather than a TensorFlow model.

  • ✗

    Vertex AI Feature Store

    Why it's wrong here

    Feature Store serves precomputed feature values for online and offline retrieval; it does not perform scaling or bucketizing transformations on raw terabytes. It is tempting because it centralises features for training and serving, and would be correct when features are already computed and only need consistent storage.

  • ✗

    BigQuery ML

    Why it's wrong here

    BigQuery ML trains and predicts inside BigQuery using SQL, so its transformations cannot be embedded in a TensorFlow serving graph. It appeals for in-warehouse modelling without data movement, and would be correct when the model itself is trained and served through BigQuery ML.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.