PMLE Scaling Prototypes into ML Models Practice Question
You need to preprocess a large dataset (terabytes) for training a TensorFlow model. The preprocessing includes scaling and bucketizing features, and the same transformations must be applied during serving. Which tool should you use?
⚠ Common exam trap
PMLE often tests the distinction between tools that can handle large-scale preprocessing and those that are meant for feature storage or model training, so candidates might incorrectly choose Vertex AI Feature Store or BigQuery ML thinking they handle preprocessing, but the key is the need for consistent transformations during serving with TensorFlow.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dataflow with Apache Beam and tf.Transform
tf.Transform is a library built on Apache Beam that allows you to define preprocessing functions (like scaling and bucketizing) once and then apply them consistently in both training and serving. Dataflow provides a scalable, fully managed runner for Apache Beam pipelines, making it ideal for processing terabyte-scale datasets. This combination ensures that the exact same transformations are used during training and serving, avoiding training-serving skew.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Dataflow with Apache Beam and tf.Transform
Why this is correct
Dataflow with Apache Beam and tf.Transform applies identical scaling and bucketising transformations across terabyte-scale preprocessing and serving. tf.Transform exports the transform graph so training and prediction use consistent feature engineering, avoiding training-serving skew that separate pipelines would introduce.
- ✗
Dataproc with Spark ML
Why it's wrong here
Spark ML transformations are not TensorFlow ops, so the training graph and serving graph would diverge unless logic is reimplemented. Dataproc suits large-scale batch ETL in Spark, and would be correct when serving uses the same Spark pipeline rather than a TensorFlow model.
- ✗
Vertex AI Feature Store
Why it's wrong here
Feature Store serves precomputed feature values for online and offline retrieval; it does not perform scaling or bucketizing transformations on raw terabytes. It is tempting because it centralises features for training and serving, and would be correct when features are already computed and only need consistent storage.
- ✗
BigQuery ML
Why it's wrong here
BigQuery ML trains and predicts inside BigQuery using SQL, so its transformations cannot be embedded in a TensorFlow serving graph. It appeals for in-warehouse modelling without data movement, and would be correct when the model itself is trained and served through BigQuery ML.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.