PMLE Scaling Prototypes into ML Models Practice Question
You are using tf.Transform to preprocess data at scale. Which TWO services are required to run tf.Transform on Google Cloud? (Choose 2)
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dataflow
tf.Transform requires Apache Beam for execution, which on GCP is typically run on Dataflow. The processed data and transform artifacts are stored in Cloud Storage.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Functions
Why it's wrong here
Cloud Functions executes short, event-driven snippets and cannot run the Apache Beam pipeline that tf.Transform requires for distributed preprocessing. It is tempting because it handles lightweight serverless tasks, and would suit triggering a pipeline or reacting to a file upload, but tf.Transform needs Dataflow for execution and Cloud Storage for staging artefacts.
- ✓
Dataflow
Why this is correct
Dataflow is required to execute the Apache Beam pipeline that tf.Transform generates, running the preprocessing at scale. It satisfies the stem's requirement by providing the distributed runner that applies the transform over large datasets, complementing the service that analyses and materialises the transform artefacts.
- ✓
Cloud Storage
Why this is correct
tf.Transform runs Apache Beam pipelines on Dataflow, which reads and writes artefacts through Cloud Storage buckets. Cloud Storage satisfies the stem's requirement by holding the input data and the exported transform function, so the pipeline can materialise the preprocessing graph at scale.
- ✗
Vertex AI Training
Why it's wrong here
Vertex AI Training runs custom training jobs, not the Apache Beam pipeline that tf.Transform requires for its preprocessing graph. It is tempting because Vertex AI Training is the standard place to run TensorFlow workloads on Google Cloud, and it would be the correct choice when the task is training or tuning a model rather than executing a Beam transformation.
- ✗
BigQuery
Why it's wrong here
BigQuery is a warehouse, not a required execution service for tf.Transform; the pipeline runs on Dataflow and reads and writes through Cloud Storage. BigQuery is tempting because it stores analytical data, and would be correct when the source dataset itself resides in BigQuery tables.
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.