Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You are using tf.Transform to preprocess data at scale. Which TWO services are required to run tf.Transform on Google Cloud? (Choose 2)

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dataflow

tf.Transform requires Apache Beam for execution, which on GCP is typically run on Dataflow. The processed data and transform artifacts are stored in Cloud Storage.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Functions

    Why it's wrong here

    Cloud Functions executes short, event-driven snippets and cannot run the Apache Beam pipeline that tf.Transform requires for distributed preprocessing. It is tempting because it handles lightweight serverless tasks, and would suit triggering a pipeline or reacting to a file upload, but tf.Transform needs Dataflow for execution and Cloud Storage for staging artefacts.

  • ✓

    Dataflow

    Why this is correct

    Dataflow is required to execute the Apache Beam pipeline that tf.Transform generates, running the preprocessing at scale. It satisfies the stem's requirement by providing the distributed runner that applies the transform over large datasets, complementing the service that analyses and materialises the transform artefacts.

  • ✓

    Cloud Storage

    Why this is correct

    tf.Transform runs Apache Beam pipelines on Dataflow, which reads and writes artefacts through Cloud Storage buckets. Cloud Storage satisfies the stem's requirement by holding the input data and the exported transform function, so the pipeline can materialise the preprocessing graph at scale.

  • ✗

    Vertex AI Training

    Why it's wrong here

    Vertex AI Training runs custom training jobs, not the Apache Beam pipeline that tf.Transform requires for its preprocessing graph. It is tempting because Vertex AI Training is the standard place to run TensorFlow workloads on Google Cloud, and it would be the correct choice when the task is training or tuning a model rather than executing a Beam transformation.

  • ✗

    BigQuery

    Why it's wrong here

    BigQuery is a warehouse, not a required execution service for tf.Transform; the pipeline runs on Dataflow and reads and writes through Cloud Storage. BigQuery is tempting because it stores analytical data, and would be correct when the source dataset itself resides in BigQuery tables.

About these practice questions

This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.