Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

A data science team is building a feature engineering pipeline that processes large-scale data from BigQuery daily. They need to compute aggregate features and store the results in Vertex AI Feature Store for both online serving and offline training. Which Google Cloud service is best suited for this batch computation?

⚠ Common exam trap

PMLE often tests the boundary between orchestration (Cloud Composer/Airflow) and computation (Dataflow/Beam) — candidates who pick Composer because it 'runs pipelines' miss that it does not process the data itself.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dataflow

Dataflow is Google Cloud's managed Apache Beam service, purpose-built for batch and streaming data pipelines that read from BigQuery, transform data, and write to sinks like Vertex AI Feature Store. It handles autoscaling, sharding, and windowing natively, making it the canonical choice for daily batch feature engineering at scale. Its native BigQuery and Feature Store I/O connectors mean minimal glue code.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Composer

    Why it's wrong here

    Cloud Composer orchestrates workflows but does not itself perform the large-scale aggregation; it only triggers and schedules tasks. It is tempting because it coordinates pipelines, and it would be correct when the requirement is dependency management and scheduling across multiple services rather than the computation itself.

  • ✗

    Dataproc

    Why it's wrong here

    Dataproc requires provisioning and managing a cluster, adding operational overhead for a scheduled SQL aggregation that BigQuery executes natively. It is tempting because it handles large-scale Spark and Hadoop batch workloads, which is the right choice when existing Spark jobs or custom libraries must run.

  • ✗

    Cloud Functions

    Why it's wrong here

    Cloud Functions enforces short execution timeouts and limited memory, so it cannot process daily large-scale BigQuery aggregations. It is tempting because it is serverless and event-driven, which suits lightweight, per-event transformations rather than heavy scheduled batch feature computation.

  • ✓

    Dataflow

    Why this is correct

    Dataflow runs Apache Beam pipelines that read from BigQuery, compute aggregates at scale, and write to Vertex AI Feature Store for both online serving and offline training. Its managed batch processing satisfies the daily large-scale computation requirement without provisioning servers.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.