PMLE Scaling Prototypes into ML Models Practice Question
A data science team is building a feature engineering pipeline that processes large-scale data from BigQuery daily. They need to compute aggregate features and store the results in Vertex AI Feature Store for both online serving and offline training. Which Google Cloud service is best suited for this batch computation?
⚠ Common exam trap
PMLE often tests the boundary between orchestration (Cloud Composer/Airflow) and computation (Dataflow/Beam) — candidates who pick Composer because it 'runs pipelines' miss that it does not process the data itself.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dataflow
Dataflow is Google Cloud's managed Apache Beam service, purpose-built for batch and streaming data pipelines that read from BigQuery, transform data, and write to sinks like Vertex AI Feature Store. It handles autoscaling, sharding, and windowing natively, making it the canonical choice for daily batch feature engineering at scale. Its native BigQuery and Feature Store I/O connectors mean minimal glue code.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Composer
Why it's wrong here
Cloud Composer orchestrates workflows but does not itself perform the large-scale aggregation; it only triggers and schedules tasks. It is tempting because it coordinates pipelines, and it would be correct when the requirement is dependency management and scheduling across multiple services rather than the computation itself.
- ✗
Dataproc
Why it's wrong here
Dataproc requires provisioning and managing a cluster, adding operational overhead for a scheduled SQL aggregation that BigQuery executes natively. It is tempting because it handles large-scale Spark and Hadoop batch workloads, which is the right choice when existing Spark jobs or custom libraries must run.
- ✗
Cloud Functions
Why it's wrong here
Cloud Functions enforces short execution timeouts and limited memory, so it cannot process daily large-scale BigQuery aggregations. It is tempting because it is serverless and event-driven, which suits lightweight, per-event transformations rather than heavy scheduled batch feature computation.
- ✓
Dataflow
Why this is correct
Dataflow runs Apache Beam pipelines that read from BigQuery, compute aggregates at scale, and write to Vertex AI Feature Store for both online serving and offline training. Its managed batch processing satisfies the daily large-scale computation requirement without provisioning servers.
Go deeper
Related to this question
About these practice questions
One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.