Courseiva

PMLE Scaling Prototypes into ML Models Practice Question

You need to perform a large-scale feature computation on streaming data from Pub/Sub, transforming raw events into features, and writing results to Vertex AI Feature Store for online serving. Which Google Cloud architecture is most appropriate?

⚠ Common exam trap

PMLE often tests the misconception that any compute service (Cloud Functions, Cloud Run) can handle streaming feature engineering, when in fact managed Apache Beam on Dataflow is the canonical answer for large-scale streaming transformations into Feature Store.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Dataflow streaming pipeline with Apache Beam to read from Pub/Sub, compute features, and write to Feature Store

Dataflow with Apache Beam is the purpose-built managed service for streaming ETL on Google Cloud, natively integrating with Pub/Sub as a source and Vertex AI Feature Store as a sink. Beam's windowing, watermarks, and exactly-once semantics handle the large-scale, continuous feature computation required for online serving. It provides autoscaling and managed infrastructure, which is exactly what a production streaming feature pipeline needs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use Dataproc with Spark Streaming to read from Pub/Sub and write to Feature Store

    Why it's wrong here

    Dataproc with Spark Streaming lacks native integration for writing directly to Vertex AI Feature Store’s online serving endpoints; it requires a custom connector or intermediate batch layer, introducing latency that undermines the low-latency online serving requirement. It is tempting because Spark Streaming excels at large-scale, stateful stream processing for batch-oriented feature pipelines, and would be correct if the output target were a data warehouse or filesystem rather than an online feature store.

  • ✗

    Use Cloud Functions triggered by Pub/Sub to compute features and update Feature Store

    Why it's wrong here

    Cloud Functions imposes execution timeouts and per-instance concurrency limits that break sustained large-scale streaming aggregation. It is tempting because Pub/Sub triggers are native, and it would suit low-volume, short-lived event handling rather than continuous feature computation into Feature Store.

  • ✓

    Use Dataflow streaming pipeline with Apache Beam to read from Pub/Sub, compute features, and write to Feature Store

    Why this is correct

    Dataflow runs Apache Beam, providing the horizontal scaling and windowing needed for high-volume Pub/Sub streams, and its native Feature Store sink writes computed features directly for low-latency online serving, meeting the streaming transformation and serving constraints.

  • ✗

    Use Cloud Run to consume Pub/Sub messages and update Feature Store via a service

    Why it's wrong here

    Cloud Run's request-scoped instances lack the sustained, ordered stream processing and windowing that Dataflow provides for large-scale feature computation. It is tempting as a serverless Pub/Sub consumer, and would be correct for lightweight event handling rather than continuous feature pipelines.

About these practice questions

One of 775 original PMLE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

Same concept, more angles

4 more ways this is tested on PMLE

These questions test the same concept from different angles. Work through them to make sure you can recognise it however the exam phrases it.

Variation 1. A data science team is building a feature engineering pipeline that processes large-scale data from BigQuery daily. They need to compute aggregate features and store the results in Vertex AI Feature Store for both online serving and offline training. Which Google Cloud service is best suited for this batch computation?

medium
  • A.Cloud Composer
  • B.Dataproc
  • C.Cloud Functions
  • ✓ D.Dataflow

Why D: Dataflow is Google Cloud's managed Apache Beam service, purpose-built for batch and streaming data pipelines that read from BigQuery, transform data, and write to sinks like Vertex AI Feature Store. It handles autoscaling, sharding, and windowing natively, making it the canonical choice for daily batch feature engineering at scale. Its native BigQuery and Feature Store I/O connectors mean minimal glue code.

Variation 2. A data engineering team needs to compute rolling window features (7-day average, 30-day sum) from a high-volume stream of e-commerce events stored in BigQuery. They must output the features to Vertex AI Feature Store for online serving. Which approach is MOST cost-effective and scalable?

hard
  • A.Use Cloud Composer (Airflow) with a daily DAG to run SQL queries on BigQuery and export results
  • B.Schedule a query in BigQuery using scheduled queries and export results to Feature Store
  • C.Use Cloud Functions triggered by Pub/Sub to compute features on the fly
  • ✓ D.Use Dataflow with Apache Beam, reading from BigQuery, computing windowed aggregations, and writing to Vertex AI Feature Store

Why D: Dataflow (Apache Beam) is ideal for processing large-scale batch and streaming data. It can read from BigQuery, perform windowing computations, and write to Feature Store's online store. Cloud Functions have timeouts, Cloud Composer is not optimal for streaming, and BigQuery scheduled queries are not designed for streaming-feature computation.

Variation 3. You are building a machine learning pipeline on Google Cloud. You need to perform feature engineering on large datasets stored in BigQuery and store the resulting features in Vertex AI Feature Store for both online and offline use. Which TWO Google Cloud services should you use?

easy
  • A.Cloud Functions
  • ✓ B.Dataflow
  • C.BigQuery ML
  • ✓ D.Dataproc
  • E.Cloud Build

Why B: Dataflow can read from BigQuery, compute features via Apache Beam, and write to Feature Store. Alternatively, Dataproc can also do this but Dataflow is more serverless. Cloud Functions is not suitable for large-scale. Cloud Build is for CI/CD. BigQuery ML is for in-database ML.

Variation 4. A data engineer wants to compute feature aggregates over a large dataset stored in BigQuery and write the results to Vertex AI Feature Store. The pipeline must handle both batch and streaming data. Which Google Cloud service should they use?

easy
  • A.BigQuery scheduled queries
  • B.Cloud Functions triggered by Pub/Sub
  • C.Cloud Dataproc with Spark
  • ✓ D.Cloud Dataflow with Apache Beam

Why D: Cloud Dataflow with Apache Beam is a unified stream and batch data processing service. It can read from BigQuery, compute aggregates, and write to Vertex AI Feature Store, handling both batch and streaming data with the same pipeline code. This makes it the ideal choice for the requirement.

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.