PMLE Scaling Prototypes into ML Models Practice Question
A data engineering team needs to compute rolling window features (7-day average, 30-day sum) from a high-volume stream of e-commerce events stored in BigQuery. They must output the features to Vertex AI Feature Store for online serving. Which approach is MOST cost-effective and scalable?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Dataflow with Apache Beam, reading from BigQuery, computing windowed aggregations, and writing to Vertex AI Feature Store
Dataflow (Apache Beam) is ideal for processing large-scale batch and streaming data. It can read from BigQuery, perform windowing computations, and write to Feature Store's online store. Cloud Functions have timeouts, Cloud Composer is not optimal for streaming, and BigQuery scheduled queries are not designed for streaming-feature computation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use Cloud Composer (Airflow) with a daily DAG to run SQL queries on BigQuery and export results
Why it's wrong here
Daily batch DAGs cannot compute 7-day and 30-day rolling windows over a continuous event stream with the freshness online serving demands, and Composer's scheduler overhead adds cost without streaming semantics. It suits scheduled ETL orchestration, not continuous feature computation; Dataflow with sliding windows would be the correct choice.
- ✗
Schedule a query in BigQuery using scheduled queries and export results to Feature Store
Why it's wrong here
Scheduled queries recompute the entire window on each run, so cost and latency scale with table size rather than with newly arrived events. It is tempting because BigQuery scheduled queries suit periodic batch refreshes of static datasets, where a nightly full recomputation is acceptable and streaming freshness is not required.
- ✗
Use Cloud Functions triggered by Pub/Sub to compute features on the fly
Why it's wrong here
Cloud Functions triggered by Pub/Sub recompute each event independently, so they cannot maintain the cross-event state a 7-day average or 30-day sum requires; each invocation sees only its own payload. This pattern suits stateless per-event transforms such as enrichment or validation, where no windowing across events is needed.
- ✓
Use Dataflow with Apache Beam, reading from BigQuery, computing windowed aggregations, and writing to Vertex AI Feature Store
Why this is correct
Dataflow with Apache Beam provides managed, autoscaling stream processing that reads from BigQuery, applies windowed aggregations for the 7-day and 30-day features, and writes directly to Vertex AI Feature Store. This avoids custom serving infrastructure while meeting the online-serving and scalability constraints.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
This PMLE question is part of Courseiva's 775-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.