mediumMultiple Choice
PMLE Practice Question: An ML team is scaling a prototype to production
An ML team is scaling a prototype to production. The data pipeline currently reads from Cloud Storage and transforms data with a custom Python script. They need to handle higher throughput and add monitoring. Which approach should they take?
⚠ Common exam trap
Google Cloud often tests the distinction between orchestration (Cloud Composer) and execution (Dataflow), leading candidates to choose an orchestrator when a dedicated processing engine is required for scaling and monitoring.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Migrate the pipeline to Apache Beam on Dataflow with Cloud Monitoring
Apache Beam on Dataflow provides a unified programming model for batch and streaming data processing, enabling automatic scaling to handle higher throughput. Cloud Monitoring integrates natively with Dataflow to track pipeline metrics, latency, and error rates, addressing the monitoring requirement. This approach is purpose-built for production-grade data pipelines, unlike ad-hoc solutions.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Deploy the Python script on a large Compute Engine instance with a cron job
Why it's wrong here
A single large instance with cron scales vertically only and provides no pipeline monitoring beyond OS-level metrics, so throughput ceilings and silent failures persist. It tempts because it is a quick lift-and-shift of the existing script, yet production needs autoscaling and per-stage observability.
- ✓
Migrate the pipeline to Apache Beam on Dataflow with Cloud Monitoring
Why this is correct
Apache Beam on Dataflow provides autoscaling runners that distribute the Python transforms across workers, satisfying the higher-throughput constraint that a single script cannot meet. Dataflow's native integration with Cloud Monitoring supplies the required pipeline metrics and alerting without custom instrumentation.
- ✗
Rewrite the pipeline to use Pub/Sub and Cloud Functions for processing
Why it's wrong here
Pub/Sub with Cloud Functions suits event-driven, short-lived tasks, not bulk transformation of Cloud Storage data; function timeouts and per-invocation limits cap throughput, and pipeline monitoring is absent. It tempts because it is serverless, yet Dataflow offers autoscaling and native job metrics.
- ✗
Use Cloud Composer to orchestrate the Python script at scale
Why it's wrong here
Cloud Composer orchestrates and schedules the existing script but does not itself provide scalable distributed transformation or built-in pipeline monitoring. It tempts because orchestration is a production concern, yet the throughput requirement calls for a managed service such as Dataflow that autoscales and exposes job metrics.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PMLE question from scratch — 775 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PMLE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PMLE exam.