easyMultiple Choice
PDE Practice Question: Wants to automate their batch data processing…
An organization wants to automate their batch data processing pipeline using Cloud Composer. The pipeline consists of multiple tasks: extract from Cloud Storage, transform with Dataflow, and load into BigQuery. Which Airflow operator should be used to run Dataflow jobs?
⚠ Common exam trap
Google Cloud often tests the distinction between Dataflow and Dataproc operators, so the trap here is that candidates might confuse DataprocSubmitJobOperator (for Hadoop/Spark) with Dataflow operators, especially when the question mentions 'transform' without specifying the processing framework.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
DataflowCreatePythonJobOperator
B is correct because the DataflowCreatePythonJobOperator is specifically designed to submit and manage Apache Beam pipelines written in Python as Dataflow jobs in Google Cloud. This operator handles the creation of a Dataflow job from a Python file, which aligns with the requirement to run Dataflow transformations within a Cloud Composer DAG.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
BigQueryInsertJobOperator
Why it's wrong here
BigQueryInsertJobOperator runs BigQuery SQL or load jobs; it cannot start a Dataflow pipeline, so the transform task would never execute. It is tempting because BigQuery is the pipeline's load destination, and this operator is correct when a task must run a query or load job inside BigQuery itself.
- ✓
DataflowCreatePythonJobOperator
Why this is correct
DataflowCreatePythonJobOperator launches a Python-defined Dataflow pipeline directly from Airflow, satisfying the requirement to run Dataflow jobs within the Composer DAG. It handles pipeline submission and job monitoring natively, unlike generic operators, making it the precise fit for the transform stage between Cloud Storage extraction and BigQuery loading.
- ✗
GCSToBigQueryOperator
Why it's wrong here
GCSToBigQueryOperator loads objects from Cloud Storage straight into BigQuery, bypassing Dataflow entirely, so the transform stage is skipped. It is tempting because the pipeline's extract and load endpoints match its inputs and outputs, and it is correct when no transformation step is required between them.
- ✗
DataprocSubmitJobOperator
Why it's wrong here
DataprocSubmitJobOperator submits Dataproc jobs, not Dataflow jobs; the stem's transform step runs on Dataflow, so this operator cannot launch it. It is tempting because Dataproc and Dataflow are both managed Google Cloud data-processing services, and DataprocSubmitJobOperator is the correct choice when the transform runs on Dataproc clusters.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.