PDE Maintaining and Automating Data Workloads Practice Question
You need to schedule a Dataproc Spark job to run at 2 AM every day, and upon completion, trigger a BigQuery load job. Which Cloud Composer operator should you use to run the Spark job?
⚠ Common exam trap
Watch out — candidates often confuse operators that manage cluster lifecycle (like DataprocClusterCreateOperator) with operators that submit jobs, or they mistakenly think DataflowPythonOperator can run Spark jobs because both are data processing frameworks.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
DataprocSubmitJobOperator
The DataprocSubmitJobOperator is specifically designed to submit a job (e.g., a Spark job) to an existing Dataproc cluster. In this scenario, you need to run a Spark job on a scheduled basis, and Cloud Composer (Airflow) provides this operator to submit the job to Dataproc. After the Spark job completes, you can chain a BigQuery load operator to trigger the load, matching the requirement exactly.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
DataflowPythonOperator
Why it's wrong here
DataflowPythonOperator launches Apache Beam pipelines on Dataflow, not Spark jobs on Dataproc. It is tempting because both are managed data-processing services, and it would be correct when the workload is a Beam pipeline rather than a Spark application.
- ✗
BigQueryOperator
Why it's wrong here
BigQueryOperator runs BigQuery jobs such as load or query tasks; it cannot submit a Spark job to a Dataproc cluster. It is tempting because the stem's downstream step is a BigQuery load, and it would be correct for that second task, not the Spark submission.
- ✗
DataprocClusterCreateOperator
Why it's wrong here
DataprocClusterCreateOperator only provisions a cluster; it neither submits the Spark job nor triggers the downstream BigQuery load, so the 2 AM schedule would create idle clusters. It is tempting because cluster creation is a genuine prerequisite step, and it would be correct within a workflow that first builds a cluster before a DataprocSubmitJobOperator runs the job.
- ✓
DataprocSubmitJobOperator
Why this is correct
DataprocSubmitJobOperator submits a Spark job to a Dataproc cluster and waits for completion, so it fits the scheduled 2 AM run. Its downstream task can then trigger the BigQuery load job within the same DAG.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.