PDE Maintaining and Automating Data Workloads Practice Question
A data platform team uses Cloud Composer 2 to orchestrate a DAG that runs a Dataproc Serverless batch job producing a partitioned BigQuery table. The DAG passes the output location to a downstream task that runs a dbt model. The team wants failures in the dbt task to automatically trigger a retry of only that task, and they want the DAG to expose the Dataproc job ID in the Airflow UI for troubleshooting. Which approach BEST satisfies both requirements?
⚠ Common exam trap
The trap here is setting retries on the upstream Dataproc operator or in default_args, when the requirement is to retry only the failing dbt task.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the DataprocSubmitJobOperator and pass the returned job reference to XCom, then set retries on the dbt task and read the XCom value in its template fields.
DataprocSubmitJobOperator returns a job reference that Airflow pushes to XCom, which makes the job ID available for downstream tasks and visible in the Airflow UI. Configuring retries specifically on the dbt task ensures that a transient failure there triggers only that task to rerun, leaving the already-successful Dataproc submission intact. This combination precisely meets both the observability and granular retry goals.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use the DataprocSubmitJobOperator and configure the DAG's default_args with retries, which applies the retry policy to all tasks including the dbt task.
Why it's wrong here
default_args.retries applies to every task in the DAG, so a failure in the Dataproc submission would also trigger retries of upstream work, contradicting the goal of retrying only the dbt task. It also does not specifically expose the job ID in the UI beyond what the operator already provides. This approach is broader than required and can cause unnecessary reprocessing.
- ✗
Use a single PythonOperator that submits the Dataproc job via the API, waits for completion, and then invokes dbt, with retries set on the PythonOperator.
Why it's wrong here
Combining submission, waiting, and dbt execution into one task means a dbt failure marks the entire task failed, forcing the Dataproc job to be resubmitted on retry. The job ID would only be logged as text, not exposed through XCom or the operator's native interface. This fails the requirement for granular retries and does not provide structured job ID visibility.
- ✓
Use the DataprocSubmitJobOperator and pass the returned job reference to XCom, then set retries on the dbt task and read the XCom value in its template fields.
Why this is correct
DataprocSubmitJobOperator returns the job reference, which Airflow automatically pushes to XCom, making the job ID visible and available downstream. Setting retries on the dbt task ensures only that task is retried on failure, and using XCom in template fields lets the dbt task consume the job ID or output path. This satisfies both the observability and granular retry requirements.
- ✗
Use the DataprocSubmitJobOperator with retries set on it, and add a separate downstream PythonOperator that runs dbt without its own retry configuration.
Why it's wrong here
Placing retries on the Dataproc operator does not help when the dbt task fails, because Airflow retries the failed task, not upstream tasks. The dbt task without retry configuration would fail permanently on a transient error. While the Dataproc operator does expose the job ID, the retry behavior does not meet the requirement for retrying only the dbt task.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.