PDE Maintaining and Automating Data Workloads Practice Question
A company runs a critical batch pipeline using Cloud Dataflow. The pipeline processes financial transactions and runs every hour. Recently, some runs have failed due to transient errors (e.g., network timeouts). The engineer wants to automatically retry failed runs without manual intervention. The pipeline is launched from a Cloud Composer DAG using DataflowPythonOperator. What is the BEST way to handle retries?
⚠ Common exam trap
Many candidates confuse Dataflow-level retry options (like --maxRetryAttempts) with Airflow task-level retries, or assume that a sensor or external trigger is required to detect and retry failures, when in fact Airflow's native retry parameter is the simplest and most appropriate solution for transient errors in a DAG-managed pipeline.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set the 'retries' parameter in the DAG's default_args to a positive integer.
Cloud Composer (Apache Airflow) natively supports task-level retries via the 'retries' parameter in default_args. When a DataflowPythonOperator fails due to a transient error, Airflow automatically re-executes the task up to the specified number of retries, without requiring custom sensors or external triggers. This is the simplest and most reliable mechanism for handling transient failures in a DAG-driven pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Add a DataflowJobStatusSensor in the DAG that waits for job completion and retries if failed.
Why it's wrong here
A sensor only polls job state; it does not re-run the failed Dataflow job, so transient failures persist. Setting retries on the operator (or DAG default_args) re-executes the task automatically. Sensors suit gating downstream tasks on completion, not orchestrating retry of the launching task itself.
- ✓
Set the 'retries' parameter in the DAG's default_args to a positive integer.
Why this is correct
Setting retries in the DAG's default_args makes Cloud Composer automatically re-run the DataflowPythonOperator task after transient failures such as network timeouts, without manual intervention. Airflow's task-level retry mechanism satisfies the requirement for unattended recovery of hourly runs.
- ✗
Configure the Dataflow pipeline to automatically retry on failure using the --numberOfWorkerHarnessThreads option.
Why it's wrong here
--numberOfWorkerHarnessThreads controls concurrent threads per worker for throughput, not job-level retry behaviour. It cannot restart a failed pipeline. Composer operator retry settings re-execute the task on transient errors. This flag suits tuning parallelism in high-volume pipelines, not failure recovery.
- ✗
Use a Cloud Function triggered by Cloud Scheduler to re-launch the pipeline if the Dataflow job fails.
Why it's wrong here
A Cloud Function polling job status duplicates orchestration Composer already provides, adding latency and external state. Setting retries on the DataflowPythonOperator re-runs the task natively within the DAG. Scheduler-triggered functions suit workloads outside Composer, not pipelines already orchestrated by it.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.