PDE Maintaining and Automating Data Workloads Practice Question
Your team runs a Cloud Composer 2 environment in project `analytics-prod`. A nightly DAG loads Cloud Storage files into BigQuery, but a recent Cloud Storage outage caused several tasks to fail after exhausting their retries. You want failed task instances to automatically re-run without manual intervention once the upstream dependency recovers. What should you do?
⚠ Common exam trap
The trap here is assuming that increasing retry counts or delays will recover tasks that have already reached the failed state, when only clearing and re-queuing those task instances makes the scheduler run them again.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Define a task-level `on_failure_callback` that calls the Airflow REST API to clear the failed task instances after a delay.
Failed task instances that have exhausted retries stay in a failed state until they are cleared, at which point the scheduler re-queues them. Triggering that clear programmatically, for example from a failure callback that waits for the dependency to recover, restores the pipeline automatically. Backfill settings, retry counts, and timeout tuning only affect future scheduling or in-flight attempts, not instances already marked failed.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set the Airflow configuration option `[core] default_task_retries` to a higher value in the environment configuration and restart the schedulers.
Why it's wrong here
Raising default retries increases how many times a task is attempted before being marked failed, but the tasks in this scenario already exhausted their retries. It does not resurrect task instances that have already reached the failed state, so the pipeline still requires a manual re-run and the outage recovery goal is not met.
- ✗
Increase the DAG's `retry_delay` and `execution_timeout` so tasks wait longer for Cloud Storage to recover before failing.
Why it's wrong here
Retry delay and execution timeout govern timing within the existing retry budget and per-attempt duration. Once retries are exhausted and a task is marked failed, adjusting these values does not cause the failed instance to run again. The scenario needs re-execution of already-failed tasks, not longer waits during attempts that have already completed.
- ✗
Configure the DAG with `catchup=True` and set `max_active_runs=1` so missed schedule intervals are backfilled.
Why it's wrong here
Catchup controls whether past schedule intervals are created as DAG runs, not whether individual failed task instances are retried. Enabling it here would generate backfill runs for historic intervals and could duplicate data loads rather than re-executing only the tasks that failed during the outage, so it does not address automatic recovery of the failed tasks.
- ✓
Define a task-level `on_failure_callback` that calls the Airflow REST API to clear the failed task instances after a delay.
Why this is correct
Clearing failed task instances through the Airflow REST API re-queues them so the scheduler runs them again once the upstream Cloud Storage dependency is healthy. Wrapping this in an `on_failure_callback` with a delay automates recovery without manual operator intervention, which is exactly the requirement after retries are exhausted.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.