mediumMultiple Choice
PDE Practice Question: A financial services company uses Cloud Composer…
A financial services company uses Cloud Composer to orchestrate daily batch jobs. One job extracts data from MongoDB to Cloud Storage, then loads into BigQuery, and finally runs a Dataflow pipeline for aggregations. The Dataflow job fails intermittently. They want to automatically restart only the failed Dataflow job without re-running the earlier extraction and load. Which Airflow operator configuration should they use?
⚠ Common exam trap
Google Cloud often tests the distinction between task-level retry mechanisms and dependency/trigger rules, so the trap here is confusing `retries` (which restarts the failed task) with `trigger_rule` or `depends_on_past` (which only affect task scheduling or downstream execution).
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set retries=2 on the Dataflow operator
Setting retries=2 on the Dataflow operator instructs Airflow to automatically restart only that specific task upon failure, without affecting upstream tasks (MongoDB extraction, BigQuery load). This isolates the retry to the Dataflow job, preserving the earlier completed work and avoiding redundant data movement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Implement a SlaMiss sensor
Why it's wrong here
An SlaMiss sensor only records that a task exceeded its expected duration; it neither detects failure nor triggers retries. It is tempting because SLA monitoring sounds like automated recovery, but it is for alerting on lateness, not restarting a failed Dataflow job.
- ✗
Use a DAG with depends_on_past=True
Why it's wrong here
depends_on_past=True sequences task instances across DAG runs; it does not retry a failed task within the current run. It is tempting because it enforces ordering, but it addresses run-to-run dependencies, not automatic restart of the failed Dataflow job.
- ✓
Set retries=2 on the Dataflow operator
Why this is correct
Airflow retries apply at task level, so setting retries=2 on the Dataflow operator restarts only that failed task while upstream extract and load tasks remain successful and are not re-run. This satisfies the requirement to avoid repeating earlier extraction and loading.
- ✗
Set trigger_rule='one_success' for downstream tasks
Why it's wrong here
trigger_rule='one_success' governs when a downstream task may run, not whether the failed Dataflow task itself retries. It is tempting because it controls task execution flow, but it cannot restart the failed task without upstream re-runs.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.