Courseiva
mediumMultiple Choice

PDE Practice Question: A financial services company uses Cloud Composer…

A financial services company uses Cloud Composer to orchestrate daily batch jobs. One job extracts data from MongoDB to Cloud Storage, then loads into BigQuery, and finally runs a Dataflow pipeline for aggregations. The Dataflow job fails intermittently. They want to automatically restart only the failed Dataflow job without re-running the earlier extraction and load. Which Airflow operator configuration should they use?

⚠ Common exam trap

Google Cloud often tests the distinction between task-level retry mechanisms and dependency/trigger rules, so the trap here is confusing `retries` (which restarts the failed task) with `trigger_rule` or `depends_on_past` (which only affect task scheduling or downstream execution).

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set retries=2 on the Dataflow operator

Setting retries=2 on the Dataflow operator instructs Airflow to automatically restart only that specific task upon failure, without affecting upstream tasks (MongoDB extraction, BigQuery load). This isolates the retry to the Dataflow job, preserving the earlier completed work and avoiding redundant data movement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Implement a SlaMiss sensor

    Why it's wrong here

    An SlaMiss sensor only records that a task exceeded its expected duration; it neither detects failure nor triggers retries. It is tempting because SLA monitoring sounds like automated recovery, but it is for alerting on lateness, not restarting a failed Dataflow job.

  • ✗

    Use a DAG with depends_on_past=True

    Why it's wrong here

    depends_on_past=True sequences task instances across DAG runs; it does not retry a failed task within the current run. It is tempting because it enforces ordering, but it addresses run-to-run dependencies, not automatic restart of the failed Dataflow job.

  • ✓

    Set retries=2 on the Dataflow operator

    Why this is correct

    Airflow retries apply at task level, so setting retries=2 on the Dataflow operator restarts only that failed task while upstream extract and load tasks remain successful and are not re-run. This satisfies the requirement to avoid repeating earlier extraction and loading.

  • ✗

    Set trigger_rule='one_success' for downstream tasks

    Why it's wrong here

    trigger_rule='one_success' governs when a downstream task may run, not whether the failed Dataflow task itself retries. It is tempting because it controls task execution flow, but it cannot restart the failed task without upstream re-runs.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.