PDE Maintaining and Automating Data Workloads Practice Question
A data engineer manages a Cloud Composer 2 environment. A DAG that downloads a large reference dataset each night occasionally exceeds the default task timeout because the source API is slow. The engineer wants the task to fail fast and be retried automatically rather than hanging for hours, and wants failed runs to be visible for alerting. Which configuration should be applied to the task?
⚠ Common exam trap
Many exam-takers confuse concurrency controls such as max_active_runs or worker autoscaling with a per-task execution timeout.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set execution_timeout on the task to a bounded value and set retries with a retry_delay so failed attempts are re-queued and surfaced in task state.
execution_timeout is the task-level setting that bounds wall-clock runtime and triggers a failure, while retries and retry_delay make the scheduler automatically re-queue the task. Together they produce fast failure, automatic retry, and a recorded state that alerting can observe, which is precisely what the scenario requires for a slow external API.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Configure the worker's celery_worker_autoscale setting to add more workers so the slow API call finishes sooner.
Why it's wrong here
Adding workers increases parallel capacity but does not make a single slow external API call complete faster, nor does it impose a timeout. Autoscaling addresses queue throughput for many tasks, not the wall-clock duration of one task waiting on a remote service. This leaves the hanging task unchanged and adds cost without solving the stated issue.
- ✗
Set depends_on_past to true and add a retry_exponential_backoff so the task waits for the prior run before starting.
Why it's wrong here
depends_on_past serializes runs relative to the previous run's success, which would delay rather than bound execution. retry_exponential_backoff spaces out retries but only applies after a failure has already been recorded, and it does not impose a maximum task duration. Combined, these settings do not make a hung task fail fast, so they miss the core requirement.
- ✗
Increase the DAG's schedule interval and add a max_active_runs limit of one so overlapping runs cannot occur.
Why it's wrong here
max_active_runs controls concurrency between DAG runs, not the duration of a single task. Adjusting the schedule interval changes how often the DAG is triggered but does nothing to bound a slow API call. Neither setting causes a hung task to fail fast or be retried, so this does not address the timeout problem described.
- ✓
Set execution_timeout on the task to a bounded value and set retries with a retry_delay so failed attempts are re-queued and surfaced in task state.
Why this is correct
execution_timeout caps how long a task may run before Airflow marks it failed, and the retries plus retry_delay settings cause the scheduler to re-queue it automatically. The failure state is recorded and available for alerting, which matches the requirement to fail fast, retry, and remain visible.
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.