Courseiva

PDE Maintaining and Automating Data Workloads Practice Question

A financial services firm runs an Apache Airflow workload on Cloud Composer 2 that ingests market data, runs dbt transformations, and loads curated tables into BigQuery. The DAG currently uses a single PythonOperator that runs a long shell command, and the team wants to make failures easier to diagnose and retries more granular. They also want to avoid rerunning already-successful upstream steps. Which change BEST meets these goals?

⚠ Common exam trap

The trap here is assuming that scaling the Composer environment or enabling catchup improves retry granularity, when those settings affect capacity and scheduling rather than task-level state.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Split the monolithic PythonOperator into a chain of task-level operators such as BigQueryInsertJobOperator and DataprocSubmitJobOperator, with retries configured per task and depends_on_past set appropriately.

Decomposing a monolithic task into discrete operators is the standard Airflow pattern for granular retries and clearer observability. Each task maintains its own state, so a failure in the load step can be retried without rerunning ingestion or transformation. Purpose-built operators emit structured logs and status, and task-level retry configuration gives precise control over failure handling.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Move the shell command into a Cloud Function and invoke it from the PythonOperator using an HTTP call, keeping the DAG structure unchanged.

    Why it's wrong here

    Wrapping the shell command in a Cloud Function changes where the code runs but leaves the DAG as a single task, so Airflow still sees one success or failure point. A failure in the function would mark the whole task failed and require rerunning the entire pipeline. This does not deliver granular retries or preserve upstream successes, and it adds a new runtime dependency to manage.

  • ✓

    Split the monolithic PythonOperator into a chain of task-level operators such as BigQueryInsertJobOperator and DataprocSubmitJobOperator, with retries configured per task and depends_on_past set appropriately.

    Why this is correct

    Breaking the DAG into discrete task-level operators lets Airflow retry only the failed step and preserves successful upstream results, since each task tracks its own state. Using purpose-built operators such as BigQueryInsertJobOperator gives clearer logs and status per step, and configuring retries at the task level provides granular control. This directly addresses diagnosability and avoids rerunning completed work, which is the goal described.

  • ✗

    Increase the worker count and worker machine type on the Cloud Composer environment so the single PythonOperator task has more resources and completes faster.

    Why it's wrong here

    Adding workers or larger machines improves throughput for parallel tasks but does nothing to improve failure diagnosis or retry granularity for a single monolithic task. The entire shell command still fails or succeeds as one unit, so a failure in the final load step would force a rerun of the whole script. This does not meet the requirement to avoid rerunning already-successful upstream steps.

  • ✗

    Set the DAG's max_active_runs to 1 and enable catchup so that failed runs are automatically retried on the next schedule interval.

    Why it's wrong here

    max_active_runs and catchup control how many DAG runs execute concurrently and whether historical intervals are scheduled, not how individual task failures are retried. Catchup would actually create additional backfill runs, increasing load. These settings do not decompose the monolithic task or provide per-step retries, so the diagnosis and rerun-avoidance goals remain unmet.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.