PDE Maintaining and Automating Data Workloads Practice Question
A data engineer needs to run a recurring SQL transformation in BigQuery every night at 02:00 and, if it fails, retry automatically and send a notification. The team wants the least operational overhead and no external orchestrator. What should they use?
⚠ Common exam trap
The trap here is assuming a general-purpose orchestrator is always the right answer for scheduled work, when BigQuery's built-in scheduled query already covers a single recurring SQL statement.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A BigQuery scheduled query configured with a schedule, a destination table, and notification settings.
BigQuery scheduled queries are the native mechanism for running SQL on a recurring schedule. They accept a schedule, a query, an optional destination table, and configuration for failure notification through Pub/Sub, all without any infrastructure to operate. Composer, Cloud Scheduler with API calls, and Dataflow all can run SQL, but each introduces components to manage, which conflicts with the requirement for the least operational overhead.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A Cloud Composer DAG that runs a BigQueryInsertJobOperator on a cron schedule.
Why it's wrong here
Composer can schedule the query, but it requires provisioning and maintaining an Airflow environment, which is far more operational overhead than the scenario calls for. For a single recurring SQL statement with retry and notification, a full orchestration service is unnecessary. Composer is better suited when there are multiple interdependent tasks to coordinate.
- ✗
A Cloud Scheduler job that calls the BigQuery jobs.insert API with a service account.
Why it's wrong here
Cloud Scheduler can invoke the API on a cron, but you must construct and manage the HTTP request, handle authentication, and build retry and notification logic yourself. That is more custom code and more failure modes than the built-in scheduled query feature, which already packages scheduling, retries, and notifications together.
- ✓
A BigQuery scheduled query configured with a schedule, a destination table, and notification settings.
Why this is correct
BigQuery scheduled queries natively support a recurring schedule, a SQL statement, a destination table, and optional Pub/Sub notification on failure. They run entirely inside BigQuery with no cluster or orchestrator to manage, which matches the requirement for minimal operational overhead. This is the purpose-built feature for recurring SQL transformations on a schedule.
- ✗
A Dataflow batch pipeline that reads the source table and writes the transformed result.
Why it's wrong here
Dataflow is designed for large-scale data processing with custom transforms, not for running a single SQL statement on a nightly schedule. Using it here adds a pipeline to build, deploy, and monitor when the transformation is expressible as SQL. It solves a scale problem the scenario does not have.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.