PDE · domain
Maintaining and Automating Data Workloads
This domain covers keeping production data pipelines healthy on Google Cloud: monitoring and alerting with Cloud Monitoring, tuning BigQuery and Dataproc performance, protecting data with DLP and IAM, and automating recurring work with Cloud Composer, Cloud Scheduler, and Dataflow. Questions present operational symptoms and ask you to pick the right service, metric, or configuration.
Focused practice
Practice Maintaining and Automating Data Workloads questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Maintaining and Automating Data Workloads
Be able to read operational symptoms and choose the correct Google Cloud control: DLP for de-identification, partition and cluster filters for BigQuery pruning, oldest_unacked_message_age for Pub/Sub lag, and preemptible-safe Dataproc sizing. The most important skill is matching each symptom to the exact metric or configuration that fixes it.
Selecting Cloud DLP infoType detectors and de-identification transforms (masking, tokenization) for sensitive data in Cloud Storage and BigQuery
Diagnosing BigQuery scans with INFORMATION_SCHEMA.JOBS, partitioning, clustering, and materialized views
Creating Cloud Monitoring alerting policies on Pub/Sub metrics such as oldest_unacked_message_age with appropriate filters
Tuning Dataproc clusters, including preemptible worker ratios, autoscaling policies, and Spark shuffle configuration
Watch out for
Common Maintaining and Automating Data Workloads exam traps
- ▸Assuming a partitioned table is always pruned; filters must reference the partitioning column directly, or BigQuery scans the whole table.
- ▸Alerting on the wrong Pub/Sub metric, such as message count instead of oldest_unacked_message_age, so slow-consumer backlogs go unnoticed.
- ▸Using too many preemptible Dataproc workers without fallback, causing shuffle failures and slow jobs when VMs are reclaimed.
Question index
All Maintaining and Automating Data Workloads questions (83)
Click any question to see the full explanation, or start a practice session above.
You are setting up Dataplex data quality rules for a BigQuery table. You want to define rules that check for non-null values in key columns and also validate that a column's values fall within a certain range. Which TWO rule types must you use? (Choose 2)
Hard2Your team runs a Cloud Composer (Airflow) environment that executes a nightly BigQuery ELT DAG. A downstream task must only run after an upstream task that loads a partitioned table completes, and you want the downstream task to wait for the load's completion signal without polling BigQuery repeatedly. Which Airflow mechanism should you implement to coordinate these tasks within the DAG?
Medium3You need to estimate the cost of a BigQuery query before running it. Which command or feature should you use?
Easy4Your team runs a Cloud Composer 2 environment in project `analytics-prod`. A nightly DAG loads Cloud Storage files into BigQuery, but a recent Cloud Storage outage caused several tasks to fail after exhausting their retries. You want failed task instances to automatically re-run without manual intervention once the upstream dependency recovers. What should you do?
Medium5Your Dataflow streaming pipeline writes to BigQuery and occasionally fails with `QuotaExceededError` on streaming inserts during peak hours. You want to reduce insert-driven quota pressure without changing the pipeline's output schema or losing exactly-once semantics. Which change should you make?
Medium6You are designing a data quality pipeline that must inspect PII in BigQuery tables and de-identify sensitive columns before sharing with analysts. Which GCP service should you use?
Medium7You need to schedule a simple workflow that fetches data from an API every hour, transforms it using Cloud Functions, and writes the result to Cloud Storage. The workflow has no complex branching or retry logic beyond basic retries. Which orchestration service is the MOST cost-effective and simplest to implement?
Easy8A company runs a Dataflow pipeline that processes a high-volume data stream. They notice that the pipeline's worker CPU utilisation is near 100% and the system lag is increasing. Which three actions can improve performance? (Choose three.)
Hard9Your team runs a Cloud Composer 2 environment to orchestrate BigQuery ELT jobs. A DAG that loads a critical fact table must only run after an upstream ingestion DAG completes, and you want Composer itself to trigger it automatically without any external scheduler. Which mechanism should you use?
Medium10You are using Cloud Composer to orchestrate a data pipeline that runs a Dataproc job to process data, followed by a BigQuery load. You notice that the Dataproc job sometimes takes longer than expected, causing the BigQuery load to start before the Dataproc job finishes, resulting in incomplete data. Which Airflow feature should you use to ensure the BigQuery load only runs after the Dataproc job completes successfully?
Medium11A retail analytics team runs a Cloud Composer DAG that starts a Dataproc job, waits for it, and then runs a BigQuery load. The Dataproc job sometimes fails due to a transient YARN resource shortage, and the team wants the DAG to automatically retry the Dataproc submission a few times before alerting. Which Airflow configuration on the Dataproc task best meets this need?
Medium12You need to schedule a Dataproc Spark job to run at 2 AM every day, and upon completion, trigger a BigQuery load job. Which Cloud Composer operator should you use to run the Spark job?
Easy13A data engineering team manages BigQuery datasets across multiple projects. They want to automatically detect and respond when a scheduled query fails, and they want the response to create an incident in their existing ticketing system. The team prefers minimal custom infrastructure and wants to use native Google Cloud tooling. Which approach should they use?
Medium14You manage a BigQuery reservation with 500 baseline slots and autoscaling up to 2000 slots. Your team runs a mix of interactive queries and batch load jobs. During peak hours, you notice that interactive queries are throttled when autoscaling slots are consumed by long-running batch loads. How can you ensure interactive queries get priority access to slots?
Hard15Your streaming Dataflow pipeline reads from Pub/Sub, enriches data with a side input, and writes to BigQuery. You need to update the enrichment logic without draining the pipeline, to minimize data loss and maintain exactly-once semantics. What should you do?
Medium16A data engineer needs to run a recurring SQL transformation in BigQuery every night at 02:00 and, if it fails, retry automatically and send a notification. The team wants the least operational overhead and no external orchestrator. What should they use?
Easy17A company runs a Dataflow streaming pipeline that processes financial transactions. They need to apply a new transformation that enriches the data with a lookup from Cloud Bigtable without stopping the pipeline. The pipeline must be updated in a way that minimises data loss and preserves exactly-once semantics. What is the recommended approach?
Hard18Your organization has a BigQuery flat-rate reservation with 500 slots. During peak hours, queries are queued and you need additional capacity temporarily. You want to add slots for a burst of activity without committing to a long-term purchase. What should you do?
Medium19A team wants to enforce data quality rules on BigQuery tables using Dataplex. They need to run column-level checks for null values and row-level checks for value ranges on a schedule. Which Dataplex feature should they use?
Medium20A data engineer needs to monitor a Pub/Sub-based streaming pipeline. Which two Cloud Monitoring metrics should be used to detect a backlog of unprocessed messages? (Choose two.)
Medium21You manage several Cloud Composer 2 environments that run production DAGs. You must define an alerting strategy that detects when a DAG run fails and when a task is stuck retrying for an unusually long time, using Cloud Monitoring. (Choose two.)
Hard22Your team runs a Cloud Composer 2 environment (composer-2.1.0-airflow-2.6.3) that executes a daily BigQuery ETL workflow. The workflow must not run on weekends. You want to implement this with minimal code and without modifying the DAG's task logic. What should you do?
Medium23You manage a Cloud Dataflow streaming pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline uses the BigQueryIO write transform with STREAMING_INSERTS. You need to ensure exactly-once processing semantics for the BigQuery writes. What should you do?
Hard24A data engineer schedules a Cloud Composer 2 environment to run a DAG that triggers a Dataflow batch job every night. The DAG sometimes fails because the Dataflow job takes longer than the default task timeout. The engineer wants the DAG to wait for the Dataflow job to finish rather than timing out. Which change should the engineer make?
Easy25A financial services firm stores customer transaction data in BigQuery. Compliance requires that a nightly Cloud Composer DAG verify that the previous day's partition is complete before downstream reporting DAGs run, and that the reporting DAG never start if the verification fails. The engineer wants the dependency expressed inside orchestration rather than by polling from the reporting DAG. What should the engineer do?
Medium26A financial services company runs a Dataflow streaming pipeline that reads from Pub/Sub and writes enriched records to BigQuery. Compliance requires that the raw Pub/Sub messages be retained for seven years so that any record can be reprocessed if the enrichment logic is later found to be incorrect. The pipeline currently has no archival step. What should the data engineer do to satisfy the retention requirement with the least operational overhead?
Hard27You are using Cloud Workflows to orchestrate a series of API calls. You need to handle errors and retries. Which THREE features of Cloud Workflows can you use? (Choose THREE.)
Easy28Your organization runs a Cloud Composer 2 environment that executes dozens of DAGs. Several DAGs share a connection to an external REST API that enforces a rate limit of 100 requests per minute. During peak hours, DAGs fail with HTTP 429 errors. You want to prevent these failures without changing the external API's limits. What should you do?
Hard29You have a Dataflow batch pipeline that processes data from Cloud Storage and writes to BigQuery. The pipeline uses a custom DoFn that sometimes throws exceptions due to malformed input records. You want to ensure that the pipeline continues processing valid records while logging the malformed ones for later analysis, without failing the entire job. Which Dataflow feature should you use?
Hard30A data engineer must ensure that a Cloud Composer DAG which loads a BigQuery table runs every day at 02:00 UTC and that a dependent downstream report DAG runs only after the load succeeds. The report DAG lives in the same Composer environment but is a separate DAG file. Which feature should be used to coordinate the two DAGs?
Easy31An engineer needs to create a reusable Dataflow pipeline that can be executed with different parameters without modifying code. Which Dataflow feature should they use?
Easy32Your company uses Cloud Composer to run a daily ETL workflow. The workflow consists of several tasks that must run in a specific order. You want to receive an alert if any task fails. Which Cloud Monitoring feature should you use?
Easy33A streaming Dataflow pipeline needs to be updated without draining the existing pipeline. Which update strategy should be used?
Easy34A data engineer is migrating a Composer 1 environment to Cloud Composer 2 and notices that a DAG relying on a legacy operator for a deprecated service no longer works. The engineer wants a durable fix that keeps the pipeline running and avoids repeating this problem in future upgrades. Which action should be taken?
Hard35Which BigQuery feature allows you to estimate the cost of a query before running it, by returning the number of bytes that would be processed?
Easy36A data engineer has a BigQuery SQL script that must run every day at 06:00, load its results into a reporting table, and retry automatically if the query fails due to transient errors. The team has no existing orchestration tooling, wants the lowest operational overhead, and needs the schedule and the SQL to be managed entirely inside Google Cloud. Which approach should the engineer use?
Easy37A data engineer is building a batch pipeline that runs daily using Cloud Composer. The pipeline has three tasks: extract data from Cloud Storage, transform data using Dataflow, and load the transformed data into BigQuery. The engineer wants to ensure that the Dataflow job only starts after the extraction task completes successfully, and the load task only starts after the Dataflow job finishes. How should the engineer define the task dependencies in the Airflow DAG?
Medium38A data engineer wants to quickly estimate the cost of running a BigQuery query before executing it. Which command-line tool or command should they use?
Easy39You need to schedule a BigQuery query to run every day at 6:00 AM and write the results to a new table. The query is simple and does not require complex dependencies. You want a low-maintenance, serverless solution with minimal configuration. What should you do?
Easy40You want to monitor the latency of messages in a Pub/Sub subscription. Which Cloud Monitoring metric should you use to see the age of the oldest unacknowledged message?
Easy41You are building a data pipeline that runs daily batch jobs on Dataproc, then loads results into BigQuery. You want to orchestrate the entire workflow, including dependencies between steps, retries, and monitoring. Which Google Cloud service is most appropriate?
Medium42A Dataflow streaming pipeline writes to BigQuery and has run in production for months. The team wants to add a transformation and deploy the change with zero data loss and no interruption to the running pipeline. Which deployment approach should they use?
Hard43Which Dataflow feature allows you to package a pipeline into a reusable template that can be deployed with different parameters at runtime?
Easy44You need to inspect a BigQuery table for sensitive data such as credit card numbers and apply masking. Which GCP service should you use to identify and de-identify the data?
Medium45Your team uses Cloud Composer to orchestrate a daily ETL workflow that extracts data from an on-premises database, transforms it in Dataproc, and loads it into BigQuery. The workflow occasionally fails due to transient network issues. You want to automatically retry failed tasks without manual intervention. What should you do?
Medium46A team runs a Dataflow batch pipeline that writes to BigQuery. They want to reduce cost and improve reliability of the write step. The pipeline currently uses BigQueryIO with the default write method. Which two changes should they make to achieve these goals? (Choose two.)
Medium47A company is migrating their on-premises data warehouse to BigQuery. They have a mix of batch and streaming ingestion. The data team wants to optimize query costs. Which THREE practices should they adopt?
Medium48You want to optimize BigQuery costs for a large dataset that is frequently queried by time range. You also need to ensure that predictable workloads have dedicated slot capacity. Which TWO strategies should you combine? (Choose 2)
Medium49Your company uses Cloud Composer to orchestrate a data pipeline that includes Dataproc Spark jobs and BigQuery load operations. You need to pass the output file path from the Spark job to the next BigQuery task in the DAG. Which two mechanisms can you use to share data between tasks? (Choose TWO.)
Medium50Your organization runs a Cloud Composer (Apache Airflow) environment that executes a DAG every night to move data from Cloud Storage into BigQuery. The DAG has been succeeding for months, but last week the nightly load silently produced a BigQuery table with zero rows while the Airflow task still reported success. You need to add a safeguard that fails the DAG task whenever the loaded row count is zero before downstream tasks run. What should you do?
Medium51A data engineer needs to alert when Pub/Sub subscription has messages older than 1 hour. Which Cloud Monitoring metric and filter should they use?
Hard52You are monitoring a Dataproc cluster and notice that the cluster utilisation is high, but jobs are running slowly. The cluster uses preemptible workers for cost savings. What is the most likely cause of the performance degradation?
Medium53A data engineering team maintains several Cloud Composer DAGs that load files from Cloud Storage into BigQuery. They want the DAG to start only after a new file has landed in the source bucket, rather than polling on a fixed schedule and often finding nothing. Which Airflow construct should they use to trigger the DAG based on the arrival of the object?
Easy54A data engineer needs to run a recurring nightly extract-transform-load job that pulls data from a REST API, applies Python transformations, and writes the output to a Cloud Storage bucket. The team wants a fully managed, serverless scheduler that can retry failed runs and send notifications, and they do not want to maintain any cluster or VM. Which Google Cloud service should they use to define and run this job?
Easy55You manage a Dataflow streaming pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline must be updated to add a new transformation that enriches each message with data from Cloud SQL. You need to minimize downtime and ensure the pipeline continues processing without data loss. What should you do?
Medium56A company runs a critical batch pipeline using Cloud Dataflow. The pipeline processes financial transactions and runs every hour. Recently, some runs have failed due to transient errors (e.g., network timeouts). The engineer wants to automatically retry failed runs without manual intervention. The pipeline is launched from a Cloud Composer DAG using DataflowPythonOperator. What is the BEST way to handle retries?
Medium57An organization uses BigQuery on-demand pricing. To control costs, they want to estimate the bytes processed by a query before running it. Which command or method should they use?
Easy58You have a BigQuery table that is used by multiple teams. To save costs, you want to provide a consistent view of the data as of a specific point in time without creating full copies. Which BigQuery feature should you use?
Medium59You need to deploy a reusable Dataflow pipeline that can be executed with different parameters from Cloud Composer. Which TWO components should you use? (Choose 2)
Easy60You manage a Cloud Composer environment that runs a critical DAG every hour. The DAG includes a task that calls a Cloud Function to process data. Recently, the Cloud Function started taking longer than expected, causing the DAG to exceed its SLA. You need to detect this delay and automatically retry the task if it fails due to timeout, while minimizing changes to the DAG. What should you do?
Hard61A data engineer needs to inspect a BigQuery table for sensitive data such as credit card numbers and email addresses before sharing it with a third party. The engineer also wants to de-identify the data by masking the sensitive columns. Which Google Cloud service should be used?
Easy62A company uses Cloud Pub/Sub for a real-time data pipeline. The subscription has a backlog of millions of messages that are not being processed quickly enough. In Cloud Monitoring, you observe that the 'subscription/num_undelivered_messages' metric is high and growing, while 'subscription/oldest_unacked_message_age' is also increasing. Which action is MOST likely to reduce the backlog?
Medium63A data platform team wants to grant a service account the ability to run BigQuery jobs and read data in a specific dataset, while ensuring it cannot create or delete datasets. Which IAM approach satisfies this with least privilege?
Medium64You are optimizing a BigQuery query that scans 1 TB of data every day. The query joins a large fact table (partitioned by date) with a small dimension table. You notice that the query always scans the entire fact table, even though you only need the last 7 days of data. Which optimization will MOST reduce the bytes scanned?
Hard65You are running a streaming pipeline with Dataflow that reads from Pub/Sub and writes to BigQuery. You notice that the system lag metric is increasing over time, indicating that messages are taking longer to process. What is the most likely cause and how should you address it?
Medium66A data platform team uses Cloud Composer 2 to orchestrate a DAG that runs a Dataproc Serverless batch job producing a partitioned BigQuery table. The DAG passes the output location to a downstream task that runs a dbt model. The team wants failures in the dbt task to automatically trigger a retry of only that task, and they want the DAG to expose the Dataproc job ID in the Airflow UI for troubleshooting. Which approach BEST satisfies both requirements?
Medium67Your team uses Cloud Dataproc for Spark ML training jobs. You want to reduce costs for non-critical, fault-tolerant training jobs. Which Dataproc feature should you use for worker nodes?
Medium68A data engineer needs to orchestrate a complex data pipeline that involves multiple steps including data extraction from Cloud Storage, transformation using Dataflow, and loading into BigQuery. The pipeline has dependencies between tasks and requires monitoring and retries. Which Google Cloud service should be used for orchestration?
Easy69A data engineer manages a Cloud Composer 2 environment. A DAG that downloads a large reference dataset each night occasionally exceeds the default task timeout because the source API is slow. The engineer wants the task to fail fast and be retried automatically rather than hanging for hours, and wants failed runs to be visible for alerting. Which configuration should be applied to the task?
Hard70A company wants to use Cloud DLP to inspect data in BigQuery for sensitive information and de-identify it by masking credit card numbers. They want to perform this on a schedule. Which approach should they take?
Hard71You are building a data pipeline that ingests data from on-premises into Cloud Storage, then processes it with Dataproc, and finally loads into BigQuery. You need to schedule the pipeline to run daily. The pipeline must handle occasional failures gracefully. Which THREE Google Cloud services should you use together to achieve this? (Choose 3)
Medium72A media company runs a Cloud Composer environment whose DAGs trigger Dataflow batch jobs and BigQuery loads. The operations team reports that Composer costs are rising and that DAGs occasionally stall because workers are saturated. You review the environment and find that several tasks are long-running sensors that hold worker slots while waiting on external conditions. Which TWO changes should you make to reduce worker saturation and cost? (Choose two.)
Hard73A financial services firm runs an Apache Airflow workload on Cloud Composer 2 that ingests market data, runs dbt transformations, and loads curated tables into BigQuery. The DAG currently uses a single PythonOperator that runs a long shell command, and the team wants to make failures easier to diagnose and retries more granular. They also want to avoid rerunning already-successful upstream steps. Which change BEST meets these goals?
Hard74You are designing a Cloud Composer workflow that loads data from Cloud Storage into BigQuery, runs a Dataflow job to transform the data, and then triggers a Dataproc Spark job. After each step, you need to conditionally branch based on success or failure. Which Airflow feature allows you to pass messages between tasks to enable dynamic branching?
Medium75You are designing the deployment process for a Dataflow streaming pipeline that processes financial transactions. The pipeline must be updated without losing in-flight state, such as open windows and timers, and without downtime. Your team uses the Apache Beam Java SDK and deploys from a CI/CD pipeline. Which update strategy should you use?
Hard76You need to orchestrate a simple, linear workflow that calls several Cloud Functions and API endpoints sequentially with conditional logic. The workflow should be defined as code and have minimal overhead. Which GCP service should you use?
Easy77A data engineer uses Cloud Composer to orchestrate a daily batch pipeline. A downstream task should only start after an upstream BigQuery load job finishes successfully and a specific file appears in Cloud Storage. Which combination of operators should the engineer use in the Airflow DAG?
Medium78A data engineer must give a Dataproc Serverless for Spark batch workload permission to read objects from a specific Cloud Storage bucket and write to a BigQuery dataset, following least privilege. The workload runs as a custom service account. Which approach should be used?
Easy79Your company uses Cloud Composer to orchestrate a complex data pipeline. You need to ensure that the pipeline can recover from failures and that tasks are retried automatically with exponential backoff. You also want to be alerted if a task fails after all retries. Which combination of features should you implement?
Hard80You need to schedule a recurring BigQuery query that aggregates data from a partitioned table and writes the results to a new table every day at 03:00 UTC. You want a fully managed solution with minimal operational overhead. What should you use?
Easy81A BigQuery table has a REQUIRED column 'user_id' that now needs to accept NULL values due to upstream data changes. You want to alter the schema with minimal downtime and no data loss. What should you do?
Hard82Your company stores sensitive customer data in Cloud Storage. You need to inspect the data for personally identifiable information (PII) and de-identify it before sharing with a third party. Which Google Cloud service should you use?
Medium83You manage a Cloud Composer 2 environment that runs a DAG with a task using the BigQueryInsertJobOperator. The task occasionally fails with 'rateLimitExceeded' when submitting many jobs in parallel. You want to limit the number of concurrent BigQuery jobs submitted by this DAG without affecting other DAGs in the same environment. What should you do?
HardOther domains
All PDE exam domains
Frequently asked questions
- What does the Maintaining and Automating Data Workloads domain cover on the PDE exam?
- Be able to read operational symptoms and choose the correct Google Cloud control: DLP for de-identification, partition and cluster filters for BigQuery pruning, oldest_unacked_message_age for Pub/Sub lag, and preemptible-safe Dataproc sizing. The most important skill is matching each symptom to the exact metric or configuration that fixes it.
- How many questions are in this domain?
- This page lists all 83 Maintaining and Automating Data Workloads questions in the PDE question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Maintaining and Automating Data Workloads questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.