Courseiva

PDE · topic practice

Maintaining and Automating Data Workloads practice questions

This domain covers keeping production data pipelines healthy on Google Cloud: monitoring and alerting with Cloud Monitoring, tuning BigQuery and Dataproc performance, protecting data with DLP and IAM, and automating recurring work with Cloud Composer, Cloud Scheduler, and Dataflow. Questions present operational symptoms and ask you to pick the right service, metric, or configuration.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Maintaining and Automating Data Workloads

What the exam tests

What to know about Maintaining and Automating Data Workloads

Be able to read operational symptoms and choose the correct Google Cloud control: DLP for de-identification, partition and cluster filters for BigQuery pruning, oldest_unacked_message_age for Pub/Sub lag, and preemptible-safe Dataproc sizing. The most important skill is matching each symptom to the exact metric or configuration that fixes it.

Selecting Cloud DLP infoType detectors and de-identification transforms (masking, tokenization) for sensitive data in Cloud Storage and BigQuery

Diagnosing BigQuery scans with INFORMATION_SCHEMA.JOBS, partitioning, clustering, and materialized views

Creating Cloud Monitoring alerting policies on Pub/Sub metrics such as oldest_unacked_message_age with appropriate filters

Tuning Dataproc clusters, including preemptible worker ratios, autoscaling policies, and Spark shuffle configuration

Watch out for

Common Maintaining and Automating Data Workloads exam traps

  • ▸Assuming a partitioned table is always pruned; filters must reference the partitioning column directly, or BigQuery scans the whole table.
  • ▸Alerting on the wrong Pub/Sub metric, such as message count instead of oldest_unacked_message_age, so slow-consumer backlogs go unnoticed.
  • ▸Using too many preemptible Dataproc workers without fallback, causing shuffle failures and slow jobs when VMs are reclaimed.

Practice set

Maintaining and Automating Data Workloads questions

20 questions · select your answer, then reveal the explanation

A company uses Dataflow streaming pipelines to process real-time events. They notice increasing system lag over time. Which two Cloud Monitoring metrics should be examined to diagnose the cause?

A data team needs to share a BigQuery dataset with another business unit. They want to provide a point-in-time snapshot of the data without incurring additional storage costs for the copy. Which BigQuery feature should they use?

A company runs a Dataproc cluster for ETL jobs that process data nightly. They want to reduce costs while maintaining performance. Which strategy is MOST effective?

A company uses Cloud Composer for pipeline orchestration. They need to define task dependencies where Task B and Task C can run in parallel after Task A, and Task D must run after both B and C complete. How should they define the DAG?

A data engineer notices that BigQuery queries are slower than expected. They want to identify the most expensive stages in the query execution. Which tool or command should they use?

A data engineer needs to migrate a schema from BigQuery where a column is currently REQUIRED and needs to become NULLABLE. Which TWO statements are correct? (Choose 2)

A company runs BigQuery workloads with varying demand. They want to use flat-rate pricing with baseline slots and the ability to burst during peak times. Which TWO actions should they take? (Choose 2)

A company uses Cloud Composer (Airflow) to orchestrate pipelines. They want to implement a pattern where a task polls for a file arrival in Cloud Storage and then triggers subsequent tasks. Which THREE Airflow concepts are essential? (Choose 3)

A company runs a streaming Dataflow pipeline that reads from Pub/Sub, enriches data with a side input from BigQuery, and writes to BigQuery. After updating the pipeline code (adding a new field to the output), the engineer notices that the new pipeline version is not picking up the updated code because the job was started from a template. The engineer wants to update the streaming pipeline without draining it. What should the engineer do?

You are monitoring a streaming Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. In Cloud Monitoring, you notice that the 'system_lag' metric is increasing over time and now exceeds 10 minutes. The 'data_watermark' metric shows a steady lag. What is the most likely cause of the increasing system lag?

A company wants to share a large BigQuery dataset with a partner for analysis. The partner needs read-only access to a specific snapshot of the data as of a certain point in time, and the company wants to avoid additional storage costs for the partner. What is the most cost-effective approach?

Your organization has a BigQuery flat-rate reservation with 2000 slots. During peak hours, query performance degrades because concurrent queries exceed the available slots. You want to handle these bursts without changing the base reservation. What should you do?

A data engineer is designing a Dataflow pipeline that reads from a Kafka topic (using Pub/Sub for Kafka) and writes to BigQuery. The data schema may change over time, with new fields appearing. The engineer wants to handle schema drift automatically without failing the pipeline. Which approach should the engineer use?

You are troubleshooting a Dataproc cluster that runs nightly Spark jobs. The jobs are failing with out-of-memory errors. You want to reduce costs while fixing the issue. Which combination of actions should you take? (Select the BEST answer.)

A data engineer needs to set up a Dataplex data quality scan to run weekly on a BigQuery table. The scan should check that: (1) the 'email' column is not null, (2) the 'age' column is between 0 and 120, and (3) the 'country_code' column matches a list of valid ISO codes. Which TWO Dataplex features should the engineer use?

A data engineer is building a Cloud Workflows workflow that orchestrates multiple Cloud Functions and API calls. The workflow should handle transient failures with retries and send a notification to a Pub/Sub topic if the workflow ultimately fails. Which THREE steps should the engineer include in the workflow definition?

A company uses BigQuery flat-rate pricing with 500 slots purchased as a committed use discount. During peak hours, they need additional capacity but do not want to buy more committed slots. They have a secondary project used for ad-hoc queries by analysts. How can they provide burst capacity to the primary project during peak times without increasing committed spend?

Your team uses Cloud Composer to run Apache Airflow DAGs. One DAG uses a BigQueryInsertJobOperator to run a query and then uses BigQueryCheckOperator to verify the results. The DAG is failing intermittently because the query result is not ready when the check operator runs. How should you modify the DAG to ensure the check operator runs only after the query completes successfully?

A data engineer needs to share a large BigQuery table with a different team, but wants to minimize storage costs. The table is 1 TB in size and is updated daily. The other team only needs read access to the data as of a specific point in time (e.g., end of each day). Which BigQuery feature should be used to provide a read-only copy without duplicating the entire table?

You are designing a Dataflow pipeline that reads from Pub/Sub, performs transformations, and writes to BigQuery. The pipeline must handle schema changes in the incoming data (e.g., new fields appearing). The BigQuery schema should evolve automatically to accept new fields without failing. Which approach should you use?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Maintaining and Automating Data Workloads sessions

Start a Maintaining and Automating Data Workloads only practice session

Every question in these sessions is drawn from the Maintaining and Automating Data Workloads domain — nothing else.

Related practice questions

Related PDE topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the PDE exam test about Maintaining and Automating Data Workloads?
Be able to read operational symptoms and choose the correct Google Cloud control: DLP for de-identification, partition and cluster filters for BigQuery pruning, oldest_unacked_message_age for Pub/Sub lag, and preemptible-safe Dataproc sizing. The most important skill is matching each symptom to the exact metric or configuration that fixes it.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Maintaining and Automating Data Workloads questions in a focused session?
Yes — the session launcher on this page draws every question from the Maintaining and Automating Data Workloads domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other PDE topics?
Use the topic links above to move to related areas, or go back to the PDE question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the PDE exam covers. They are not copied from any real exam or dump site.