Courseiva

PDE Maintaining and Automating Data Workloads Practice Question

You are running a streaming pipeline with Dataflow that reads from Pub/Sub and writes to BigQuery. You notice that the system lag metric is increasing over time, indicating that messages are taking longer to process. What is the most likely cause and how should you address it?

⚠ Common exam trap

The trap is picking a source-side fix (partitions) that applies to Kafka, not Pub/Sub, or blaming the sink; the exam expects you to recognize that increasing system lag in Dataflow is typically a worker capacity issue solved by autoscaling.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The Dataflow workers are CPU-bound; increase the number of workers or adjust autoscaling settings.

System lag in Dataflow measures the age of the oldest unprocessed message, and a steadily increasing value indicates the pipeline cannot keep up with input. The most common cause is CPU-bound workers, so scaling out workers or tuning autoscaling (e.g., max workers, worker type) restores throughput. This directly addresses the bottleneck rather than a downstream symptom.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The source Pub/Sub topic has insufficient throughput; increase the number of partitions.

    Why it's wrong here

    Pub/Sub auto-scales; system lag in Dataflow indicates processing bottleneck, not source throughput.

  • ✓

    The Dataflow workers are CPU-bound; increase the number of workers or adjust autoscaling settings.

    Why this is correct

    High system lag suggests worker resources are insufficient; adding workers reduces lag.

  • ✗

    The BigQuery destination table has too many columns; reduce the number of columns.

    Why it's wrong here

    Column count does not typically cause system lag; it's a processing latency issue.

  • ✗

    The pipeline uses a batch transform that should be replaced with a streaming transform.

    Why it's wrong here

    Streaming transforms are already used; system lag is about resource contention, not transform type.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.