Courseiva
hardMultiple Choice

PDE Practice Question: A company runs a critical real-time data pipeline…

A company runs a critical real-time data pipeline using Dataflow that ingests events from Cloud Pub/Sub, performs aggregations using sliding windows, and writes results to BigQuery. The pipeline is deployed in us-central1. The pipeline's latency has increased recently, and the Dataflow monitoring shows that the 'system lag' metric is consistently above 5 minutes. The pipeline is using Streaming Engine and has 10 workers with 4 vCPUs each. The pipeline processes approximately 100,000 events per second. The team has verified that the source Pub/Sub topic has sufficient publish throughput and the BigQuery table has no quota issues. The pipeline logs show that some workers are experiencing GC overhead limit exceeded errors. The pipeline code uses stateful processing with a custom keyed state for deduplication. What is the most likely cause of the increased latency?

⚠ Common exam trap

Google Cloud often tests the misconception that scaling workers (Option A) is the universal fix for latency, when in reality memory-related issues like GC overhead require tuning state management or worker resources, not just parallelism.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The stateful processing is causing large state sizes that lead to GC overhead; use a more efficient state backend or increase worker memory.

The GC overhead limit exceeded errors indicate that workers are spending too much time garbage collecting, which is a classic symptom of excessive heap memory usage. Stateful processing with custom keyed state for deduplication can cause large per-key state sizes, especially with sliding windows that maintain overlapping state for each key. This forces the JVM to constantly garbage collect, increasing system lag beyond 5 minutes. Using a more efficient state backend (e.g., reducing state size or using Dataflow's built-in deduplication) or increasing worker memory directly addresses the root cause.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The number of workers is insufficient; increasing to 20 workers will reduce latency.

    Why it's wrong here

    Adding workers does not reduce per-worker heap pressure; the GC overhead limit exceeded errors show each worker's JVM is exhausting memory, so extra workers may worsen key redistribution and shuffle costs. Scaling worker count is right when CPU or throughput per worker is the constraint, not memory.

  • ✓

    The stateful processing is causing large state sizes that lead to GC overhead; use a more efficient state backend or increase worker memory.

    Why this is correct

    Keyed deduplication state grows unbounded per key, inflating worker heap and triggering GC overhead limit errors that stall processing and raise system lag. Switching to a more efficient state backend or adding worker memory directly relieves the GC pressure causing the latency.

  • ✗

    The sliding window duration is too long; reducing it to 1 minute will improve performance.

    Why it's wrong here

    Sliding windows recompute overlapping panes, so shortening the duration changes emission frequency rather than the GC overhead limit exceeded errors, which arise from heap pressure in worker JVMs. A shorter window would be chosen when downstream consumers need fresher aggregates, not to relieve memory exhaustion.

  • ✗

    The deduplication logic is causing a bottleneck; removing it will reduce latency.

    Why it's wrong here

    Deduplication keyed state is the pipeline's designed correctness mechanism, and its memory footprint is what the GC overhead errors point to; removing it trades latency for duplicate results. Custom keyed state is the right tool when deduplication must survive restarts, so the fix is state sizing and worker memory, not deletion.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.