hardMultiple Choice
PDE Practice Question: A Dataflow streaming pipeline reads from Pub/Sub,…
A Dataflow streaming pipeline reads from Pub/Sub, applies a ParDo that uses a side input from a BigQuery table (refreshed hourly), and writes to BigQuery. The side input is large and causes increased latency and worker OOM errors. Which design change solves this?
⚠ Common exam trap
Google Cloud often tests the misconception that increasing resources (like worker size or frequency) solves memory issues, when the real solution is to avoid storing large datasets in memory altogether by using an external lookup service.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use a stateful ParDo and store the lookup data in an external cache like Cloud Bigtable, performing lookups per element.
Moving the large lookup data to an external cache like Cloud Bigtable offloads memory pressure from workers, eliminating OOM errors. The side input broadcast approach keeps the entire dataset in each worker's memory, which causes OOM when the data is large. Using an external cache allows per-element lookups without storing the entire dataset in memory, reducing latency by avoiding broadcast overhead.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use a stateful ParDo and store the lookup data in an external cache like Cloud Bigtable, performing lookups per element.
Why this is correct
Moving the lookup data into Cloud Bigtable and querying it per element removes the large side input that Apache Beam materialises in worker memory, eliminating the OOM errors and latency. This satisfies the constraint that the hourly-refreshed BigQuery table is too large for side-input use.
- ✗
Increase the side input broadcast frequency to update more often.
Why it's wrong here
Refreshing more frequently increases how often the entire large side input is re-broadcast to every worker, worsening memory pressure and latency. Higher refresh rates suit small, fast-changing reference data where freshness matters and the dataset fits comfortably in worker memory.
- ✗
Split the pipeline into two: one to load the side input, the other to process main input.
Why it's wrong here
Splitting into two pipelines does not shrink the side input; the ParDo still materialises the full BigQuery table in worker memory, so OOM persists. This pattern suits separating ingestion from enrichment, not resolving a large side input, which needs a different access mechanism such as keyed lookups.
- ✗
Use smaller worker machine types to distribute memory across more workers.
Why it's wrong here
Smaller workers reduce per-worker memory, so holding the large side input causes OOM sooner, not later. Smaller machine types suit lightweight, CPU-bound transforms with small state; a large side input requires either larger workers or avoiding full broadcast altogether.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.