PDE Designing Data Processing Systems Practice Question
A retail company ingests point-of-sale events from thousands of stores into Cloud Pub/Sub. They need to process these events in a streaming Dataflow pipeline that enriches each event with store metadata from a slowly changing BigQuery table. The enrichment table is updated only once per day. The pipeline must minimize latency and avoid querying BigQuery for every event. Which approach should they use?
⚠ Common exam trap
The trap here is assuming that a streaming pipeline must query BigQuery for each event or use a join transform, when side inputs are designed for exactly this kind of enrichment.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Load the BigQuery metadata table into a side input as a key-value map, refreshing it periodically via a scheduled pipeline.
The correct approach is to load the slowly changing metadata into a side input that is refreshed periodically. This allows the streaming pipeline to enrich each event without querying BigQuery per event, keeping latency low. Side inputs are a standard Dataflow pattern for enriching streaming data with reference data that changes infrequently, and they avoid the high cost and latency of per-element external calls.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a CoGroupByKey transform to join the streaming events with a bounded read of the metadata table.
Why it's wrong here
CoGroupByKey is used to join multiple PCollections of the same key type, but it requires both inputs to be keyed and windowed appropriately. A bounded read of the metadata table would not be continuously updated, and joining a bounded side with an unbounded stream can lead to windowing mismatches. This approach does not provide a low-latency, continuously refreshed enrichment and may drop or delay events.
- ✓
Load the BigQuery metadata table into a side input as a key-value map, refreshing it periodically via a scheduled pipeline.
Why this is correct
Side inputs in Dataflow allow you to supply additional data to each element in a pipeline without querying an external system per element. By loading the slowly changing metadata into a side input and refreshing it periodically, the pipeline can enrich events with minimal latency and without hitting BigQuery for each event. This matches the requirement to minimize latency and avoid per-event queries.
- ✗
Configure the pipeline to call the BigQuery Storage Read API for each event to fetch the latest metadata.
Why it's wrong here
The BigQuery Storage Read API is optimized for high-throughput reads of large datasets, not for per-event lookups. Calling it for each event would introduce significant latency and cost, and it is not designed for low-latency, single-row queries. This would violate the requirement to minimize latency and avoid querying BigQuery for every event, making it a poor choice for streaming enrichment.
- ✗
Use BigQueryIO.Read with a query that joins the events to the metadata table, and apply the join in the pipeline.
Why it's wrong here
BigQueryIO.Read is a bounded source, not designed for per-element lookups in a streaming pipeline. Using it to join a slowly changing table would require periodic full reads or complex side-input patterns. It would not provide low-latency enrichment and could lead to excessive BigQuery scans. This approach does not meet the requirement to avoid querying BigQuery for every event and would add significant latency.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.