Courseiva

PDE Side Input Practice Question

A data pipeline ingests streaming events into Pub/Sub and needs to join them with a slowly updating reference table (few thousand rows) from a Cloud Storage CSV file. The pipeline runs on Dataflow with Apache Beam. Which approach is most cost-effective and operationally simple?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a side input that reads the CSV once and broadcasts it to all workers

Option B is correct because a Beam side input reads the small Cloud Storage CSV once and broadcasts the reference table to all workers, letting each streaming element be enriched in memory without per-event external calls, which is both cheap and operationally simple. Since the table has only a few thousand rows, it easily fits in worker memory as a side input. Option A is wrong because issuing a BigQuery query per event is slow, expensive, and adds an external dependency. Option C is wrong because routing events through Cloud SQL and doing SQL JOINs adds infrastructure and cost. Option D is wrong because CoGroupByKey requires reading the CSV into a batch PCollection per window and shuffling both sides, which is more complex and less efficient than a side input for a small, slowly changing table.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Read the CSV in a DoFn and perform a BigQuery query each time an event is processed

    Why it's wrong here

    Querying BigQuery per event adds a network round trip and repeated query cost for every element, defeating the cost-effectiveness requirement. It would suit infrequent lookups where the reference data changes constantly, but a few thousand static rows belong in a side input or cached broadcast.

  • ✓

    Use a side input that reads the CSV once and broadcasts it to all workers

    Why this is correct

    A side input reads the CSV once and broadcasts the small reference table to every worker, so each streaming element joins locally without per-element Cloud Storage reads or external lookups. This satisfies the cost-effectiveness and operational-simplicity constraints for a few-thousand-row table, avoiding the latency and expense of querying an external store per event.

  • ✗

    Implement a custom sink that writes events to Cloud SQL and performs a SQL JOIN there

    Why it's wrong here

    Writing every event to Cloud SQL adds per-row write and query costs plus a database to operate, and the join happens outside Beam, breaking windowing and exactly-once semantics. Cloud SQL suits transactional application data; joining a small reference table in-pipeline via a side input avoids this overhead entirely.

  • ✗

    Use CoGroupByKey to join the stream and batch PCollections by a common key after reading the CSV into a batch PCollection each window

    Why it's wrong here

    CoGroupByKey works for two bounded PCollections or one unbounded with global window; this approach would require windowing on the batch side and is not the simplest or most cost-effective for a small table.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.