Courseiva

PDE Ingesting and Processing the Data Practice Question

A financial services company receives real-time stock trade data via Pub/Sub. They need to enrich each trade with reference data from a Cloud SQL table and write the results to BigQuery for real-time analytics. The enrichment must handle late-arriving data and ensure exactly-once processing. Which Dataflow streaming pipeline configuration should be used?

⚠ Common exam trap

PDE often tests the distinction between legacy streaming inserts (at-least-once, duplicates possible) and the Storage Write API's exactly-once mode — candidates who overlook the exactly-once requirement pick a template or legacy insert option.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Build a custom Dataflow pipeline using the Storage Write API with exactly-once semantics and a side input from Cloud SQL

The BigQuery Storage Write API with exactly-once semantics is the only option that satisfies both the exactly-once requirement and the need for enrichment with Cloud SQL reference data. A custom Dataflow pipeline can use the Storage Write API's exactly-once mode while applying a side input (or CoGroupByKey) to join the streaming trades with the Cloud SQL reference table. This combination handles late-arriving data via allowed lateness and windowing while guaranteeing no duplicate writes to BigQuery.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a Dataflow Flex Template that reads from Pub/Sub, joins in memory, and writes to BigQuery using legacy streaming inserts

    Why it's wrong here

    Legacy streaming inserts into BigQuery do not provide exactly-once semantics, and an in-memory join cannot handle late-arriving reference data without windowing and state. Flex Templates are correct when custom container dependencies are needed, not when the stem requires exactly-once BigQuery writes with late-data handling.

  • ✗

    Use Pub/Sub to BigQuery template with streaming inserts and a side input from Cloud SQL

    Why it's wrong here

    The Pub/Sub to BigQuery template performs no enrichment, and streaming inserts cannot guarantee exactly-once delivery; side inputs are not part of that template's configuration. Templates suit simple, code-free ingestion pipelines, so this would be correct only if no Cloud SQL join or exactly-once requirement existed.

  • ✓

    Build a custom Dataflow pipeline using the Storage Write API with exactly-once semantics and a side input from Cloud SQL

    Why this is correct

    The Storage Write API's exactly-once mode prevents duplicate writes to BigQuery, while a side input broadcasts the Cloud SQL reference table to workers for enrichment. Event-time windows with allowed lateness handle late-arriving trades, satisfying both stated requirements.

  • ✗

    Deploy a Dataproc Spark Streaming job that reads from Pub/Sub, enriches via JDBC, and writes to BigQuery

    Why it's wrong here

    Dataproc Spark Streaming lacks Dataflow's native Pub/Sub-to-BigQuery exactly-once sink and its built-in handling of late data via watermarks and triggers; JDBC enrichment also bypasses Cloud SQL's managed connectors. It is tempting when existing Spark skills or Hadoop ecosystems must be reused, but the stem demands Dataflow's streaming semantics.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.