PDE Designing Data Processing Systems Practice Question
A financial services firm is designing a data processing system on Google Cloud that must ingest change data capture (CDC) streams from an on-premises PostgreSQL database into BigQuery with sub-minute latency, preserve the ordering of changes per primary key, and apply updates and deletes so that BigQuery reflects the current state of each row. The source database cannot be modified to add triggers. Which two design elements should you include? (Choose two.)
⚠ Common exam trap
The trap here is assuming that streaming inserts into BigQuery can apply updates and deletes, when the streaming API is append-only and mutations require a merge pattern.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Datastream to read the PostgreSQL write-ahead log and stream changes into Cloud Storage or BigQuery.
Log-based CDC through Datastream reads the PostgreSQL write-ahead log without touching the source schema, and the staging-plus-merge destination pattern keyed on the primary key is what turns a change stream into correct current-state rows in BigQuery, including deletes. Together they meet the latency, ordering, and mutation requirements.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Add database triggers to the PostgreSQL source that publish change events to a Pub/Sub topic.
Why it's wrong here
The scenario states the source database cannot be modified, so adding triggers is not permitted, and triggers would also add write overhead to the transactional system. Pub/Sub is a valid transport, but the ingestion mechanism must not depend on source-side schema changes, which rules out this approach regardless of its other merits.
- ✓
Use Datastream to read the PostgreSQL write-ahead log and stream changes into Cloud Storage or BigQuery.
Why this is correct
Datastream performs log-based CDC by reading the PostgreSQL write-ahead log, so no triggers or schema changes are needed on the source, satisfying the constraint that the database cannot be modified. It delivers low-latency change records that can land in Cloud Storage or be written directly into BigQuery, providing the ingestion mechanism the design requires.
- ✗
Use a federated BigQuery external table that queries the on-premises PostgreSQL instance directly at read time.
Why it's wrong here
Federated queries read the current state of the source on demand and cannot capture intermediate changes, ordering, or deletes over time. They also push query load onto the production database and cannot deliver sub-minute incremental replication into BigQuery, so this does not meet the CDC latency or ordering requirements.
- ✗
Enable BigQuery streaming inserts with insertId deduplication to apply deletes directly to the target table.
Why it's wrong here
Streaming inserts with insertId only deduplicate best-effort retries of the same record; they cannot express a delete or an update to an existing row. BigQuery storage is append-oriented, so mutations must be applied through DML or the merge pattern, not through the streaming API, which is why this element does not satisfy the delete requirement.
- ✓
Configure the BigQuery destination to use the staging and merge pattern with a primary key so deletes and updates are applied.
Why this is correct
Datastream's BigQuery destination writes change records to a staging table and then merges them into the target using the primary key, which is what makes updates and deletes take effect rather than accumulating as append-only history. Declaring the primary key is what lets the merge correctly collapse multiple changes to the same row into the latest state.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.