hardMultiple Choice
PDE Practice Question: A financial services company operates a real-time…
A financial services company operates a real-time fraud detection pipeline using Apache Beam running on Google Cloud Dataflow. The pipeline reads transactions from Pub/Sub, enriches them with customer data from Bigtable, runs a machine learning model with side inputs from a Redis cluster, and writes results to BigQuery for downstream reporting. The data must be processed with exactly-once semantics to avoid duplicate fraud alerts or missing transactions. The pipeline currently uses a global window with 5-minute accumulation, but the team is experiencing high latency and occasional duplicates when the model side input is updated (triggered every 15 minutes via a WatchTransform). Additionally, the pipeline has a dead letter queue that outputs failed records to a separate Pub/Sub topic, but these records are never reprocessed. The team needs to ensure high reliability and data quality. Which course of action should the team take to improve solution quality?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Implement sliding windows of 5 minutes with a 2-minute allowed lateness, use side inputs with periodic refreshes using the .withUpdateFrequency transformation, and set up a Cloud Function to automatically replay dead letter records back to the main Pub/Sub topic after fixing the issue.
Sliding windows with 2-minute allowed lateness handle late-arriving data without causing duplicates, and the .withUpdateFrequency transformation refreshes the side input every 15 minutes without triggering reprocessing, reducing latency. Replaying dead letter records via a Cloud Function ensures data completeness. Option A is incorrect because fixed windows with session gaps do not address side input latency and may lose late events. Option B is incorrect because batch processing is unsuitable for real-time fraud detection and introduces significant latency. Option D is incorrect because custom triggers with early firings can cause duplicates due to side input updates, and the Read transform does not efficiently handle periodic model refreshes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use fixed windows with a 10-minute duration and session gap of 2 minutes, disable side input caching, and log all dead letter records to Cloud Storage for manual inspection.
Why it's wrong here
Fixed windows with session gaps do not reduce side input latency and may cause data loss due to window boundaries.
- ✗
Switch to a batch processing approach that runs every minute using Cloud Composer, with data loaded from Pub/Sub into BigQuery and then processed with Dataproc to run the model.
Why it's wrong here
Batch processing every minute introduces unacceptable latency for real-time fraud detection and does not leverage Dataflow's streaming capabilities.
- ✓
Implement sliding windows of 5 minutes with a 2-minute allowed lateness, use side inputs with periodic refreshes using the .withUpdateFrequency transformation, and set up a Cloud Function to automatically replay dead letter records back to the main Pub/Sub topic after fixing the issue.
Why this is correct
Sliding windows with allowed lateness capture late data, periodic side input refreshes reduce model update latency, and automatic reprocessing of dead letters ensures exactly-once semantics and data completeness.
- ✗
Keep the global window but use a custom trigger with early firings every 30 seconds and a late-firing threshold of 1 minute, and configure the side input to be broadcast every 5 minutes using a Read transform.
Why it's wrong here
Global window with custom triggers can cause duplicates if not perfectly tuned, and broadcasting side inputs every 5 minutes via Read is less efficient than incremental updates.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.