PDE Designing Data Processing Systems Practice Question
A financial services company needs to process credit card transactions in real time to detect fraudulent patterns. The pipeline must handle late-arriving data (up to 2 hours) and produce accurate results. They want to use a unified programming model that works for both batch and streaming. Which Google Cloud service should they use?
⚠ Common exam trap
The trap here is assuming that any streaming service can handle late data, but only Dataflow with Beam provides built-in event-time semantics and allowed lateness for accurate results.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Dataflow with Apache Beam
Cloud Dataflow with Apache Beam is the correct choice because it offers a unified batch and streaming programming model, supports event-time processing with watermarks and triggers, and can handle late-arriving data accurately through configurable allowed lateness. This makes it well-suited for real-time fraud detection with data arriving up to 2 hours late.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Pub/Sub with Cloud Functions
Why it's wrong here
Cloud Pub/Sub with Cloud Functions can trigger actions on messages, but it is not designed for complex event-time processing or handling late-arriving data with accuracy. Cloud Functions are stateless and have limited execution time, making them unsuitable for maintaining state across windows and processing late data. This combination lacks the advanced windowing capabilities needed.
- ✗
BigQuery with scheduled queries
Why it's wrong here
BigQuery scheduled queries run on a schedule (e.g., every hour), not in real time. They cannot process streaming data continuously or handle late-arriving data with event-time semantics. While BigQuery can analyze streaming data via the Storage Write API, it does not provide the unified programming model or windowing features required for this fraud detection scenario.
- ✗
Cloud Dataproc with Spark Streaming
Why it's wrong here
Cloud Dataproc with Spark Streaming can process real-time data, but it requires cluster management and does not provide a unified batch/streaming API as seamlessly as Beam. Handling late data with Spark Streaming requires manual watermarking and may not guarantee accurate results without complex tuning. It is less suitable for this scenario's requirement for a unified model.
- ✓
Cloud Dataflow with Apache Beam
Why this is correct
Cloud Dataflow, based on Apache Beam, provides a unified programming model for batch and streaming. It supports event-time processing, watermarks, and triggers to handle late-arriving data accurately. With windowing and allowed lateness, it can produce correct results even when data arrives up to 2 hours late, making it ideal for real-time fraud detection with late data.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.