Courseiva

PDE Designing Data Processing Systems Practice Question

A financial services company needs to process credit card transactions in real time to detect fraudulent patterns. The pipeline must handle late-arriving data (up to 2 hours) and produce accurate results. They want to use a unified programming model that works for both batch and streaming. Which Google Cloud service should they use?

⚠ Common exam trap

The trap here is assuming that any streaming service can handle late data, but only Dataflow with Beam provides built-in event-time semantics and allowed lateness for accurate results.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Dataflow with Apache Beam

Cloud Dataflow with Apache Beam is the correct choice because it offers a unified batch and streaming programming model, supports event-time processing with watermarks and triggers, and can handle late-arriving data accurately through configurable allowed lateness. This makes it well-suited for real-time fraud detection with data arriving up to 2 hours late.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Pub/Sub with Cloud Functions

    Why it's wrong here

    Cloud Pub/Sub with Cloud Functions can trigger actions on messages, but it is not designed for complex event-time processing or handling late-arriving data with accuracy. Cloud Functions are stateless and have limited execution time, making them unsuitable for maintaining state across windows and processing late data. This combination lacks the advanced windowing capabilities needed.

  • ✗

    BigQuery with scheduled queries

    Why it's wrong here

    BigQuery scheduled queries run on a schedule (e.g., every hour), not in real time. They cannot process streaming data continuously or handle late-arriving data with event-time semantics. While BigQuery can analyze streaming data via the Storage Write API, it does not provide the unified programming model or windowing features required for this fraud detection scenario.

  • ✗

    Cloud Dataproc with Spark Streaming

    Why it's wrong here

    Cloud Dataproc with Spark Streaming can process real-time data, but it requires cluster management and does not provide a unified batch/streaming API as seamlessly as Beam. Handling late data with Spark Streaming requires manual watermarking and may not guarantee accurate results without complex tuning. It is less suitable for this scenario's requirement for a unified model.

  • ✓

    Cloud Dataflow with Apache Beam

    Why this is correct

    Cloud Dataflow, based on Apache Beam, provides a unified programming model for batch and streaming. It supports event-time processing, watermarks, and triggers to handle late-arriving data accurately. With windowing and allowed lateness, it can produce correct results even when data arrives up to 2 hours late, making it ideal for real-time fraud detection with late data.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.