Courseiva

PDE Designing Data Processing Systems Practice Question

An organization is implementing a data lake on Google Cloud using Cloud Storage. They need to process both batch and streaming data with a unified pipeline. The team has experience with Apache Beam. Which architecture should they use to minimize operational overhead?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Kappa architecture with Cloud Dataflow using the same pipeline for batch and streaming

Kappa architecture uses a single streaming pipeline for both batch and streaming, simplifying operations. Dataflow implements Beam and supports both modes.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Kappa architecture with Cloud Dataflow using the same pipeline for batch and streaming

    Why this is correct

    Cloud Dataflow runs the same Apache Beam pipeline for both bounded and unbounded sources, so one codebase serves batch and streaming. This satisfies the unified-pipeline requirement while remaining fully managed, minimising operational overhead for the Beam-experienced team.

  • ✗

    Use Cloud Dataproc for batch and Cloud Dataflow for streaming

    Why it's wrong here

    Cloud Dataproc runs Hadoop and Spark clusters requiring provisioning, scaling and patching, adding operational overhead, and it does not share a pipeline with Dataflow. Dataproc suits migrating existing Spark or Hadoop workloads, but here the team's Apache Beam experience and the unified-pipeline requirement point to Dataflow alone.

  • ✗

    Lambda architecture with Cloud Dataflow for batch and Cloud Pub/Sub for streaming

    Why it's wrong here

    Lambda architecture maintains two separate processing paths — a batch layer and a speed layer — plus a serving layer to merge them, duplicating logic and operational burden. It is tempting because it handles batch and streaming, but Apache Beam's unified model already processes both with one pipeline, which is the stated goal.

  • ✗

    Use Cloud Data Fusion for both batch and streaming

    Why it's wrong here

    Data Fusion is mainly batch; streaming support is limited.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.