Courseiva
hardMultiple Select

PDE Practice Question: Which TWO statements about designing a data…

Which TWO statements about designing a data processing pipeline on Google Cloud are correct? (Choose 2.)

⚠ Common exam trap

Google Cloud often tests the distinction between fully managed services (like BigQuery for warehousing) and managed cluster services (like Dataproc), as well as the limitations of Pub/Sub ordering guarantees, to see if candidates confuse operational databases with analytical systems.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Data Fusion allows you to build and manage data pipelines visually without writing code.

Option D is correct because Cloud Data Fusion is a fully managed, code-free ETL/ELT service built on the open-source CDAP framework, providing a visual drag-and-drop pipeline designer and a broad library of connectors and transformations, so pipelines can be built and managed without writing code. Option E is correct because Dataflow, based on Apache Beam, uses a unified programming model in which the same pipeline code can run in batch or streaming mode, with the runner handling windowing, triggers, and watermarks for both. Option A is wrong because Pub/Sub does not guarantee global ordering across all subscribers; ordering is only provided per ordering key within a region when message ordering is explicitly enabled. Option B is wrong because Cloud Bigtable is a wide-column NoSQL database optimized for high-throughput, low-latency reads and writes at scale, not for data warehousing or SQL analytics. Option C is wrong because Dataproc is a managed Apache Hadoop/Spark service for running clusters and jobs, whereas BigQuery is the fully managed data warehousing and analytics service.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Pub/Sub guarantees message ordering across all subscribers globally.

    Why it's wrong here

    Pub/Sub only guarantees ordering within a single region when using ordering keys.

  • ✗

    Cloud Bigtable is ideal for data warehousing and SQL analytics.

    Why it's wrong here

    Bigtable is a NoSQL database for low-latency operations, not for SQL analytics.

  • ✗

    Dataproc is the best choice for fully managed data warehousing and analytics.

    Why it's wrong here

    Dataproc is for running Hadoop/Spark clusters; BigQuery is for data warehousing.

  • ✓

    Cloud Data Fusion allows you to build and manage data pipelines visually without writing code.

    Why this is correct

    Cloud Data Fusion provides a graphical, code-free interface for building and managing pipelines, satisfying the stem's requirement for visual design without coding. Its underlying CDAP engine handles orchestration and execution, so teams can create repeatable data integration workflows through drag-and-drop rather than hand-written code.

  • ✓

    Dataflow supports both batch and streaming modes in a single pipeline model.

    Why this is correct

    Dataflow's Apache Beam model uses the same pipeline code for both batch and streaming execution, differing only in the source and windowing semantics. This unified model satisfies the stem's requirement that a single pipeline design handle bounded and unbounded data without separate implementations.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.