Courseiva

PDE Designing Data Processing Systems Practice Question

Which Google Cloud service provides a fully managed, serverless Spark environment without requiring cluster provisioning?

⚠ Common exam trap

PDE often tests the distinction between serverless and managed services, and candidates may confuse Dataflow (Beam) with Dataproc Serverless (Spark) or think Dataproc on GKE is serverless when it still requires cluster management.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dataproc Serverless

Dataproc Serverless is a fully managed, serverless Spark environment on Google Cloud that eliminates the need to provision or manage clusters. It automatically scales resources and charges only for the duration of the workload, making it ideal for running Spark jobs without infrastructure overhead.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Dataproc on GKE

    Why it's wrong here

    Dataproc on GKE runs Spark on a self-managed Kubernetes cluster, so you still provision and size node pools. It suits teams wanting Spark with container orchestration control. The serverless requirement is met by Dataproc Serverless, which provisions Spark capacity automatically without cluster management.

  • ✗

    Dataflow

    Why it's wrong here

    Dataflow is a serverless Apache Beam pipeline runner for stream and batch processing, not a Spark execution engine. It would be the correct choice for unified stream and batch pipelines written with the Beam SDK, where no Spark runtime or cluster management is needed.

  • ✓

    Dataproc Serverless

    Why this is correct

    Dataproc Serverless runs Spark workloads without any cluster provisioning, directly satisfying the stem's serverless requirement. Unlike standard Dataproc, which needs manual cluster creation and sizing, it provisions ephemeral compute automatically per job, so no infrastructure management is needed. This makes it the fully managed Spark environment the question describes.

  • ✗

    Cloud Data Fusion

    Why it's wrong here

    Cloud Data Fusion is a managed codeless ETL service built on CDAP, orchestrating pipelines rather than executing Spark jobs directly. It fits visual data integration across sources. A serverless Spark runtime without cluster provisioning is provided by Dataproc Serverless, not Data Fusion.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.