PDE Designing Data Processing Systems Practice Question
Which Google Cloud service provides a fully managed, serverless Spark environment without requiring cluster provisioning?
⚠ Common exam trap
PDE often tests the distinction between serverless and managed services, and candidates may confuse Dataflow (Beam) with Dataproc Serverless (Spark) or think Dataproc on GKE is serverless when it still requires cluster management.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dataproc Serverless
Dataproc Serverless is a fully managed, serverless Spark environment on Google Cloud that eliminates the need to provision or manage clusters. It automatically scales resources and charges only for the duration of the workload, making it ideal for running Spark jobs without infrastructure overhead.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Dataproc on GKE
Why it's wrong here
Dataproc on GKE runs Spark on a self-managed Kubernetes cluster, so you still provision and size node pools. It suits teams wanting Spark with container orchestration control. The serverless requirement is met by Dataproc Serverless, which provisions Spark capacity automatically without cluster management.
- ✗
Dataflow
Why it's wrong here
Dataflow is a serverless Apache Beam pipeline runner for stream and batch processing, not a Spark execution engine. It would be the correct choice for unified stream and batch pipelines written with the Beam SDK, where no Spark runtime or cluster management is needed.
- ✓
Dataproc Serverless
Why this is correct
Dataproc Serverless runs Spark workloads without any cluster provisioning, directly satisfying the stem's serverless requirement. Unlike standard Dataproc, which needs manual cluster creation and sizing, it provisions ephemeral compute automatically per job, so no infrastructure management is needed. This makes it the fully managed Spark environment the question describes.
- ✗
Cloud Data Fusion
Why it's wrong here
Cloud Data Fusion is a managed codeless ETL service built on CDAP, orchestrating pipelines rather than executing Spark jobs directly. It fits visual data integration across sources. A serverless Spark runtime without cluster provisioning is provided by Dataproc Serverless, not Data Fusion.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.