PDE Designing Data Processing Systems Practice Question
Which Google Cloud service provides a serverless Spark environment where you can run Spark jobs without provisioning or managing a cluster?
⚠ Common exam trap
PDE often tests the distinction between serverless Spark (Dataproc Serverless) and other serverless data services like Dataflow, causing candidates to confuse the underlying processing engines.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dataproc Serverless
Dataproc Serverless is a Google Cloud service that allows you to run Spark jobs without provisioning or managing a cluster. It automatically scales resources and charges only for the duration of the job, making it ideal for serverless Spark workloads. This matches the requirement exactly.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Dataflow
Why it's wrong here
Dataflow runs Apache Beam pipelines, not Apache Spark jobs, so it cannot execute Spark code. It is the right pick for unified batch and streaming Beam pipelines, but the stem explicitly asks for a serverless Spark environment.
- ✓
Dataproc Serverless
Why this is correct
Dataproc Serverless runs Spark workloads on managed, ephemeral infrastructure, so no cluster provisioning, sizing or teardown is required. This satisfies the serverless constraint directly, unlike Dataproc on Compute Engine, which requires you to create and manage clusters yourself.
- ✗
Dataprep
Why it's wrong here
Dataprep is a data-preparation service that runs on Dataflow, not Spark, so it cannot execute Spark jobs. It is tempting because it processes data without cluster management, but its purpose is visually transforming datasets for analysis. Dataprep would be the right choice for cleaning and preparing data, not for running Spark workloads.
- ✗
Cloud Data Fusion
Why it's wrong here
Cloud Data Fusion is a codeless data-integration pipeline builder, not a Spark execution environment. It is chosen when visually designing ETL pipelines, but it does not offer the serverless Spark job execution the stem requires.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.