PDE Designing Data Processing Systems Practice Question
A startup needs a fully managed, serverless Spark service to run occasional data processing jobs without managing clusters. They want to pay only for the resources used during job execution. Which Google Cloud service should they use?
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dataproc Serverless
Dataproc Serverless provides a serverless Spark environment where you pay per job execution. Cloud Data Fusion is for visual ETL. Dataproc is managed but not serverless. Dataflow is serverless for Beam, not Spark.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Dataproc Serverless
Why this is correct
Dataproc Serverless runs Spark workloads without provisioning or managing a cluster, allocating resources only while the job executes and charging for that consumption. This matches the stem's requirements for occasional jobs, no cluster management and pay-per-use billing, unlike a standard Dataproc cluster.
- ✗
Dataflow
Why it's wrong here
Dataflow is a managed Apache Beam runner for streaming and batch pipelines, not a serverless Spark service; it cannot execute Spark jobs. It suits event-driven transformation pipelines. The scenario's requirement for Spark with per-job, clusterless billing points to Dataproc Serverless instead.
- ✗
Cloud Data Fusion
Why it's wrong here
Cloud Data Fusion is a codeless ETL orchestration layer that runs pipelines on Dataproc or Dataflow; it does not itself execute Spark jobs serverlessly with per-job billing. It suits visual data-integration workflows, not direct Spark execution.
- ✗
Dataproc
Why it's wrong here
Dataproc provisions clusters of virtual machines, so the startup would manage nodes and pay for idle capacity between occasional jobs. It is tempting because Dataproc runs Spark and suits sustained workloads where cluster tuning is wanted, but the stem requires serverless execution with per-job billing, which Dataproc Serverless provides instead.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.