mediumMultiple Choice
PDE Practice Question: Migrating their on-premises Apache Spark jobs to…
A company is migrating their on-premises Apache Spark jobs to Dataproc. They want to minimize code changes and take advantage of serverless infrastructure. Which Dataproc feature should they use?
⚠ Common exam trap
Google Cloud often tests the distinction between 'serverless' and 'managed' services; the trap here is that candidates may confuse Dataproc Workflow Templates or Jobs API with serverless capabilities, but those still require cluster management, whereas Dataproc Serverless Spark truly abstracts the infrastructure.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dataproc Serverless Spark
Dataproc Serverless Spark is the correct choice because it allows the company to run Spark workloads without provisioning or managing clusters, minimizing code changes by using the same Spark APIs and libraries. This serverless infrastructure automatically scales resources and handles failures, aligning with the goal of reducing operational overhead while maintaining compatibility with existing Spark jobs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Dataproc clusters with preemptible VMs
Why it's wrong here
Preemptible VMs cut compute cost on standard provisioned clusters, but the cluster still exists and must be sized and managed, so no serverless benefit is gained. This suits cost-sensitive batch workloads tolerant of instance preemption, not migration with minimal operational overhead.
- ✗
Dataproc Workflow Templates
Why it's wrong here
Workflow Templates orchestrate and schedule sequences of existing Dataproc jobs; they still run on provisioned clusters, so they deliver neither serverless execution nor reduced code changes. They suit managing multi-step, dependency-linked batch pipelines on persistent infrastructure.
- ✓
Dataproc Serverless Spark
Why this is correct
Dataproc Serverless Spark runs PySpark and Spark SQL workloads without provisioning clusters, so existing Apache Spark code executes with minimal modification. It directly satisfies both stem constraints: reducing code changes and consuming serverless infrastructure, since compute scales automatically and no cluster management is required.
- ✗
Dataproc Jobs API with custom machine types
Why it's wrong here
The Jobs API submits work to clusters you provision, and custom machine types merely tune node sizing; neither removes infrastructure management. This suits teams needing precise resource control over long-running clusters, not a serverless target requiring no code changes.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.