Courseiva

PDE Designing Data Processing Systems Practice Question

Which Google Cloud service provides a serverless Spark environment where you can run Spark jobs without provisioning or managing a cluster?

⚠ Common exam trap

PDE often tests the distinction between serverless Spark (Dataproc Serverless) and other serverless data services like Dataflow, causing candidates to confuse the underlying processing engines.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dataproc Serverless

Dataproc Serverless is a Google Cloud service that allows you to run Spark jobs without provisioning or managing a cluster. It automatically scales resources and charges only for the duration of the job, making it ideal for serverless Spark workloads. This matches the requirement exactly.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Dataflow

    Why it's wrong here

    Dataflow runs Apache Beam pipelines, not Apache Spark jobs, so it cannot execute Spark code. It is the right pick for unified batch and streaming Beam pipelines, but the stem explicitly asks for a serverless Spark environment.

  • ✓

    Dataproc Serverless

    Why this is correct

    Dataproc Serverless runs Spark workloads on managed, ephemeral infrastructure, so no cluster provisioning, sizing or teardown is required. This satisfies the serverless constraint directly, unlike Dataproc on Compute Engine, which requires you to create and manage clusters yourself.

  • ✗

    Dataprep

    Why it's wrong here

    Dataprep is a data-preparation service that runs on Dataflow, not Spark, so it cannot execute Spark jobs. It is tempting because it processes data without cluster management, but its purpose is visually transforming datasets for analysis. Dataprep would be the right choice for cleaning and preparing data, not for running Spark workloads.

  • ✗

    Cloud Data Fusion

    Why it's wrong here

    Cloud Data Fusion is a codeless data-integration pipeline builder, not a Spark execution environment. It is chosen when visually designing ETL pipelines, but it does not offer the serverless Spark job execution the stem requires.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.