Courseiva

PDE Designing Data Processing Systems Practice Question

A startup needs a fully managed, serverless Spark service to run occasional data processing jobs without managing clusters. They want to pay only for the resources used during job execution. Which Google Cloud service should they use?

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dataproc Serverless

Dataproc Serverless provides a serverless Spark environment where you pay per job execution. Cloud Data Fusion is for visual ETL. Dataproc is managed but not serverless. Dataflow is serverless for Beam, not Spark.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Dataproc Serverless

    Why this is correct

    Dataproc Serverless runs Spark workloads without provisioning or managing a cluster, allocating resources only while the job executes and charging for that consumption. This matches the stem's requirements for occasional jobs, no cluster management and pay-per-use billing, unlike a standard Dataproc cluster.

  • ✗

    Dataflow

    Why it's wrong here

    Dataflow is a managed Apache Beam runner for streaming and batch pipelines, not a serverless Spark service; it cannot execute Spark jobs. It suits event-driven transformation pipelines. The scenario's requirement for Spark with per-job, clusterless billing points to Dataproc Serverless instead.

  • ✗

    Cloud Data Fusion

    Why it's wrong here

    Cloud Data Fusion is a codeless ETL orchestration layer that runs pipelines on Dataproc or Dataflow; it does not itself execute Spark jobs serverlessly with per-job billing. It suits visual data-integration workflows, not direct Spark execution.

  • ✗

    Dataproc

    Why it's wrong here

    Dataproc provisions clusters of virtual machines, so the startup would manage nodes and pay for idle capacity between occasional jobs. It is tempting because Dataproc runs Spark and suits sustained workloads where cluster tuning is wanted, but the stem requires serverless execution with per-job billing, which Dataproc Serverless provides instead.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.