hardMultiple Choice
PDE Practice Question: A company processes large volumes of GPS sensor…
A company processes large volumes of GPS sensor data stored in Cloud Storage. Each hour, they run an Apache Spark job that aggregates the data by geohash region. The job must be cost-effective and scale automatically. Currently, they are using a Dataproc cluster with preemptible workers. Which improvement would best reduce costs while maintaining performance?
⚠ Common exam trap
Google Cloud often tests the misconception that migrating to a different processing engine (like Dataflow or BigQuery) is always the best cost-saving move, when in fact reusing existing Spark code on a serverless platform avoids migration costs and leverages the same API.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Dataproc Serverless Spark
Dataproc Serverless Spark (Option D) eliminates the need to manage a cluster, automatically scaling resources to match job demand and charging only for the resources consumed during execution. This removes the overhead of preemptible worker management and idle cluster costs, directly reducing expenses while maintaining performance for the hourly aggregation job.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a larger Dataproc cluster with standard workers
Why it's wrong here
Standard workers cost more per hour than preemptible VMs and a larger fixed cluster removes the automatic scaling the job requires. Standard workers suit workloads that cannot tolerate preemption, such as long-running stateful jobs, not hourly batch aggregation.
- ✗
Migrate the job to BigQuery scheduled queries
Why it's wrong here
BigQuery scheduled queries run SQL, not Apache Spark, so the existing geohash aggregation logic would need rewriting and Spark-specific transforms may be unsupported. BigQuery is the right target when the workload is expressible in SQL and data already resides there.
- ✗
Switch to Dataflow batch pipeline with Apache Beam
Why it's wrong here
Dataflow with Apache Beam runs a managed, autoscaling service, but migrating the existing Spark aggregation logic off Dataproc requires rewriting it, and the stem asks to improve the current cluster's cost. Dataflow suits new pipelines or portable Beam code, not reducing spend on an established Spark job.
- ✓
Use Dataproc Serverless Spark
Why this is correct
Dataproc Serverless Spark bills per-second for actual workload consumption and provisions capacity automatically, eliminating idle cluster time and manual sizing. For hourly aggregation jobs, this removes the cost of preemptible workers sitting idle between runs while still scaling to demand.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.