PDE Designing Data Processing Systems Practice Question
A company uses Dataproc to run daily Spark ML jobs. The jobs run for 2 hours each day. The team wants to reduce costs without changing job characteristics. Which strategy is MOST cost-effective?
⚠ Common exam trap
The trap is assuming 'reduce cost' means 'reduce cluster size' — candidates pick single-node or fewer workers, but the exam expects you to recognize that preemptible/Spot instances reduce cost per unit while preserving job characteristics.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use preemptible instances for worker nodes
Preemptible (Spot) VMs cost up to 80% less than standard VMs and are ideal for fault-tolerant, batch-oriented workloads like Spark ML jobs that can tolerate occasional preemption. Since the jobs run only 2 hours daily and Dataproc automatically handles node replacement when preemptible instances are reclaimed, using preemptible workers delivers the largest cost reduction without changing job characteristics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a single-node cluster to eliminate overhead
Why it's wrong here
A single-node cluster runs master and worker roles on one VM, capping the cores and memory available to Spark ML and preventing the distributed parallelism the jobs rely on. It is tempting because it removes multi-VM overhead, and it would be correct for small, non-distributed workloads.
- ✗
Enable high-availability mode to avoid restarts
Why it's wrong here
High-availability mode runs extra master nodes and a second worker replica, raising the hourly rate for a two-hour batch job that does not need fault tolerance. It is tempting because it prevents restart delays, and it would be correct for long-running streaming clusters where a master failure would lose in-flight work.
- ✓
Use preemptible instances for worker nodes
Why this is correct
Preemptible instances cost significantly less than standard Dataproc worker nodes, and Spark can tolerate their loss through retries and task rescheduling. Since the daily jobs are short and their characteristics stay unchanged, using preemptible workers cuts compute spend without altering the job design.
- ✗
Increase the number of standard workers to finish faster
Why it's wrong here
Adding standard workers raises the per-second cost of the cluster while the two-hour runtime is set by the job's data volume and algorithm, so spend increases without a proportional time saving. It is tempting because faster completion reduces billed minutes, and it would be correct for a deadline-bound job that scales linearly.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.