PDE Designing Data Processing Systems Practice Question
A Dataproc cluster uses preemptible worker nodes to reduce costs. The cluster runs a long-running Spark job that occasionally experiences worker failures. How should the job be configured to handle preemptible worker failures gracefully?
⚠ Common exam trap
PDE often tests whether candidates confuse driver-level resilience (automatic restart) with executor/task-level resilience (maxFailures), or mistakenly believe persistent disks prevent preemption.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Set spark.task.maxFailures to a high number to allow retries.
Preemptible VMs on Dataproc can be reclaimed at any time, causing executor loss mid-task. Setting spark.task.maxFailures to a higher value (default is 4) allows Spark to retry failed tasks on remaining executors instead of failing the entire job. This is the standard, low-cost way to tolerate preemptible worker churn without abandoning the cost savings.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Set spark.task.maxFailures to a high number to allow retries.
Why this is correct
Raising `spark.task.maxFailures` lets individual tasks retry after a preemptible worker is reclaimed, so the job survives transient node loss without failing the stage. However, this only addresses task-level retries; Dataproc's enhanced flexibility mode, which pairs a standard primary pool with preemptible secondary workers, is what actually enables graceful recovery.
- ✗
Disable preemptible workers for the job.
Why it's wrong here
Disabling preemptible workers removes the cost saving the cluster was built around, abandoning the scenario's premise instead of tolerating preemption. It is tempting because it eliminates failures outright, and it would be correct for jobs with strict completion deadlines where spot capacity cannot be risked.
- ✗
Use persistent disks for preemptible workers.
Why it's wrong here
Persistent disks preserve data across preemption, but the Spark executors themselves still terminate, so in-flight tasks are lost and must be recomputed. It is tempting because persistence sounds like resilience, and it would be correct for stateful workloads needing data to survive instance replacement.
- ✗
Enable automatic restart of the Spark driver on failure.
Why it's wrong here
Restarting the driver does not help when executor nodes are preempted; the driver is typically on a standard node and survives. It is tempting because driver restart is a known resilience pattern, and it would be correct when the driver itself fails on a non-preemptible master.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.