PDE Designing Data Processing Systems Practice Question
An organization runs periodic Apache Spark jobs on Dataproc to process data from Cloud Storage. They want to reduce costs by using preemptible instances for worker nodes. What is a key consideration when using preemptible instances in Dataproc?
⚠ Common exam trap
PDE often tests the misconception that preemptible instances are transparent to the job — candidates must recognize that preemption causes task re-execution and longer runtimes, and that checkpointing is a developer responsibility, not an automatic Dataproc feature.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Jobs must be designed to handle node preemption, and overall job runtime may increase
Preemptible VMs in Dataproc can be reclaimed by Compute Engine at any time (with a 30-second shutdown notice), so Spark jobs must be designed to tolerate worker loss — for example, by using checkpointing, retries, or resilient data pipelines. Because preempted nodes are removed and replacements take time to spin up, overall job runtime typically increases compared to a cluster of standard VMs. This trade-off is the key operational consideration when choosing preemptible workers for cost savings.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Preemptible instances cannot be used with the standard cluster mode
Why it's wrong here
Preemptible instances are supported in standard Dataproc clusters as secondary workers; the limitation concerns their removal, not cluster mode. The real consideration is that preemption terminates workers mid-job, so workloads need checkpointing or tolerance for lost nodes. Standard mode is not the constraint.
- ✓
Jobs must be designed to handle node preemption, and overall job runtime may increase
Why this is correct
Preemptible nodes can be reclaimed at any time, so Spark jobs must tolerate losing workers mid-execution, and recomputation plus retries typically lengthen total runtime. This matches the stem's cost-reduction goal while acknowledging the resilience and duration trade-offs preemptible instances impose.
- ✗
Preemptible instances are only available in certain regions
Why it's wrong here
Preemptible instances are available across Dataproc regions generally, so regional availability is not the defining consideration. The genuine issue is that Compute Engine reclaims them after 24 hours or on capacity demand, so Spark jobs must tolerate worker loss and shuffle recomputation.
- ✗
Jobs will automatically restart from the last checkpoint without any performance impact
Why it's wrong here
Dataproc does not automatically restart jobs from checkpoints after preemption; Spark must be configured with checkpointing, and recovery still incurs recomputation and performance cost. Preemptible workers are reclaimed at any time, so the actual consideration is designing jobs to tolerate mid-execution node loss.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.