Courseiva

PDE Designing Data Processing Systems Practice Question

A company runs Apache Spark jobs on Dataproc. They want to reduce costs by using preemptible instances for worker nodes. The jobs are fault-tolerant and can handle occasional node loss. However, the cluster must remain available for interactive querying during business hours. Which Dataproc cluster configuration meets these requirements?

⚠ Common exam trap

The trap here is conflating 'fault-tolerant job' with 'fault-tolerant cluster' — candidates assume preemptible nodes can be used anywhere because the job can retry, forgetting that master and primary worker roles must remain non-preemptible to preserve cluster availability.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a standard cluster with preemptible instances as secondary workers

Dataproc supports preemptible VMs as secondary workers, which are added to a cluster alongside non-preemptible primary workers. Secondary workers handle the bulk of shuffle and task execution, so if a preemptible node is reclaimed, the job can recover while the primary workers and master keep the cluster alive and available for interactive queries. This gives the cost benefit of preemptible pricing without risking cluster or master availability.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a single-node cluster that automatically scales with preemptible instances

    Why it's wrong here

    A single-node cluster runs master and worker roles on one VM, so preemption or scaling removes the only node and interactive querying stops. It is tempting because single-node clusters are cheap for development, and it would be correct for small non-critical experiments rather than business-hours availability.

  • ✓

    Use a standard cluster with preemptible instances as secondary workers

    Why this is correct

    Secondary workers (preemptible workers) are ideal for fault-tolerant batch jobs. They do not store HDFS data, so losing them does not affect data durability. The cluster remains available because primary workers and master nodes are regular instances.

  • ✗

    Use standard cluster with master and worker nodes as preemptible instances

    Why it's wrong here

    Making master nodes preemptible risks losing the master, which terminates the cluster and ends interactive querying. It is tempting because applying preemptible pricing across every node maximises savings, and it would suit short-lived batch clusters where a restart is acceptable.

  • ✗

    Use a high-availability cluster with preemptible instances for primary workers

    Why it's wrong here

    Even with HA, losing primary workers can cause data loss if HDFS replication is insufficient. Secondary workers are better for preemptible use.

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.