Courseiva
mediumMultiple Choice

PDE Practice Question: A data team uses Cloud Dataproc to run nightly…

A data team uses Cloud Dataproc to run nightly Spark jobs. The job volume has increased, and the cluster is often underutilized during the day. They want to reduce costs while ensuring jobs can scale when needed. Which strategy should they adopt?

⚠ Common exam trap

Google Cloud often tests the misconception that preemptible instances can be used for all nodes, but the trap here is that primary nodes require non-preemptible instances for cluster stability, while preemptible workers are only suitable for secondary (task) nodes in a fault-tolerant framework.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a cluster with a small number of primary workers and a large pool of preemptible workers, and enable autoscaling.

It combines a small number of primary (non-preemptible) workers for reliability with a large pool of preemptible workers for cost-effective scaling, and enables autoscaling to dynamically adjust the cluster size based on workload. This minimizes cost during idle periods (preemptible instances are ~80% cheaper) while ensuring jobs can scale up quickly when needed, as autoscaling adds preemptible workers automatically. Preemptible workers are ideal for fault-tolerant Spark jobs that can handle node preemptions.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use preemptible workers for both primary and secondary nodes to minimize cost.

    Why it's wrong here

    Preemptible workers cannot serve as primary nodes, which run the driver and HDFS daemons; preemption would kill the job. They are correct only as secondary workers alongside on-demand primaries. The stem needs autoscaling with preemptible secondaries, not preemptible primaries.

  • ✗

    Manually scale the cluster up before nightly jobs and down after.

    Why it's wrong here

    Manual scaling cannot respond to fluctuating job volume, so daytime idle capacity persists and nightly demand may outstrip the pre-scaled size. It is tempting because manual control suits predictable, scheduled workloads, but here autoscaling with a secondary worker pool or scheduled deletion would match capacity to demand.

  • ✓

    Use a cluster with a small number of primary workers and a large pool of preemptible workers, and enable autoscaling.

    Why this is correct

    Autoscaling adds and removes worker nodes based on YARN pending-resource demand, so the cluster shrinks during idle daytime hours and expands for nightly Spark jobs. Preemptible workers cut compute cost for fault-tolerant batch work, while a small primary-worker core preserves HDFS durability. This directly satisfies the cost-reduction and elastic-scaling constraints.

  • ✗

    Use custom machine types with local SSDs for primary workers to improve I/O.

    Why it's wrong here

    Custom machine types with local SSDs address I/O performance, not the idle-daytime cost or elastic scaling the stem requires. Ephemeral clusters or autoscaling with preemptible workers cut spend. Custom types suit steady workloads needing specific CPU-to-memory ratios or high disk throughput.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.