Courseiva

PDE Maintaining and Automating Data Workloads Practice Question

You are monitoring a Dataproc cluster and notice that the cluster utilisation is high, but jobs are running slowly. The cluster uses preemptible workers for cost savings. What is the most likely cause of the performance degradation?

⚠ Common exam trap

PDE often tests the misconception that high utilization means the cluster is healthy, when in fact preemptible worker churn causes retries that inflate utilization while slowing jobs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The preemptible workers are being preempted frequently, causing task retries and slowdowns.

Preemptible (spot) workers in Dataproc can be reclaimed by Compute Engine at any time, and when they are preempted mid-task, the task must be retried on remaining workers. Frequent preemption causes repeated task retries, stragglers, and overall job slowdown even though the cluster appears highly utilized. This is the classic symptom of relying heavily on preemptible workers.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The primary workers are using standard disks instead of SSDs.

    Why it's wrong here

    Disk type on primary workers does not explain slowdowns tied to preemptible workers being reclaimed mid-task. Switching to SSDs is tempting when diagnosing I/O-bound stages such as shuffle-heavy Spark jobs, and would be correct if metrics showed disk saturation rather than repeated worker loss.

  • ✓

    The preemptible workers are being preempted frequently, causing task retries and slowdowns.

    Why this is correct

    Frequent preemption of preemptible VMs forces Dataproc to rerun interrupted tasks on remaining workers, inflating job duration despite high CPU utilisation. Since the stem specifies preemptible workers used for cost savings, this directly explains the slowdown: capacity vanishes mid-job, triggering retries and stragglers rather than a genuine resource shortage.

  • ✗

    The cluster is under-provisioned; increase the number of preemptible workers.

    Why it's wrong here

    Adding preemptible workers does not stop preemption; those workers are reclaimed when capacity is needed, so the slowdown persists. Scaling out is tempting when utilisation is genuinely capacity-bound, and would be correct if the cluster ran on standard, non-preemptible workers that stayed available.

  • ✗

    The cluster is using an older image version; upgrade to the latest.

    Why it's wrong here

    An outdated image version does not cause preemptible workers to be reclaimed mid-job; preemption, not image age, explains the slowdown. Upgrading images is tempting when diagnosing compatibility or dependency faults, and would be right if logs showed version-related errors rather than repeated worker loss.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.