Courseiva
mediumMultiple Choice

PDE Practice Question: A data pipeline uses Cloud Composer (Airflow) to…

A data pipeline uses Cloud Composer (Airflow) to orchestrate Dataproc jobs. Each job submits a Spark application that reads from BigQuery and writes to Cloud Storage. The pipeline runs nightly and takes 6 hours. Management wants to reduce costs. Which approach is most effective?

⚠ Common exam trap

Google Cloud often tests the misconception that 'upgrading' storage class or changing billing granularity saves money, when in fact the correct answer involves leveraging cheaper compute resources (preemptible VMs) that are designed for fault-tolerant batch jobs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use preemptible VMs for the Dataproc cluster

Preemptible VMs are significantly cheaper (up to 80% discount) than standard VMs and are ideal for fault-tolerant, batch workloads like nightly Dataproc jobs. Since the pipeline runs nightly and takes 6 hours, it can tolerate the occasional preemption of worker nodes by using Spark's built-in resilience (e.g., task retries). This directly reduces compute cost without sacrificing completion, assuming the cluster is configured with enough preemptible workers to handle the workload.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use preemptible VMs for the Dataproc cluster

    Why this is correct

    Dataproc preemptible VMs cost substantially less than standard instances, and Spark's resilient distributed datasets tolerate their eviction by recomputing lost partitions. For a nightly six-hour batch job, that discount directly cuts spend without changing the pipeline's logic.

  • ✗

    Switch to Cloud Dataproc billing per second instead of per minute

    Why it's wrong here

    Per-second billing changes only the granularity of Dataproc charges, not the six hours of cluster runtime, so savings are marginal. It is tempting because it sounds like a cost optimisation, and would help short-lived or bursty clusters where sub-minute runtime is billed.

  • ✗

    Increase the memory of the driver node to improve performance

    Why it's wrong here

    Driver memory governs the Spark driver JVM, not shuffle or executor capacity, so it cannot shorten a six-hour job dominated by BigQuery reads and Cloud Storage writes. It is tempting because driver OOM errors are common, and increasing driver memory is the right fix when the driver crashes collecting results.

  • ✗

    Upgrade the Cloud Storage class from Standard to Nearline

    Why it's wrong here

    Nearline applies a 30-day minimum storage duration and per-object retrieval charges; nightly-written objects are read and rewritten constantly, so early deletion fees and access costs outweigh any per-GiB saving. Nearline is correct for data accessed less than once a month, such as monthly archives or backups.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.