A data pipeline uses Cloud Composer (Airflow) to orchestrate Dataproc jobs. Each job submits a Spark application that reads from BigQuery and writes to Cloud Storage. The pipeline runs nightly and takes 6 hours. Management wants to reduce costs. Which approach is most effective?
Dataproc preemptible VMs cost substantially less than standard instances, and Spark's resilient distributed datasets tolerate their eviction by recomputing lost partitions. For a nightly six-hour batch job, that discount directly cuts spend without changing the pipeline's logic.
Why this answer
Preemptible VMs are significantly cheaper (up to 80% discount) than standard VMs and are ideal for fault-tolerant, batch workloads like nightly Dataproc jobs. Since the pipeline runs nightly and takes 6 hours, it can tolerate the occasional preemption of worker nodes by using Spark's built-in resilience (e.g., task retries). This directly reduces compute cost without sacrificing completion, assuming the cluster is configured with enough preemptible workers to handle the workload.
Exam trap
Google Cloud often tests the misconception that 'upgrading' storage class or changing billing granularity saves money, when in fact the correct answer involves leveraging cheaper compute resources (preemptible VMs) that are designed for fault-tolerant batch jobs.
How to eliminate wrong answers
Option B is wrong because Dataproc already bills per second after a 1-minute minimum, so switching to per-second billing is not a change that reduces costs further. Option C is wrong because increasing driver memory does not reduce costs; it may actually increase costs by requiring a larger, more expensive VM, and performance gains are unlikely if the bottleneck is not driver memory. Option D is wrong because upgrading from Standard to Nearline storage increases cost (Nearline has higher retrieval and minimum storage duration fees) and is intended for infrequently accessed data, not for nightly write workloads where data is read soon after writing.