Your team is migrating a batch ETL job from an on-premises Hadoop cluster to Dataproc. The job reads CSV files from Cloud Storage, joins them with a slowly changing dimension table in BigQuery, and writes aggregated results back to BigQuery. The on-premises job used Hive on Tez and took six hours. You need to reduce runtime on Dataproc while minimizing cost. Which approach should you take?
Dataproc Serverless for Spark provisions resources on demand, scales with the workload, and shuts down when the batch completes, so you pay only for the execution time. The BigQuery connector reads the dimension table efficiently, and Spark's optimizer can broadcast the dimension if it is small. This combination reduces runtime versus the legacy Tez job and minimizes cost by avoiding an idle cluster.
Why this answer
Dataproc Serverless for Spark runs batch workloads on ephemeral, autoscaling infrastructure that is released when the batch finishes, so cost tracks actual execution rather than idle cluster time. Reading CSV from Cloud Storage and using the BigQuery connector for the dimension table lets Spark optimize the join, often broadcasting the small dimension. This reduces runtime compared with the legacy Tez engine and avoids paying for a cluster that sits idle between runs.
Exam trap
The trap here is defaulting to a persistent cluster with preemptible workers for cost savings, when preemption can extend runtime and an always-on cluster bills for idle time.