Courseiva

PDE Ingesting and Processing the Data Practice Question

Your team is migrating a batch ETL job from an on-premises Hadoop cluster to Dataproc. The job reads CSV files from Cloud Storage, joins them with a slowly changing dimension table in BigQuery, and writes aggregated results back to BigQuery. The on-premises job used Hive on Tez and took six hours. You need to reduce runtime on Dataproc while minimizing cost. Which approach should you take?

⚠ Common exam trap

The trap here is defaulting to a persistent cluster with preemptible workers for cost savings, when preemption can extend runtime and an always-on cluster bills for idle time.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Run the job as a Dataproc Serverless for Spark batch, reading the CSV files from Cloud Storage and using the BigQuery connector to read the dimension table, with autoscaling enabled.

Dataproc Serverless for Spark runs batch workloads on ephemeral, autoscaling infrastructure that is released when the batch finishes, so cost tracks actual execution rather than idle cluster time. Reading CSV from Cloud Storage and using the BigQuery connector for the dimension table lets Spark optimize the join, often broadcasting the small dimension. This reduces runtime compared with the legacy Tez engine and avoids paying for a cluster that sits idle between runs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Run the job as a Dataproc Serverless for Spark batch, reading the CSV files from Cloud Storage and using the BigQuery connector to read the dimension table, with autoscaling enabled.

    Why this is correct

    Dataproc Serverless for Spark provisions resources on demand, scales with the workload, and shuts down when the batch completes, so you pay only for the execution time. The BigQuery connector reads the dimension table efficiently, and Spark's optimizer can broadcast the dimension if it is small. This combination reduces runtime versus the legacy Tez job and minimizes cost by avoiding an idle cluster.

  • ✗

    Create a Dataproc cluster with a fixed number of standard workers, install Hive, and run the same Hive on Tez script to preserve the existing logic.

    Why it's wrong here

    Replicating the legacy Hive on Tez execution on Dataproc preserves the original six-hour runtime and adds cluster startup overhead. A fixed number of standard workers cannot scale with data volume, so cost stays high regardless of actual utilization. The goal is to reduce runtime and cost, and lifting the same engine and script unchanged does not achieve either.

  • ✗

    Create a long-running Dataproc cluster with many preemptible workers and run the job with Spark SQL, caching the BigQuery dimension table as a Spark DataFrame.

    Why it's wrong here

    A long-running cluster with many preemptible workers can be cost-effective for steady workloads, but preemptible workers can be reclaimed at any time, causing Spark stages to recompute and lengthen runtime unpredictably. For a six-hour batch job with a join, losing executors mid-stage can be costly. The scenario asks to minimize cost while reducing runtime, and an always-on cluster with volatile workers does not reliably achieve either.

  • ✗

    Export the CSV files to a temporary BigQuery table, perform the join and aggregation entirely in BigQuery SQL, and schedule the query with Cloud Scheduler.

    Why it's wrong here

    Moving the entire job into BigQuery SQL can be efficient, but exporting CSVs to a temporary table adds load time and storage cost, and Cloud Scheduler only triggers the query rather than managing compute lifecycle. The scenario specifically asks for a Dataproc migration, and this option sidesteps Dataproc entirely while adding an unnecessary export step that increases both runtime and cost.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.