Courseiva

PDE Ingesting and Processing the Data Practice Question

Your company is migrating an on-premises Hadoop cluster to Google Cloud. You need to transform large datasets using Spark SQL. Which Google Cloud service should you use?

⚠ Common exam trap

Google Cloud certification exams often test the distinction between managed Spark (Dataproc) and serverless SQL (BigQuery) or Beam-based processing (Dataflow), trapping candidates who see 'SQL' and immediately think of BigQuery without recognizing the Spark SQL execution context.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dataproc

Dataproc is the managed Spark and Hadoop service on Google Cloud, purpose-built for running existing Spark SQL workloads with minimal changes. It allows you to spin up a cluster, run your Spark SQL transformations on large datasets stored in Cloud Storage or BigQuery, and then tear it down, making it the direct equivalent of an on-premises Hadoop cluster in the cloud.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Dataflow

    Why it's wrong here

    Dataflow executes Apache Beam pipelines, not Spark SQL, so migrating Spark SQL code requires reimplementation rather than reuse. It is tempting because it is a managed, autoscaling processing service, and would be correct when building new streaming or batch pipelines in Beam rather than porting Spark jobs.

  • ✓

    Dataproc

    Why this is correct

    Dataproc is a managed Spark and Hadoop service, so existing Spark SQL jobs run with minimal refactoring. It satisfies the migration constraint by providing managed cluster provisioning and autoscaling while preserving the Spark SQL transformation logic used on-premises.

  • ✗

    BigQuery

    Why it's wrong here

    BigQuery runs standard SQL on its own engine, not Spark SQL, so existing Spark SQL jobs would need rewriting rather than lifting and shifting. It is tempting because it is a fully managed analytics warehouse, and would be correct when the transformations can be expressed in BigQuery SQL against data already loaded there.

  • ✗

    Cloud Dataprep

    Why it's wrong here

    Cloud Dataprep is a visually-driven data-wrangling service for preparing data, not a Spark SQL execution engine, so it cannot run your existing Spark SQL transformations. It is tempting because it does transform datasets, and would be correct for interactive, recipe-based cleaning of messy files before loading them.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.