PDE Ingesting and Processing the Data Practice Question
Your company is migrating an on-premises Hadoop cluster to Google Cloud. You need to transform large datasets using Spark SQL. Which Google Cloud service should you use?
⚠ Common exam trap
Google Cloud certification exams often test the distinction between managed Spark (Dataproc) and serverless SQL (BigQuery) or Beam-based processing (Dataflow), trapping candidates who see 'SQL' and immediately think of BigQuery without recognizing the Spark SQL execution context.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dataproc
Dataproc is the managed Spark and Hadoop service on Google Cloud, purpose-built for running existing Spark SQL workloads with minimal changes. It allows you to spin up a cluster, run your Spark SQL transformations on large datasets stored in Cloud Storage or BigQuery, and then tear it down, making it the direct equivalent of an on-premises Hadoop cluster in the cloud.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Dataflow
Why it's wrong here
Dataflow executes Apache Beam pipelines, not Spark SQL, so migrating Spark SQL code requires reimplementation rather than reuse. It is tempting because it is a managed, autoscaling processing service, and would be correct when building new streaming or batch pipelines in Beam rather than porting Spark jobs.
- ✓
Dataproc
Why this is correct
Dataproc is a managed Spark and Hadoop service, so existing Spark SQL jobs run with minimal refactoring. It satisfies the migration constraint by providing managed cluster provisioning and autoscaling while preserving the Spark SQL transformation logic used on-premises.
- ✗
BigQuery
Why it's wrong here
BigQuery runs standard SQL on its own engine, not Spark SQL, so existing Spark SQL jobs would need rewriting rather than lifting and shifting. It is tempting because it is a fully managed analytics warehouse, and would be correct when the transformations can be expressed in BigQuery SQL against data already loaded there.
- ✗
Cloud Dataprep
Why it's wrong here
Cloud Dataprep is a visually-driven data-wrangling service for preparing data, not a Spark SQL execution engine, so it cannot run your existing Spark SQL transformations. It is tempting because it does transform datasets, and would be correct for interactive, recipe-based cleaning of messy files before loading them.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.