hardMultiple Select
PDE Practice Question: Migrating an on-premises Hadoop cluster to Google…
A company is migrating an on-premises Hadoop cluster to Google Cloud. They need to run existing Spark jobs with minimal modification. Which THREE strategies should they consider? (Choose THREE.)
⚠ Common exam trap
Candidates often assume BigQuery or Dataflow are the only Google Cloud data processing options, overlooking that Dataproc is specifically designed for minimal-change migrations of existing Spark/Hadoop workloads.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use Cloud Dataproc with Spark and Hive components.
Option B is correct because Cloud Dataproc is a managed Hadoop/Spark service that natively runs Apache Spark and Hive, so existing Spark jobs can be migrated with minimal code changes. Option C is correct because Dataproc jobs commonly read and write data in Cloud Storage (gs://) instead of HDFS, which is the standard cloud-native replacement for on-premises HDFS storage and requires only path changes. Option E is correct because the Dataproc Jobs API lets you submit Spark jobs programmatically to a cluster, preserving the existing job-submission workflow with minimal modification. Option A is not appropriate because migrating everything to BigQuery would require rewriting analytics logic and does not run existing Spark jobs as-is. Option D is not appropriate because rewriting Spark jobs as Dataflow pipelines is a significant code and framework change, contradicting the minimal-modification requirement.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Migrate to BigQuery for all analytics.
Why it's wrong here
BigQuery is a serverless analytics warehouse with its own SQL dialect, so existing Spark code cannot run unmodified; jobs would need rewriting. It suits new cloud-native analytics workloads, not lift-and-shift Spark execution. Dataproc, Dataproc Serverless or Dataproc on GKE preserve Spark APIs.
- ✓
Use Cloud Dataproc with Spark and Hive components.
Why this is correct
Cloud Dataproc provides managed Spark and Hive, so existing jobs run with minimal modification — satisfying the stem's constraint. Unlike re-platforming to BigQuery or Dataflow, which require rewriting jobs against different APIs, Dataproc preserves the Hadoop ecosystem's execution model and tooling on Google Cloud.
- ✓
Store data in Cloud Storage instead of HDFS.
Why this is correct
Cloud Storage decouples storage from compute, letting Dataproc clusters read and write data without HDFS reconfiguration. Spark's Hadoop-compatible connector (`gs://`) means existing jobs need only path changes, satisfying the minimal-modification constraint. Ephemeral clusters can then be deleted without data loss, unlike HDFS on persistent nodes.
- ✗
Rewrite Spark jobs as Dataflow pipelines.
Why it's wrong here
Dataflow runs Apache Beam, not Spark, so every job must be rewritten against the Beam API, violating the minimal-modification requirement. Dataflow is the right choice for new pipelines needing unified batch and streaming with autoscaling. Dataproc lets existing Spark code run largely unchanged.
- ✓
Use Dataproc Jobs API to submit jobs.
Why this is correct
The Dataproc Jobs API submits Spark jobs directly to a Dataproc cluster, accepting existing JARs and Python files without rewriting them. This satisfies the minimal-modification constraint, since Hadoop and Spark workloads run largely unchanged on Dataproc's managed infrastructure.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.