Courseiva
hardMultiple Select

PDE Practice Question: Migrating an on-premises Hadoop cluster to Google…

A company is migrating an on-premises Hadoop cluster to Google Cloud. They need to run existing Spark jobs with minimal modification. Which THREE strategies should they consider? (Choose THREE.)

⚠ Common exam trap

Candidates often assume BigQuery or Dataflow are the only Google Cloud data processing options, overlooking that Dataproc is specifically designed for minimal-change migrations of existing Spark/Hadoop workloads.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use Cloud Dataproc with Spark and Hive components.

Option B is correct because Cloud Dataproc is a managed Hadoop/Spark service that natively runs Apache Spark and Hive, so existing Spark jobs can be migrated with minimal code changes. Option C is correct because Dataproc jobs commonly read and write data in Cloud Storage (gs://) instead of HDFS, which is the standard cloud-native replacement for on-premises HDFS storage and requires only path changes. Option E is correct because the Dataproc Jobs API lets you submit Spark jobs programmatically to a cluster, preserving the existing job-submission workflow with minimal modification. Option A is not appropriate because migrating everything to BigQuery would require rewriting analytics logic and does not run existing Spark jobs as-is. Option D is not appropriate because rewriting Spark jobs as Dataflow pipelines is a significant code and framework change, contradicting the minimal-modification requirement.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Migrate to BigQuery for all analytics.

    Why it's wrong here

    BigQuery is a serverless analytics warehouse with its own SQL dialect, so existing Spark code cannot run unmodified; jobs would need rewriting. It suits new cloud-native analytics workloads, not lift-and-shift Spark execution. Dataproc, Dataproc Serverless or Dataproc on GKE preserve Spark APIs.

  • ✓

    Use Cloud Dataproc with Spark and Hive components.

    Why this is correct

    Cloud Dataproc provides managed Spark and Hive, so existing jobs run with minimal modification — satisfying the stem's constraint. Unlike re-platforming to BigQuery or Dataflow, which require rewriting jobs against different APIs, Dataproc preserves the Hadoop ecosystem's execution model and tooling on Google Cloud.

  • ✓

    Store data in Cloud Storage instead of HDFS.

    Why this is correct

    Cloud Storage decouples storage from compute, letting Dataproc clusters read and write data without HDFS reconfiguration. Spark's Hadoop-compatible connector (`gs://`) means existing jobs need only path changes, satisfying the minimal-modification constraint. Ephemeral clusters can then be deleted without data loss, unlike HDFS on persistent nodes.

  • ✗

    Rewrite Spark jobs as Dataflow pipelines.

    Why it's wrong here

    Dataflow runs Apache Beam, not Spark, so every job must be rewritten against the Beam API, violating the minimal-modification requirement. Dataflow is the right choice for new pipelines needing unified batch and streaming with autoscaling. Dataproc lets existing Spark code run largely unchanged.

  • ✓

    Use Dataproc Jobs API to submit jobs.

    Why this is correct

    The Dataproc Jobs API submits Spark jobs directly to a Dataproc cluster, accepting existing JARs and Python files without rewriting them. This satisfies the minimal-modification constraint, since Hadoop and Spark workloads run largely unchanged on Dataproc's managed infrastructure.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.