Courseiva

PDE Designing Data Processing Systems Practice Question

You are migrating on-premises Hadoop jobs to Google Cloud. The existing jobs use Spark for ETL and Hive for querying. You want to minimize changes to the existing code and maintain the ability to use Hive queries with the same metastore across multiple clusters. Which service combination should you use?

⚠ Common exam trap

The trap is picking a modern serverless option (Dataflow or BigQuery) that requires code rewrites, when the requirement explicitly says 'minimize changes' and 'same metastore across multiple clusters' — only Dataproc plus Dataproc Metastore satisfies both.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Dataproc with Cloud Storage and Dataproc Metastore

Cloud Dataproc runs managed Spark and Hive clusters, so existing Spark ETL jobs and Hive queries migrate with minimal code changes. Pairing Dataproc with Cloud Storage for data and Dataproc Metastore for a shared Hive metastore across clusters preserves the same table definitions and schema across multiple clusters.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Dataflow with Beam SQL

    Why it's wrong here

    Beam SQL executes SQL over Dataflow pipelines but provides no Hive metastore, so existing Hive queries and shared table definitions cannot be reused. It is tempting because Dataflow runs Apache Beam, and Beam SQL would suit new pipelines needing SQL transforms without Hive compatibility across clusters.

  • ✗

    Cloud Dataproc with Dataproc on GKE

    Why it's wrong here

    Dataproc on GKE runs Spark and Hive workloads on Kubernetes clusters, yet it does not supply the shared, persistent Hive metastore that multiple clusters must reference. It is tempting because Dataproc on GKE preserves Spark and Hive code, and it would be correct when containerised orchestration is required rather than a shared metastore.

  • ✗

    Cloud BigQuery with external tables on Cloud Storage

    Why it's wrong here

    BigQuery external tables query Cloud Storage data but cannot execute HiveQL or connect to an existing Hive metastore, so the jobs' code would need rewriting. It is tempting because BigQuery offers serverless SQL over GCS, and external tables would be correct for new analytical queries without Hive dependencies.

  • ✓

    Cloud Dataproc with Cloud Storage and Dataproc Metastore

    Why this is correct

    Dataproc Metastore provides a managed Hive metastore service compatible with the Hive Metastore API, so existing Hive queries run unchanged and share metadata across multiple Dataproc clusters. Cloud Storage replaces HDFS as the storage layer, satisfying the requirement to minimise code changes during migration.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.