PDE Designing Data Processing Systems Practice Question
You are migrating on-premises Hadoop jobs to Google Cloud. The existing jobs use Spark for ETL and Hive for querying. You want to minimize changes to the existing code and maintain the ability to use Hive queries with the same metastore across multiple clusters. Which service combination should you use?
⚠ Common exam trap
The trap is picking a modern serverless option (Dataflow or BigQuery) that requires code rewrites, when the requirement explicitly says 'minimize changes' and 'same metastore across multiple clusters' — only Dataproc plus Dataproc Metastore satisfies both.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Dataproc with Cloud Storage and Dataproc Metastore
Cloud Dataproc runs managed Spark and Hive clusters, so existing Spark ETL jobs and Hive queries migrate with minimal code changes. Pairing Dataproc with Cloud Storage for data and Dataproc Metastore for a shared Hive metastore across clusters preserves the same table definitions and schema across multiple clusters.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Dataflow with Beam SQL
Why it's wrong here
Beam SQL executes SQL over Dataflow pipelines but provides no Hive metastore, so existing Hive queries and shared table definitions cannot be reused. It is tempting because Dataflow runs Apache Beam, and Beam SQL would suit new pipelines needing SQL transforms without Hive compatibility across clusters.
- ✗
Cloud Dataproc with Dataproc on GKE
Why it's wrong here
Dataproc on GKE runs Spark and Hive workloads on Kubernetes clusters, yet it does not supply the shared, persistent Hive metastore that multiple clusters must reference. It is tempting because Dataproc on GKE preserves Spark and Hive code, and it would be correct when containerised orchestration is required rather than a shared metastore.
- ✗
Cloud BigQuery with external tables on Cloud Storage
Why it's wrong here
BigQuery external tables query Cloud Storage data but cannot execute HiveQL or connect to an existing Hive metastore, so the jobs' code would need rewriting. It is tempting because BigQuery offers serverless SQL over GCS, and external tables would be correct for new analytical queries without Hive dependencies.
- ✓
Cloud Dataproc with Cloud Storage and Dataproc Metastore
Why this is correct
Dataproc Metastore provides a managed Hive metastore service compatible with the Hive Metastore API, so existing Hive queries run unchanged and share metadata across multiple Dataproc clusters. Cloud Storage replaces HDFS as the storage layer, satisfying the requirement to minimise code changes during migration.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.