You are migrating on-premises Hadoop jobs to Google Cloud. The existing jobs use Spark for ETL and Hive for querying. You want to minimize changes to the existing code and maintain the ability to use Hive queries with the same metastore across multiple clusters. Which service combination should you use?
Dataproc Metastore provides a managed Hive metastore service compatible with the Hive Metastore API, so existing Hive queries run unchanged and share metadata across multiple Dataproc clusters. Cloud Storage replaces HDFS as the storage layer, satisfying the requirement to minimise code changes during migration.
Why this answer
Cloud Dataproc runs managed Spark and Hive clusters, so existing Spark ETL jobs and Hive queries migrate with minimal code changes. Pairing Dataproc with Cloud Storage for data and Dataproc Metastore for a shared Hive metastore across clusters preserves the same table definitions and schema across multiple clusters.
Exam trap
The trap is picking a modern serverless option (Dataflow or BigQuery) that requires code rewrites, when the requirement explicitly says 'minimize changes' and 'same metastore across multiple clusters' — only Dataproc plus Dataproc Metastore satisfies both.
How to eliminate wrong answers
Option A is wrong because Dataflow with Beam SQL requires rewriting Spark/Hive logic into Beam pipelines and does not provide a Hive metastore. Option B is wrong because Dataproc on GKE changes the runtime model and does not by itself provide a shared Hive metastore across clusters. Option C is wrong because BigQuery with external tables replaces Hive semantics and requires rewriting queries, and it does not preserve a Hive metastore.