PDE Designing Data Processing Systems Practice Question
A company uses Cloud Dataproc to run Spark ML training jobs. They want to persist the trained models and metadata in a Hive-compatible metastore. Which Dataproc feature should they use?
⚠ Common exam trap
PDE often tests the confusion between metadata storage services (Dataproc Metastore vs. Data Catalog) and database services (Bigtable), causing candidates to choose a non-Hive-compatible option for metastore requirements.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Dataproc Metastore
Dataproc Metastore is a fully managed, Hive-compatible metastore service on Google Cloud that integrates natively with Dataproc clusters, allowing Spark and Hive jobs to persist and share table metadata and schemas. It provides a serverless, scalable alternative to running a self-managed Hive metastore on a cluster, and it supports the Hive Metastore Thrift API so existing Spark ML workflows can store model metadata without code changes. This directly meets the requirement for a Hive-compatible metastore for trained models and metadata.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Cloud Hive Metastore (self-managed)
Why it's wrong here
A self-managed Hive metastore runs on a cluster that is deleted with the job, so metadata is lost unless separately persisted; Dataproc Metastore is the managed service that survives cluster deletion. It is tempting because it is Hive-compatible, making it correct when you must control the metastore host yourself.
- ✗
Cloud Bigtable
Why it's wrong here
Cloud Bigtable is a wide-column NoSQL store for low-latency key lookups, not a Hive-compatible metastore, so Spark cannot use it to persist table definitions and partitions. It is tempting because it is a scalable datastore, making it correct when the workload needs high-throughput time-series or key-value access.
- ✓
Dataproc Metastore
Why this is correct
Dataproc Metastore provides a fully managed, Hive-compatible metastore service that persists table metadata and model artefacts independently of cluster lifecycle. This satisfies the requirement to retain trained models and metadata in a Hive-compatible store across ephemeral Dataproc clusters.
- ✗
Cloud Data Catalog
Why it's wrong here
Cloud Data Catalog is a metadata inventory for discovery and governance; it does not provide a Hive-compatible metastore that Spark can read and write tables through. It is tempting because it stores metadata, making it correct when the requirement is cataloguing and searching assets rather than persisting Hive tables.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.