Courseiva

PDE Designing Data Processing Systems Practice Question

A company uses Cloud Dataproc to run Spark ML training jobs. They want to persist the trained models and metadata in a Hive-compatible metastore. Which Dataproc feature should they use?

⚠ Common exam trap

PDE often tests the confusion between metadata storage services (Dataproc Metastore vs. Data Catalog) and database services (Bigtable), causing candidates to choose a non-Hive-compatible option for metastore requirements.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Dataproc Metastore

Dataproc Metastore is a fully managed, Hive-compatible metastore service on Google Cloud that integrates natively with Dataproc clusters, allowing Spark and Hive jobs to persist and share table metadata and schemas. It provides a serverless, scalable alternative to running a self-managed Hive metastore on a cluster, and it supports the Hive Metastore Thrift API so existing Spark ML workflows can store model metadata without code changes. This directly meets the requirement for a Hive-compatible metastore for trained models and metadata.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Hive Metastore (self-managed)

    Why it's wrong here

    A self-managed Hive metastore runs on a cluster that is deleted with the job, so metadata is lost unless separately persisted; Dataproc Metastore is the managed service that survives cluster deletion. It is tempting because it is Hive-compatible, making it correct when you must control the metastore host yourself.

  • ✗

    Cloud Bigtable

    Why it's wrong here

    Cloud Bigtable is a wide-column NoSQL store for low-latency key lookups, not a Hive-compatible metastore, so Spark cannot use it to persist table definitions and partitions. It is tempting because it is a scalable datastore, making it correct when the workload needs high-throughput time-series or key-value access.

  • ✓

    Dataproc Metastore

    Why this is correct

    Dataproc Metastore provides a fully managed, Hive-compatible metastore service that persists table metadata and model artefacts independently of cluster lifecycle. This satisfies the requirement to retain trained models and metadata in a Hive-compatible store across ephemeral Dataproc clusters.

  • ✗

    Cloud Data Catalog

    Why it's wrong here

    Cloud Data Catalog is a metadata inventory for discovery and governance; it does not provide a Hive-compatible metastore that Spark can read and write tables through. It is tempting because it stores metadata, making it correct when the requirement is cataloguing and searching assets rather than persisting Hive tables.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.