Courseiva
easyMultiple Choice

PDE Practice Question: A team needs to migrate an existing on-premises…

A team needs to migrate an existing on-premises Hadoop Hive workload to Google Cloud. They want to minimize code changes and use a managed service for transient clusters. Which service should they choose?

⚠ Common exam trap

It's easy for candidates to confuse Cloud Dataflow's ability to process batch data with Hadoop compatibility, but Dataflow does not support Hive or transient Hadoop clusters, making Dataproc the only correct option for minimizing code changes.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Dataproc

Cloud Dataproc is the correct choice because it is a managed Spark and Hadoop service that supports Hive workloads natively, allowing you to run existing Hive scripts with minimal changes. It also supports transient clusters, which can be automatically scaled up and down, aligning with the requirement for transient clusters.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cloud Dataflow

    Why it's wrong here

    Cloud Dataflow runs Apache Beam pipelines, requiring the HiveQL logic to be rewritten as Beam transforms, and it provisions no Hadoop cluster. It is tempting because Dataflow is the managed service for batch and streaming pipelines, which suits new pipeline development rather than lift-and-shift Hive migration.

  • ✗

    Cloud Dataprep

    Why it's wrong here

    Cloud Dataprep is a visual, serverless data-wrangling tool for analysts preparing datasets; it does not execute HiveQL or run transient Hadoop clusters. It is tempting because it migrates data pipelines without code, which suits teams transforming CSV or BigQuery data rather than lifting Hive workloads.

  • ✓

    Cloud Dataproc

    Why this is correct

    Cloud Dataproc runs Apache Hive natively on managed, ephemeral clusters, so existing HiveQL scripts migrate with minimal rewriting. It satisfies both constraints: the managed service removes cluster administration, and transient clusters spin up only for job duration, cutting cost versus persistent on-premises infrastructure.

  • ✗

    BigQuery

    Why it's wrong here

    BigQuery is a serverless analytics warehouse queried with GoogleSQL, so existing HiveQL scripts need rewriting and no transient Hadoop cluster is provisioned. It is tempting because it is the usual managed destination for migrated Hive tables, which suits teams willing to refactor queries rather than minimise code changes.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.