Courseiva
mediumMultiple SelectObjective-mapped

PDE Practice Question: Is moving on-premises Hadoop workloads to Google…

An organization is moving on-premises Hadoop workloads to Google Cloud. They need to minimize code changes and manage transient clusters for cost savings. Which two Google Cloud services should they consider? (Choose TWO.)

⚠ Common exam trap

Many candidates confuse Cloud Dataflow (a Google Cloud service that runs Beam pipelines) with Dataproc, not realizing that Dataflow requires rewriting Hadoop jobs into Beam pipelines, while Dataproc on GKE and Cloud Dataproc directly support unmodified Hadoop/Spark code.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

Dataproc on GKE

Cloud Dataproc (option D) is a managed service for running Spark and Hadoop clusters. It supports transient clusters that can be created on-demand and deleted when idle, minimizing costs. It also allows direct migration of on-premises Hadoop code with minimal changes because it supports standard Hadoop/Spark APIs. Dataproc on GKE (option C) provides similar benefits but runs containerized workloads on GKE, offering additional ephemeral cluster capabilities and integration with Kubernetes. Both options minimize code changes and enable transient clusters for cost savings, while BigQuery (option B) requires rewriting SQL queries and Cloud Dataflow (option E) requires converting to Beam pipelines. Compute Engine with self-managed Hadoop (option A) does not provide transient cluster management by default.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • Compute Engine with self-managed Hadoop

    Why it's wrong here

    Requires manual setup and does not minimize changes compared to Dataproc.

  • BigQuery

    Why it's wrong here

    BigQuery is not Hadoop-compatible; requires code rewrite.

  • Dataproc on GKE

    Why this is correct

    Allows running Spark workloads on GKE, leveraging container orchestration.

  • Cloud Dataproc

    Why this is correct

    Dataproc is a managed Hadoop/Spark service supporting transient clusters.

  • Cloud Dataflow

    Why it's wrong here

    Dataflow is not Hadoop-compatible; would require code rewrite.

About these practice questions

Courseiva writes every PDE question from scratch — 890 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.