Courseiva

Cloud Digital Leader Google Cloud Products and Services Practice Question

A company needs to run a Hadoop/Spark workload on Google Cloud. They must use existing YARN applications and need to optimise for cost by using preemptible VMs for task nodes. Which three services should they use?

⚠ Common exam trap

GCDL often tests the trap of selecting BigQuery or Dataflow for Hadoop/Spark workloads, when only Dataproc (with Compute Engine and Cloud Storage) supports YARN-based Spark/Hadoop jobs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Compute Engine

Cloud Dataproc (B) is the right managed service because it natively runs Hadoop/Spark clusters on Google Cloud and supports existing YARN applications, including the ability to designate preemptible VMs specifically as secondary worker (task) nodes to reduce cost. Compute Engine (A) is correct because Dataproc clusters are provisioned on Compute Engine VM instances, so the underlying compute for master, primary, and preemptible secondary workers is Compute Engine. Cloud Storage (C) is correct because Dataproc uses Cloud Storage (gs://) as its default Hadoop-compatible file system (via the Cloud Storage connector), letting the workload store input/output data durably and cheaply instead of HDFS on persistent disks. BigQuery (D) is not appropriate here because it is a serverless analytics data warehouse, not a platform for running YARN/Spark applications. Dataflow (E) is also not appropriate because it is a managed Apache Beam runner for data pipelines, not a Hadoop/Spark YARN cluster environment.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Compute Engine

    Why this is correct

    Compute Engine provides the virtual machines that form the worker and master nodes of a Cloud Dataproc cluster. When you run a Hadoop/Spark workload on Google Cloud, Cloud Dataproc orchestrates the deployment, but the actual CPU, memory, and local storage attached to each cluster node are Compute Engine instances. You can also run Hadoop/Spark directly on your own Compute Engine VMs without Dataproc, making Compute Engine the fundamental compute infrastructure for such workloads.

  • ✓

    Cloud Dataproc

    Why this is correct

    Cloud Dataproc is the fully managed service that simplifies running open-source Apache Hadoop and Spark clusters on Google Cloud. It creates and manages cluster nodes as Compute Engine VMs, auto-scales them, and integrates with Cloud Storage for persistent data. Instead of manually installing Hadoop or Spark on raw VMs, Dataproc handles the cluster lifecycle, software configuration, and monitoring, making it the recommended way to execute Hadoop/Spark workloads.

  • ✓

    Cloud Storage

    Why this is correct

    Cloud Storage replaces HDFS as the default data layer in a Cloud Dataproc cluster, allowing you to decouple compute from storage. Since HDFS is ephemeral and tied to cluster lifecycle, storing input and output data in Cloud Storage enables you to delete clusters after jobs finish without losing data, and to share data across clusters. Dataproc reads and writes data directly from and to Cloud Storage buckets via the gs:// connector, avoiding expensive HDFS replication overhead and enabling elastic cluster scaling.

  • ✗

    BigQuery

    Why it's wrong here

    BigQuery is a serverless, highly scalable enterprise data warehouse that runs analytics using SQL, not Apache Hadoop or Spark. It cannot run Hadoop/Spark jobs or host YARN/executors; it is a separate service for storing and querying structured data. While you can use BigQuery to analyze data that originated from Hadoop workloads, it does not provide the same execution environment and is not an alternative for running Spark code.

  • ✗

    Dataflow

    Why it's wrong here

    Dataflow is a unified stream and batch data processing service based on the Apache Beam programming model, not on Hadoop or Spark. It is intended for building data pipelines that transform and enrich data, but it does not run Java/Scala Spark jobs or provide a Hadoop-compatible distributed filesystem. Choosing Dataflow would require rewriting your Hadoop/Spark code into Beam transforms, so it is not a drop-in replacement for a Hadoop/Spark workload.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

Go deeper

Related to this question

About these practice questions

This GCDL question is part of Courseiva's 848-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This GCDL practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the GCDL exam.