Cloud Digital Leader Google Cloud Products and Services Practice Question
A company needs to run a Hadoop/Spark workload on Google Cloud. They must use existing YARN applications and need to optimise for cost by using preemptible VMs for task nodes. Which three services should they use?
⚠ Common exam trap
GCDL often tests the trap of selecting BigQuery or Dataflow for Hadoop/Spark workloads, when only Dataproc (with Compute Engine and Cloud Storage) supports YARN-based Spark/Hadoop jobs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Compute Engine
Cloud Dataproc (B) is the right managed service because it natively runs Hadoop/Spark clusters on Google Cloud and supports existing YARN applications, including the ability to designate preemptible VMs specifically as secondary worker (task) nodes to reduce cost. Compute Engine (A) is correct because Dataproc clusters are provisioned on Compute Engine VM instances, so the underlying compute for master, primary, and preemptible secondary workers is Compute Engine. Cloud Storage (C) is correct because Dataproc uses Cloud Storage (gs://) as its default Hadoop-compatible file system (via the Cloud Storage connector), letting the workload store input/output data durably and cheaply instead of HDFS on persistent disks. BigQuery (D) is not appropriate here because it is a serverless analytics data warehouse, not a platform for running YARN/Spark applications. Dataflow (E) is also not appropriate because it is a managed Apache Beam runner for data pipelines, not a Hadoop/Spark YARN cluster environment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Compute Engine
Why this is correct
Compute Engine provides the virtual machines that form the worker and master nodes of a Cloud Dataproc cluster. When you run a Hadoop/Spark workload on Google Cloud, Cloud Dataproc orchestrates the deployment, but the actual CPU, memory, and local storage attached to each cluster node are Compute Engine instances. You can also run Hadoop/Spark directly on your own Compute Engine VMs without Dataproc, making Compute Engine the fundamental compute infrastructure for such workloads.
- ✓
Cloud Dataproc
Why this is correct
Cloud Dataproc is the fully managed service that simplifies running open-source Apache Hadoop and Spark clusters on Google Cloud. It creates and manages cluster nodes as Compute Engine VMs, auto-scales them, and integrates with Cloud Storage for persistent data. Instead of manually installing Hadoop or Spark on raw VMs, Dataproc handles the cluster lifecycle, software configuration, and monitoring, making it the recommended way to execute Hadoop/Spark workloads.
- ✓
Cloud Storage
Why this is correct
Cloud Storage replaces HDFS as the default data layer in a Cloud Dataproc cluster, allowing you to decouple compute from storage. Since HDFS is ephemeral and tied to cluster lifecycle, storing input and output data in Cloud Storage enables you to delete clusters after jobs finish without losing data, and to share data across clusters. Dataproc reads and writes data directly from and to Cloud Storage buckets via the gs:// connector, avoiding expensive HDFS replication overhead and enabling elastic cluster scaling.
- ✗
BigQuery
Why it's wrong here
BigQuery is a serverless, highly scalable enterprise data warehouse that runs analytics using SQL, not Apache Hadoop or Spark. It cannot run Hadoop/Spark jobs or host YARN/executors; it is a separate service for storing and querying structured data. While you can use BigQuery to analyze data that originated from Hadoop workloads, it does not provide the same execution environment and is not an alternative for running Spark code.
- ✗
Dataflow
Why it's wrong here
Dataflow is a unified stream and batch data processing service based on the Apache Beam programming model, not on Hadoop or Spark. It is intended for building data pipelines that transform and enrich data, but it does not run Java/Scala Spark jobs or provide a Hadoop-compatible distributed filesystem. Choosing Dataflow would require rewriting your Hadoop/Spark code into Beam transforms, so it is not a drop-in replacement for a Hadoop/Spark workload.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
Learn chapter
Modernizing Applications with GCP
Key term
Serverless
Serverless is a cloud computing model where the cloud provider manages the servers, and you only pay for the actual compute time your code uses, without having to worry about provisioning or maintaining infrastructure.
Key term
Data warehouse
A data warehouse is a central repository that stores large amounts of structured data from multiple sources, optimized for querying and analysis rather than day-to-day transactions.
About these practice questions
This GCDL question is part of Courseiva's 848-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This GCDL practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the GCDL exam.