Cloud Digital Leader Practice Question: Google Cloud products, services, and solutions
A data analytics team needs to run a one-time transformation on 10 TB of data stored in Cloud Storage, then load the results into BigQuery. The transformation is a custom Java application that reads files, processes them, and writes to a new location. Which service should they use to minimize operational overhead?
⚠ Common exam trap
The trap is to choose Cloud Functions due to its serverless nature, but it cannot handle 10 TB in a single invocation. Another trap is confusing Dataproc Serverless (which is for Spark jobs) with Dataflow (which is for Beam pipelines in Java).
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Dataflow with Apache Beam Java SDK
Cloud Dataflow with Apache Beam Java SDK is the appropriate service for running custom Java transformations on large datasets in a fully managed, serverless manner. It reads from Cloud Storage, processes data, and can load results to BigQuery without provisioning clusters. Other options like Cloud Functions have execution limits, GKE requires cluster management, and Dataproc Serverless is intended for Spark jobs, not arbitrary custom Java applications.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Cloud Dataflow with Apache Beam Java SDK
Why this is correct
Cloud Dataflow with Apache Beam is designed for building robust, scalable, and often continuous data processing pipelines, not primarily for executing a standalone, one-time custom Java application directly. While it supports Java and handles large datasets, adapting a custom Java application to the Apache Beam model for a single execution introduces unnecessary development and conceptual overhead. It would be the correct choice if the transformation was a complex, recurring, or streaming job requiring a unified programming model and managed orchestration for high scalability and fault tolerance.
- ✗
Google Kubernetes Engine (GKE) with a custom container
Why it's wrong here
Running a custom container on Google Kubernetes Engine for a one-time transformation forces the team to provision, secure, and maintain a Kubernetes cluster, including node pools and autoscaling policies, even though the job runs only once. After the job finishes, the cluster remains an idle cost unless manually torn down, adding operational overhead that defeats the purpose of serverless ephemeral processing. GKE is better suited for long-running, always-on services or complex multi-tier workloads requiring orchestration, not for a short-lived batch ETL task.
- ✗
Dataproc Serverless with Spark job
Why it's wrong here
Dataproc Serverless runs Apache Spark jobs in a fully managed, autoscaling environment that eliminates the need to create or manage a Dataproc cluster, making it the ideal choice for an occasional one-time transformation. The service launches a dedicated job execution environment, scales resources based on the workload, and shuts down automatically after completion, so the team only pays for the job's duration. Because it natively supports Spark's data processing APIs, the team can deploy an existing Spark transformation with minimal rewrites and no infrastructure management.
- ✗
Cloud Functions triggered by Cloud Storage events
Why it's wrong here
Cloud Functions is an event-driven function as a service designed for lightweight, quick tasks such as webhooks or real-time notifications, and it imposes hard limits on execution time and memory allocation that would constrain a large-scale transformation. A Cloud Storage event handler is invoked by object creation or changes, not by an explicit one-time job submission, requiring the team to stage a trigger file and work around the platform's stateless, single-function execution model. For substantial data processing, this leads to timeouts and out-of-memory failures rather than the distributed compute needed.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
Learn chapter
Compute Comparison: VMs vs Containers vs Serverless
Key term
Serverless
Serverless is a cloud computing model where the cloud provider manages the servers, and you only pay for the actual compute time your code uses, without having to worry about provisioning or maintaining infrastructure.
Key term
Dataproc
Dataproc is a managed cloud service for running Apache Spark and Apache Hadoop clusters, allowing you to process large datasets quickly and economically.
About these practice questions
Courseiva writes every GCDL question from scratch — 848 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This GCDL practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the GCDL exam.