PDE Maintaining and Automating Data Workloads Practice Question
A data engineer needs to run a recurring nightly extract-transform-load job that pulls data from a REST API, applies Python transformations, and writes the output to a Cloud Storage bucket. The team wants a fully managed, serverless scheduler that can retry failed runs and send notifications, and they do not want to maintain any cluster or VM. Which Google Cloud service should they use to define and run this job?
⚠ Common exam trap
The trap here is treating Dataproc Serverless as a complete scheduling solution, when it only runs the compute and still needs an external trigger for a nightly cadence.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Cloud Scheduler triggering a Cloud Function that performs the API call, transformation, and Cloud Storage write.
Cloud Scheduler plus Cloud Functions delivers a fully managed, serverless combination for a recurring ETL task. Cloud Scheduler handles the cron cadence and can target a function over HTTP or Pub/Sub, while Cloud Functions executes the Python code with automatic retries, integrated logging, and the ability to publish notifications. No cluster or VM is required, satisfying the operational constraints.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A Compute Engine instance running a cron job that executes a Python script and writes to Cloud Storage.
Why it's wrong here
A Compute Engine VM running cron is not fully managed and requires the team to patch the OS, monitor the instance, and handle failures manually. It does not provide built-in retry semantics or notification integration out of the box. While it can run the Python transformation, the operational burden and lack of managed scheduling contradict the stated requirements, making it an unsuitable choice here.
- ✗
Cloud Composer with a single-node environment and a KubernetesPodOperator that runs the Python transformation.
Why it's wrong here
Cloud Composer is managed, but it provisions a GKE cluster and its supporting infrastructure, which conflicts with the requirement to maintain no cluster. A single-node Composer environment still runs the Airflow scheduler, web server, and workers on managed nodes that the team must size and pay for. For a single linear Python job, this is heavier than necessary and introduces cluster maintenance overhead.
- ✓
Cloud Scheduler triggering a Cloud Function that performs the API call, transformation, and Cloud Storage write.
Why this is correct
Cloud Scheduler provides a fully managed cron service that can trigger an HTTP or Pub/Sub target on a schedule, and Cloud Functions runs the Python code serverlessly with automatic retries and logging. This combination meets the requirement of no cluster or VM to maintain, supports retries on failure, and can publish to a Pub/Sub topic for notifications. It is the lightest managed option for a recurring single-step ETL job.
- ✗
Dataproc Serverless with a PySpark batch that calls the REST API and writes results to Cloud Storage.
Why it's wrong here
Dataproc Serverless is serverless in the sense that no cluster is provisioned, but it is designed for Spark workloads and does not include a scheduler. The team would still need Cloud Scheduler or Composer to trigger the batch on a nightly cadence. For a simple Python transformation that does not require Spark, this adds unnecessary complexity and cost compared to a function-based approach.
Quick reference
Cloud Service Model Comparison
| Model | You Manage | Provider Manages | Examples |
|---|---|---|---|
| IaaS | OS, runtime, apps, data | Hardware, hypervisor, networking | EC2, Azure VMs, GCP Compute Engine |
| PaaS | Apps and data | OS, runtime, middleware, hardware | Elastic Beanstalk, Azure App Service |
| SaaS | Data and settings only | Everything else | Microsoft 365, Salesforce, Workday |
| FaaS / Serverless | Function code only | Infra, scaling, runtime | Lambda, Azure Functions, Cloud Run |
| CaaS | Containers and apps | Kubernetes, OS, hardware | EKS, AKS, GKE |
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.