Courseiva

PDE Maintaining and Automating Data Workloads Practice Question

A data engineer needs to run a recurring nightly extract-transform-load job that pulls data from a REST API, applies Python transformations, and writes the output to a Cloud Storage bucket. The team wants a fully managed, serverless scheduler that can retry failed runs and send notifications, and they do not want to maintain any cluster or VM. Which Google Cloud service should they use to define and run this job?

⚠ Common exam trap

The trap here is treating Dataproc Serverless as a complete scheduling solution, when it only runs the compute and still needs an external trigger for a nightly cadence.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Cloud Scheduler triggering a Cloud Function that performs the API call, transformation, and Cloud Storage write.

Cloud Scheduler plus Cloud Functions delivers a fully managed, serverless combination for a recurring ETL task. Cloud Scheduler handles the cron cadence and can target a function over HTTP or Pub/Sub, while Cloud Functions executes the Python code with automatic retries, integrated logging, and the ability to publish notifications. No cluster or VM is required, satisfying the operational constraints.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    A Compute Engine instance running a cron job that executes a Python script and writes to Cloud Storage.

    Why it's wrong here

    A Compute Engine VM running cron is not fully managed and requires the team to patch the OS, monitor the instance, and handle failures manually. It does not provide built-in retry semantics or notification integration out of the box. While it can run the Python transformation, the operational burden and lack of managed scheduling contradict the stated requirements, making it an unsuitable choice here.

  • ✗

    Cloud Composer with a single-node environment and a KubernetesPodOperator that runs the Python transformation.

    Why it's wrong here

    Cloud Composer is managed, but it provisions a GKE cluster and its supporting infrastructure, which conflicts with the requirement to maintain no cluster. A single-node Composer environment still runs the Airflow scheduler, web server, and workers on managed nodes that the team must size and pay for. For a single linear Python job, this is heavier than necessary and introduces cluster maintenance overhead.

  • ✓

    Cloud Scheduler triggering a Cloud Function that performs the API call, transformation, and Cloud Storage write.

    Why this is correct

    Cloud Scheduler provides a fully managed cron service that can trigger an HTTP or Pub/Sub target on a schedule, and Cloud Functions runs the Python code serverlessly with automatic retries and logging. This combination meets the requirement of no cluster or VM to maintain, supports retries on failure, and can publish to a Pub/Sub topic for notifications. It is the lightest managed option for a recurring single-step ETL job.

  • ✗

    Dataproc Serverless with a PySpark batch that calls the REST API and writes results to Cloud Storage.

    Why it's wrong here

    Dataproc Serverless is serverless in the sense that no cluster is provisioned, but it is designed for Spark workloads and does not include a scheduler. The team would still need Cloud Scheduler or Composer to trigger the batch on a nightly cadence. For a simple Python transformation that does not require Spark, this adds unnecessary complexity and cost compared to a function-based approach.

Quick reference

Cloud Service Model Comparison

ModelYou ManageProvider ManagesExamples
IaaSOS, runtime, apps, dataHardware, hypervisor, networkingEC2, Azure VMs, GCP Compute Engine
PaaSApps and dataOS, runtime, middleware, hardwareElastic Beanstalk, Azure App Service
SaaSData and settings onlyEverything elseMicrosoft 365, Salesforce, Workday
FaaS / ServerlessFunction code onlyInfra, scaling, runtimeLambda, Azure Functions, Cloud Run
CaaSContainers and appsKubernetes, OS, hardwareEKS, AKS, GKE

About these practice questions

This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.