Courseiva

PDE Maintaining and Automating Data Workloads Practice Question

A financial services company runs a Dataflow streaming pipeline that reads from Pub/Sub and writes enriched records to BigQuery. Compliance requires that the raw Pub/Sub messages be retained for seven years so that any record can be reprocessed if the enrichment logic is later found to be incorrect. The pipeline currently has no archival step. What should the data engineer do to satisfy the retention requirement with the least operational overhead?

⚠ Common exam trap

The trap here is assuming Pub/Sub retention can be stretched to meet a multi-year compliance window, when its retention is bounded and it is not an archival system.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Add a Cloud Storage sink to the Dataflow pipeline that writes each raw message to a Cloud Storage bucket with a lifecycle policy moving objects to Coldline or Archive storage after 30 days.

Retaining raw messages for years requires a durable, low-cost object store, and Cloud Storage with lifecycle-based class transitions is the standard fit on Google Cloud. By branching the existing Dataflow pipeline to write unmodified messages to a bucket, the company keeps a complete, reprocessable archive without building a separate ingestion path or managing long-lived Pub/Sub backlogs.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Configure a Pub/Sub dead-letter topic and set the maximum delivery attempts high enough that failed messages persist for the retention period.

    Why it's wrong here

    Dead-letter topics only receive messages that a subscription failed to acknowledge after the configured attempts. Successfully processed messages never reach the dead-letter topic, so the archive would be incomplete and useless for full reprocessing. Its retention is also subject to Pub/Sub limits, not years-long compliance windows.

  • ✗

    Enable BigQuery table snapshots on the destination table and set a seven-year expiration on each snapshot.

    Why it's wrong here

    Snapshots capture the state of the enriched BigQuery table, not the original Pub/Sub messages. If the enrichment logic is wrong, the raw input needed to recompute correct output is not present in any snapshot. Snapshot expiration also defaults to a limited window and does not substitute for durable raw-message archival.

  • ✓

    Add a Cloud Storage sink to the Dataflow pipeline that writes each raw message to a Cloud Storage bucket with a lifecycle policy moving objects to Coldline or Archive storage after 30 days.

    Why this is correct

    Writing the unmodified messages to Cloud Storage preserves the raw payload independently of the BigQuery output, and a lifecycle policy automatically transitions old objects to colder, cheaper classes so seven years of retention stays affordable. The Dataflow sink adds little operational burden because the pipeline already runs continuously and only needs one additional write branch.

  • ✗

    Increase the Pub/Sub topic message retention duration to the maximum allowed and rely on the subscription backlog for seven years.

    Why it's wrong here

    Pub/Sub message retention has a maximum measured in days, not years, so it cannot cover a seven-year requirement. Even if it could, retaining messages in a subscription backlog is expensive and does not provide a durable archive that can be queried or selectively reprocessed. This approach fails the compliance obligation outright.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.