Courseiva

PDE Ingesting and Processing the Data Practice Question

Your organization runs a batch Dataflow pipeline that reads from BigQuery, transforms records, and writes Parquet files to Cloud Storage partitioned by event date. The pipeline currently writes all files into a single directory and downstream Hive-style queries scan the entire dataset. You need to restructure the output so that queries scan only the relevant date partitions, while keeping the pipeline idempotent on re-runs. What should you do?

⚠ Common exam trap

The trap here is assuming that embedding the date in a file name or file extension is enough for partition pruning, when engines actually prune on directory paths such as event_date=YYYY-MM-DD.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use FileIO.writeDynamic with a destination function that maps each record to a path like gs://bucket/events/event_date=YYYY-MM-DD/, and write to a temporary location before atomically moving files into place.

Hive-style partition pruning depends on directory structure, not file naming. FileIO.writeDynamic with a destination function that emits paths containing event_date=YYYY-MM-DD creates the directories that query engines use to skip irrelevant data. Writing first to a temporary location and then moving completed files into the final partition directory prevents partially written partitions from being visible if the pipeline fails and is re-run, preserving idempotency.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use AvroIO with a custom filename policy that includes the event date, and enable the pipeline's --diskSizeGb option to improve write throughput.

    Why it's wrong here

    A custom filename policy can embed the date in file names, but it still writes into a single directory, so partition pruning at the directory level is not possible. Increasing disk size affects worker local storage, not output layout. The scenario explicitly requires Hive-style partitioned directories, which this option does not produce, and it does not address idempotency on re-runs.

  • ✓

    Use FileIO.writeDynamic with a destination function that maps each record to a path like gs://bucket/events/event_date=YYYY-MM-DD/, and write to a temporary location before atomically moving files into place.

    Why this is correct

    writeDynamic lets each record choose its output directory based on the event date, producing Hive-style partition paths that downstream engines prune. Writing to a temporary location and then moving files into the final partition directory ensures that a failed re-run does not leave partial data visible to queries, which preserves idempotency. This combination directly satisfies both the partition-pruning and idempotency requirements.

  • ✗

    Write all files to a single directory and create a BigQuery external table over the bucket, then rely on BigQuery's automatic partition pruning from the file names.

    Why it's wrong here

    BigQuery external tables over Cloud Storage do not infer Hive-style partitions from file names unless you explicitly declare partition columns and use a partitioned external table definition. Even then, the files must already be organized into date-based directories for pruning to work. Leaving all files in one directory means every query scans the full dataset, defeating the goal of scanning only relevant dates.

  • ✗

    Add a GroupByKey transform keyed by event date before writing, then use TextIO to write one file per date into a flat directory with the date embedded in the file name.

    Why it's wrong here

    Grouping by date and embedding the date in a flat file name does not create directory-level partitions, so Hive-style partition pruning cannot skip files. Downstream engines would still list and open every file in the flat directory. This approach also risks memory pressure from GroupByKey on large dates and does not provide the atomic move needed for idempotent re-runs.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.