Courseiva
Storing the Data →mediumMultiple Choice

PDE Storing the Data Practice Question

A data engineer manages a BigQuery dataset holding a 40 TB partitioned table of point-of-sale transactions. Analysts frequently run queries that filter on a region column and a sale_date column together, but each query scans the full table because the region predicate is not reducing bytes billed. The engineer wants to reduce bytes scanned without changing the analytical queries or the write pipeline. What should the engineer do?

⚠ Common exam trap

The trap here is assuming that partitioning alone always prunes scans, when partitioning only helps for predicates on the partition column.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Create a clustered table on the region column, because clustering co-locates rows with similar values inside each partition so block pruning can skip irrelevant data.

Clustering is the right lever because it physically sorts data within each partition by the region column, enabling block pruning when region predicates are present. Partitioning already handles sale_date, so adding clustering on the other frequently filtered column reduces scanned bytes while leaving the queries and ingestion pipeline untouched.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Define a materialized view that aggregates the table by region and sale_date and let the optimizer rewrite all analyst queries to use it.

    Why it's wrong here

    Materialized views accelerate specific aggregation patterns and only get used when the query matches the view definition. The analysts run arbitrary filters and selections, so the optimizer cannot rewrite those queries to a pre-aggregated view, and the view would not be used, leaving bytes scanned unchanged.

  • ✗

    Enable the require_partition_filter option on the table so that every query must include a sale_date predicate.

    Why it's wrong here

    require_partition_filter only forces queries to include a partition filter; it does not reduce bytes scanned by a region predicate. Analysts already filter on sale_date, so this setting adds a constraint without changing scan size, and it would break any legitimate query that intentionally spans all partitions.

  • ✗

    Convert the table to an external table over Cloud Storage so that only the matching Parquet row groups are read at query time.

    Why it's wrong here

    External tables over Cloud Storage do not give BigQuery the same block-level pruning as native clustered storage, and moving the data out of managed storage changes the pipeline. The queries also lose the performance and slot efficiency of native partitioned storage, so bytes scanned would not reliably drop for region predicates.

  • ✓

    Create a clustered table on the region column, because clustering co-locates rows with similar values inside each partition so block pruning can skip irrelevant data.

    Why this is correct

    Clustering on region sorts data within each partition by that column, so BigQuery can prune storage blocks when a region predicate is present. Because the table is already partitioned on sale_date, adding clustering on region shrinks bytes scanned for queries that filter on both columns, and the write pipeline keeps inserting data unchanged.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.