PDE Storing the Data Practice Question
A data team needs to run complex analytical queries on a dataset that is frequently updated with new rows. They want to minimize query costs and avoid scanning old data that is rarely queried. Which BigQuery feature should they use?
⚠ Common exam trap
A common mistake is to choose a performance optimization feature (clustering, materialized views, BI Engine) when the question explicitly asks about minimizing costs and avoiding scanning old data. The correct focus is on data lifecycle management with partition expiration.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Partitioned tables with partition expiration
Partitioned tables with partition expiration allow you to divide a table into segments based on a date/timestamp column, and automatically delete partitions that are older than a specified duration. This minimizes query costs by only scanning relevant partitions and eliminates storage costs for old, rarely queried data without manual intervention.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Partitioned tables with partition expiration
Why this is correct
Partitioning by date lets BigQuery prune irrelevant partitions, so queries scan only recent data, and partition expiration automatically deletes old partitions to cut storage and query costs. This suits frequently appended datasets where historical rows are rarely queried.
- ✗
BigQuery materialized views
Why it's wrong here
Materialised views precompute and incrementally refresh query results, but they do not restrict which base partitions a query scans, so old data is still read unless the query filters on a partition column. They suit accelerating repeated aggregate queries, not age-based pruning.
- ✗
Clustered tables
Why it's wrong here
Clustering sorts data within partitions but doesn't limit scan based on time; partitioning is needed to avoid scanning old data.
- ✗
BigQuery BI Engine
Why it's wrong here
BI Engine is an in-memory analysis cache that accelerates dashboards and BI queries; it does not partition storage, so old rows are still scanned and billed. It is tempting because it reduces latency for repeated interactive queries, which suits sub-second visualisation workloads rather than cost control on rarely queried historical data.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.