PDE Storing the Data Practice Question
A data engineer is designing a BigQuery table for a fraud detection system. The table will store 500 TB of transaction records and will be queried constantly by filtering on a transaction timestamp column and a customer_id column. The engineer needs to minimize the amount of data scanned by these queries. What should the engineer do?
⚠ Common exam trap
The trap here is assuming that any performance feature, such as a materialized view or the Storage Write API, will reduce bytes scanned, when only partitioning and clustering directly affect query scan cost.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Partition the table by the transaction timestamp column and cluster by the customer_id column.
Partitioning by transaction timestamp and clustering by customer_id is the recommended approach for large BigQuery tables that are frequently filtered on a time column and a high-cardinality key. Partition pruning limits scanned partitions, and clustering sorts data within partitions so that filters on customer_id read fewer blocks, lowering both cost and latency.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Create a materialized view that pre-aggregates transactions by customer_id and timestamp.
Why it's wrong here
A materialized view can accelerate specific aggregations, but it does not reduce the scan cost of arbitrary filtered queries against the base table. For a fraud detection system needing flexible filters, pre-aggregation may not cover all query patterns and adds maintenance overhead without guaranteeing lower bytes scanned.
- ✗
Use the BigQuery Storage Write API to stream data into the table with a fixed schema.
Why it's wrong here
The Storage Write API is an ingestion mechanism that improves streaming throughput and exactly-once semantics, but it has no effect on how queries scan the stored data. It does not create partitioning or clustering, so it will not reduce the bytes scanned by timestamp and customer_id filters.
- ✗
Set the table's expiration time to 30 days to limit the amount of historical data stored.
Why it's wrong here
Table expiration deletes old data automatically, which reduces storage volume but not the scan cost for the data that remains. Fraud detection typically requires long retention, and this approach would remove historical records rather than optimize query performance for the full dataset.
- ✓
Partition the table by the transaction timestamp column and cluster by the customer_id column.
Why this is correct
Partitioning by transaction timestamp restricts scans to the relevant time ranges, and clustering by customer_id further reduces the data read within each partition. This combination is the standard BigQuery optimization for high-volume tables filtered on time and a high-cardinality key, directly minimizing bytes scanned and cost.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.