PDE Storing the Data Practice Question
A financial analytics team ingests trade records into BigQuery every minute. Queries filter almost exclusively on trade_date and account_id, and the table grows by roughly 400 GB per day. Analysts report that monthly reports scanning the last 30 days are slow and expensive. The team wants to reduce bytes scanned without changing the ingestion pipeline. What should they do?
⚠ Common exam trap
The trap here is assuming a materialized view or BI Engine is a substitute for partitioning and clustering when the real issue is scanning too much table data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Partition the table by trade_date and cluster it by account_id.
Partitioning by the trade_date column allows BigQuery to skip partitions outside the query's date range, and clustering by account_id further prunes blocks within each retained partition. This combination directly reduces bytes scanned for the described monthly reports, improving both latency and cost while leaving the ingestion pipeline untouched.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable BigQuery BI Engine reservation for the reporting dataset.
Why it's wrong here
BI Engine caches and accelerates compatible queries, but it is a memory-based acceleration layer with capacity limits and does not reduce the underlying bytes scanned for a 30-day, 400 GB-per-day table. It also does not address storage layout, so query cost from scanning remains largely unchanged for large monthly reports.
- ✓
Partition the table by trade_date and cluster it by account_id.
Why this is correct
Partitioning by trade_date lets BigQuery prune all partitions outside the 30-day window, and clustering by account_id keeps account-filtered blocks co-located within each partition. Together they cut bytes scanned for both the date range and the account predicate, directly addressing the slow, expensive monthly reports without touching ingestion.
- ✗
Convert the table to a BigQuery external table over Cloud Storage Parquet files.
Why it's wrong here
External tables do not provide the same partition pruning and clustering benefits as native partitioned tables, and they can be slower for repeated analytical queries. Moving ingestion to Cloud Storage would also change the pipeline, which the team explicitly wants to avoid, so this does not solve the cost and latency problem.
- ✗
Create a materialized view that pre-aggregates daily totals per account.
Why it's wrong here
A materialized view accelerates repeated aggregations, but the monthly reports are not limited to pre-aggregated totals and may need row-level detail. It also adds refresh cost and does not prune the 30-day scan for arbitrary queries, so it does not reliably meet the goal of reducing bytes scanned across the reporting workload.
Go deeper
Related to this question
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.