PDE Preparing and Using Data for Analysis Practice Question
Your analytics team queries a large BigQuery table `events` that is partitioned by `event_date` and clustered by `user_id`. A new dashboard runs a query filtering on `user_id = 'abc123'` but does not include a filter on `event_date`. You want to minimize the bytes scanned by this query. What should you do?
⚠ Common exam trap
The trap here is assuming that clustering alone can avoid scanning all partitions when the query lacks a filter on the partitioning column.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Add a filter on `event_date` to the query, such as `event_date >= '2024-01-01'`.
The table is partitioned by `event_date`; without a filter on that column, BigQuery cannot prune partitions and scans the entire table. Adding a date filter enables partition pruning, reducing bytes scanned. Clustering on `user_id` helps only within the scanned partitions, so it does not replace the need for a partition filter.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Add a filter on `event_date` to the query, such as `event_date >= '2024-01-01'`.
Why this is correct
Partition pruning is driven by filters on the partitioning column. Without an `event_date` predicate, BigQuery must scan all partitions, even though clustering on `user_id` helps within each partition. Adding a date filter restricts the scan to relevant partitions, dramatically reducing bytes processed and cost.
- ✗
Use `SELECT *` instead of selecting specific columns.
Why it's wrong here
Selecting all columns increases the amount of data read, not decreases it. BigQuery is columnar, so selecting only needed columns reduces bytes scanned. Using `SELECT *` would likely increase cost and slow the query, contrary to the goal of minimizing bytes scanned.
- ✗
Add a `LIMIT` clause to the query, such as `LIMIT 1000`.
Why it's wrong here
A `LIMIT` clause restricts the number of rows returned but does not reduce the amount of data scanned. BigQuery still processes all partitions and columns referenced in the query before applying the limit. It does not help with partition pruning or clustering benefits.
- ✗
Create a clustered table on `user_id` instead of `event_date`.
Why it's wrong here
Clustering on `user_id` alone would not provide partition pruning. While clustering can improve filter performance, it does not eliminate the need to scan all data if no partition filter is present. Recreating the table as clustered on `user_id` would not reduce bytes scanned for a query without a partition filter.
About these practice questions
This PDE question is part of Courseiva's 747-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.