PDE Storing the Data Practice Question
A company uses BigQuery for analytics on petabyte-scale data. They want to improve query performance by denormalizing schemas and reducing joins. Which TWO BigQuery features should they use? (Choose 2)
⚠ Common exam trap
A common mistake is to think that only clustering or partitioning can replace denormalization, but nested/repeated fields are needed for schema denormalization. Clustering is a complementary physical optimization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Clustering on frequently filtered columns
Option E is correct because BigQuery's native support for nested and repeated fields (ARRAY<STRUCT<...>>) lets you store parent-child relationships inside a single table, which is the standard way to denormalize schemas and eliminate joins while preserving relational structure. Option A is correct because clustering on frequently filtered columns physically co-locates related rows within partitions, so BigQuery scans far less data and returns results faster for filtered queries on the denormalized table. Option B is not a BigQuery feature for denormalization; subqueries still execute as separate query blocks and do not remove join costs. Option C is wrong because external tables read data from Cloud Storage and generally offer worse performance than native tables, not better denormalized performance. Option D is wrong because partitioning by date improves pruning on time filters but does not denormalize schemas or reduce joins.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Clustering on frequently filtered columns
Why this is correct
Clustering physically co-locates rows sharing the clustered column values, so filters on those columns scan far fewer blocks. This directly supports denormalisation by pruning data before joins, cutting bytes read and improving performance on petabyte-scale tables.
- ✗
Using subqueries instead of JOINs
Why it's wrong here
Subqueries still execute as separate relational operations and often incur the same shuffle and join costs, so they do not denormalise the schema. They are useful for readability or isolating aggregation logic, but nested and repeated columns are what actually eliminate joins.
- ✗
External tables reading from Cloud Storage
Why it's wrong here
External tables query data held in Cloud Storage rather than native BigQuery storage, so they add network latency and cannot be denormalised into nested, repeated columns. They suit federating or querying raw files without loading, not restructuring petabyte schemas to eliminate joins.
- ✗
Table partitioning by date
Why it's wrong here
Partitioning by date prunes scanned data for date-filtered queries but leaves the schema normalised, so joins remain. It is the right choice for reducing bytes scanned and cost on time-series tables, not for denormalising schemas or removing join operations.
- ✓
Nested and repeated fields (ARRAY<STRUCT<...>>)
Why this is correct
Nested and repeated fields store related child records inside a single parent row as ARRAY<STRUCT<...>>, eliminating joins by keeping denormalised data co-located. This directly satisfies the requirement to denormalise schemas and reduce join operations at petabyte scale.
Go deeper
Related to this question
About these practice questions
One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.