Courseiva
Storing the Data →mediumMultiple Choice

PDE Storing the Data Practice Question

A data engineer is designing a BigQuery table for a clickstream dataset with frequent queries aggregating over user sessions. Each user session has multiple events, and the engineer wants to avoid joins for performance. Which schema design pattern should they use?

⚠ Common exam trap

PDE often tests the trade-off between normalization and denormalization in BigQuery; candidates may default to normalized designs or clustering/partitioning without recognizing that nested repeated fields eliminate joins.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use nested and repeated fields to store events within each session row

BigQuery's nested and repeated fields (using STRUCT and ARRAY) allow you to store multiple events within a single session row, eliminating the need for joins when aggregating over sessions. This denormalized pattern is ideal for clickstream data and improves query performance by keeping related data together.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Use a normalized schema with separate tables for sessions and events, then join on session ID

    Why it's wrong here

    Normalising sessions and events requires the joins the engineer wants to avoid, and BigQuery charges by bytes scanned across those joins. It is tempting for transactional workloads needing update consistency, but nested repeated RECORD fields storing events inside each session row are the correct pattern here.

  • ✗

    Store each event as a separate row with session key and use clustering on session ID

    Why it's wrong here

    Storing events as separate rows still requires grouping by session ID to aggregate, so it does not remove the join-like work the scenario targets. Clustering merely co-locates those rows physically. It would suit append-heavy event logging where per-event filtering dominates, not session-level aggregation.

  • ✗

    Use partitioning on event timestamp and clustering on user ID

    Why it's wrong here

    Partitioning by event timestamp prunes by time, and clustering by user ID orders within partitions, but neither nests events under their session, so session aggregation still needs a join or GROUP BY. This suits time-range scans filtered by user, not join-free session rollups.

  • ✓

    Use nested and repeated fields to store events within each session row

    Why this is correct

    Nested and repeated fields let BigQuery store each session's events inside one row, eliminating the join between sessions and events that the stem explicitly wants to avoid. Aggregations over sessions then scan a single denormalised table, and BigQuery's columnar storage reads only the referenced nested attributes.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.