Courseiva
Storing the Data →mediumMultiple Choice

PDE Storing the Data Practice Question

A media analytics company ingests 8 TB of new JSON event logs into BigQuery every day and keeps all history for 5 years. Analysts almost never filter on the raw event payload column and only occasionally select it, but they frequently filter on event_date and user_id. Storage cost is the top concern, and query performance on the frequently filtered columns must stay fast. What should the data engineer do?

⚠ Common exam trap

The trap here is assuming that partitioning, clustering, or long-term storage pricing reduces the cost of a large column that queries still scan, when only removing that column from the table actually does.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Store the raw event payload in a Cloud Storage bucket and keep only the structured, frequently queried columns in the BigQuery table, referencing the object path.

The workload filters on structured columns but almost never touches the large payload, so the payload is dead weight in the columnar store. Keeping only the structured columns in BigQuery preserves fast partitioned and clustered filtering, while relocating the payload to Cloud Storage removes the dominant storage and scan cost. Partitioning, clustering, the JSON type, and long-term storage all leave that large column inside the table, so none of them eliminate the core expense.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Store the raw event payload in a BigQuery column of type JSON so it is parsed at query time and billed as a separate data type.

    Why it's wrong here

    The JSON data type stores semi-structured data natively and offers more efficient encoding than STRING, but it does not remove the column from the physical scan. Since the payload is rarely filtered or selected, keeping it inside the table still contributes to the bytes billed and stored on every full-row scan, so this does not address the primary cost concern in this scenario.

  • ✗

    Partition the table by event_date and cluster it by user_id, then rely on the default columnar storage to avoid scanning the payload.

    Why it's wrong here

    Partitioning and clustering reduce the bytes scanned for filters on event_date and user_id, which is valuable, but BigQuery still reads every column referenced by the query. A SELECT * or any query touching the payload column will read that column's data. Because the payload is the largest column by far, the dominant storage and scan cost remains, so this alone does not solve the problem.

  • ✗

    Enable the BigQuery long-term storage pricing tier so data older than 90 days is billed at a lower rate automatically.

    Why it's wrong here

    Long-term storage automatically discounts partitions or tables that are not modified for 90 consecutive days, which helps retention cost, but it does nothing about the daily 8 TB of new payload data stored and scanned in the active window. The largest, least-queried column still dominates both storage and query bytes billed, so this does not meet the stated goal.

  • ✓

    Store the raw event payload in a Cloud Storage bucket and keep only the structured, frequently queried columns in the BigQuery table, referencing the object path.

    Why this is correct

    Moving the rarely filtered payload out of BigQuery and into Cloud Storage Standard or Nearline removes the largest column from the columnar table, cutting both active storage cost and bytes scanned for the common queries. The structured columns remain in BigQuery for fast filtering on event_date and user_id, and the object path can be used with external tables or remote functions when the payload is genuinely needed.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.