Courseiva
Storing the Data →mediumMultiple Choice

PDE Storing the Data Practice Question

A media company ingests thousands of small JSON files per hour into a Cloud Storage bucket and wants to analyze them with BigQuery. Analysts frequently filter by event date and by device type, and they want to minimize query cost. The team wants a managed approach that avoids writing custom transformation code. Which BigQuery feature should the engineer use?

⚠ Common exam trap

The trap here is assuming that any external or federated access to Cloud Storage is automatically cheaper, when repeated analytical queries over raw small files usually cost more than loading into partitioned native storage.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use a BigQuery load job to create a native table partitioned by event date and clustered by device type.

Loading the JSON into a native, partitioned, and clustered table moves the data into BigQuery managed columnar storage, where filters on event date prune partitions and filters on device type benefit from clustering. This reduces bytes scanned and therefore cost, and it requires no custom transformation code. External and federated approaches keep querying the raw small files, which is less efficient and more expensive for repeated analysis.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Create an external table over the bucket and rely on automatic schema detection.

    Why it's wrong here

    External tables query the files in place, but each query scans the underlying objects, and many small JSON files reduce parallelism and increase bytes scanned. Cost is driven by the bytes read from Cloud Storage, and there is no partitioning or clustering benefit unless the external table is configured with those features. Automatic schema detection also has limitations and can misinfer types across heterogeneous JSON.

  • ✓

    Use a BigQuery load job to create a native table partitioned by event date and clustered by device type.

    Why this is correct

    Loading the JSON into a native table consolidates the small files into BigQuery managed storage, and partitioning by event date lets queries prune to the relevant date range. Clustering by device type further reduces bytes scanned for filters on that column. A standard load job requires no custom transformation code, matching the managed requirement, and native storage gives the best query performance and cost profile.

  • ✗

    Create a federated query that reads the bucket through the Cloud Storage connector.

    Why it's wrong here

    Federated queries against Cloud Storage are functionally similar to external tables and still scan the raw objects on each query. They do not consolidate small files, do not provide managed partitioning beyond what you configure, and can be slower than native storage for repeated analytical workloads. This option does not address the cost or performance goals described.

  • ✗

    Define a BigLake external table with a metadata cache and partition pruning enabled.

    Why it's wrong here

    BigLake external tables add fine-grained security and a metadata cache, and they can prune partitions, but they still read the underlying Cloud Storage objects for query execution. The many small JSON files remain a performance and cost concern because bytes scanned come from object storage rather than managed columnar storage. This adds capability the scenario does not require while leaving the small-file problem unsolved.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.