PDE Ingesting and Processing the Data Practice Question
A media company stores 400 TB of compressed JSON clickstream logs in a Cloud Storage bucket. Analysts need to run ad hoc SQL over this data several times a day, and the team wants to avoid the cost and delay of loading it into BigQuery. Which BigQuery capability should you use?
⚠ Common exam trap
The trap here is treating the BigLake connection as a data movement mechanism, when it is actually a governance and metadata layer that lets BigQuery read the Cloud Storage objects in place without copying them.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create an external table over the Cloud Storage files using the BigLake connection, and query it directly.
External tables over Cloud Storage, accessed through a BigLake connection, keep the data in its original bucket while exposing it to BigQuery SQL. This removes both the load latency and the duplicate storage cost that the team is trying to avoid. Native loading, bucket-to-bucket copying, and streaming inserts all either duplicate the data or provide no query surface at all.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Load the JSON files into a BigQuery native table partitioned by ingestion time.
Why it's wrong here
Loading the files creates a second copy of the 400 TB inside BigQuery managed storage, which incurs storage charges and requires a load job before queries can run. The team explicitly wants to avoid that cost and delay, so a native load, however well partitioned, answers the wrong constraint even though it would perform well for repeated analytics.
- ✓
Create an external table over the Cloud Storage files using the BigLake connection, and query it directly.
Why this is correct
BigQuery external tables backed by a BigLake connection let you run standard SQL over data that remains in Cloud Storage, so the clickstream logs are queried in place without an ingestion job. This satisfies the ad hoc querying need and avoids duplication of the 400 TB, while the connection provides fine-grained access control and metadata caching.
- ✗
Use the Storage Transfer Service to copy the bucket into a second bucket in the same region as the dataset.
Why it's wrong here
Storage Transfer Service moves or copies objects between storage locations; it does not make them queryable with SQL. Relocating the bucket may reduce network egress for later loads, but analysts still have no table to query, so this step alone leaves the ad hoc SQL requirement unmet and simply adds another copy of the data.
- ✗
Create a BigQuery dataset and use the bq command-line tool to stream the JSON records with the insertAll API.
Why it's wrong here
Streaming inserts are designed for small, continuous record-level writes and are billed per row; pushing 400 TB of historical logs this way is prohibitively expensive and slow. It also materializes all data in native storage, contradicting the goal of querying files in place, so it is a poor fit for periodic ad hoc analysis of large archives.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.