Courseiva

PDE Ingesting and Processing the Data Practice Question

A logistics company uploads nightly shipment manifests as newline-delimited JSON files into a Cloud Storage bucket. Analysts need to query these files with standard SQL immediately after upload, but the schema evolves frequently and the company wants to avoid managing a load job. Which approach should they use?

⚠ Common exam trap

The trap here is assuming that querying files in Cloud Storage requires moving them into BigQuery storage first, when external tables can query them in place.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Create a BigQuery external table over the Cloud Storage bucket using the BigLake connection with JSON format and schema autodetect.

External tables let BigQuery read files directly from Cloud Storage, so analysts can run standard SQL over the manifests without a load job, and BigLake-backed external tables support schema autodetection for newline-delimited JSON. Because the schema changes often, autodetect combined with in-place querying avoids brittle load pipelines while still exposing the data through the BigQuery SQL surface.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Create a BigQuery external table over the Cloud Storage bucket using the BigLake connection with JSON format and schema autodetect.

    Why this is correct

    BigQuery external tables over Cloud Storage let you query files in place with standard SQL, and BigLake connections add governance plus support for schema autodetect on newline-delimited JSON. Because the data stays in Cloud Storage, no load job or pipeline maintenance is required, and new files matching the URI prefix are picked up automatically, which fits a frequently evolving manifest schema.

  • ✗

    Mount the bucket in Dataproc and run a Hive external table over the JSON files.

    Why it's wrong here

    Hive external tables live in the Dataproc/Hive metastore and are queried with HiveQL or Spark SQL, not the standard BigQuery SQL the analysts expect. It also requires maintaining a persistent Dataproc cluster, which is heavier and costlier than querying the files directly from BigQuery. This does not meet the immediate, low-maintenance requirement.

  • ✗

    Use the Storage Transfer Service to copy the JSON into a BigQuery dataset, then query the dataset.

    Why it's wrong here

    Storage Transfer Service moves objects between storage systems; it does not parse JSON or create BigQuery tables with inferred schemas. Pointing it at BigQuery is not a supported direct transfer for arbitrary JSON, and it would not handle the evolving schema. The analysts would still be unable to query the manifests without additional loading logic.

  • ✗

    Load each file into a native BigQuery table with a scheduled query that runs bq load every night.

    Why it's wrong here

    Native tables require the data to be copied into BigQuery storage, and bq load jobs must be scheduled and monitored. With a rapidly evolving schema, each load job can fail on type mismatches, and you would need to alter the table schema manually. This adds operational overhead the company explicitly wants to avoid for a simple manifest query pattern.

About these practice questions

One of 747 original PDE practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.