Courseiva

PDE Preparing and Using Data for Analysis Practice Question

You are building a data pipeline that ingests JSON files from Cloud Storage into BigQuery. The JSON files contain deeply nested arrays and objects. You need to load these files with minimal transformation so that analysts can query individual nested fields using dot notation and also unnest arrays when needed. Which approach should you use?

⚠ Common exam trap

The trap here is assuming that BigQuery automatically flattens nested JSON during load, when in fact it preserves nested structures as RECORD and REPEATED types if the schema is defined accordingly.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Load the JSON files into BigQuery with a schema that defines nested fields as RECORD and arrays as REPEATED, allowing queries to reference nested fields with dot notation and use UNNEST for arrays.

BigQuery supports semi-structured data through nested and repeated fields. Defining nested objects as RECORD and arrays as REPEATED preserves the original structure, enabling dot notation for nested fields and UNNEST for arrays. This avoids flattening during load, aligning with minimal transformation and providing flexible query capabilities.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    First load the JSON files into a Cloud SQL PostgreSQL instance, then use federated queries to access the data from BigQuery.

    Why it's wrong here

    Cloud SQL federated queries allow querying external data, but they do not natively expose nested JSON structures as BigQuery RECORD and REPEATED types. This adds unnecessary complexity, latency, and cost. The scenario requires loading into BigQuery directly with nested support, not federating from a relational database.

  • ✓

    Load the JSON files into BigQuery with a schema that defines nested fields as RECORD and arrays as REPEATED, allowing queries to reference nested fields with dot notation and use UNNEST for arrays.

    Why this is correct

    BigQuery natively supports nested and repeated fields via RECORD and REPEATED types. When JSON is loaded with such a schema, nested objects become RECORDs and arrays become REPEATED fields. Analysts can then use dot notation to access nested fields and UNNEST to flatten arrays in queries, meeting the requirement without transformation.

  • ✗

    Convert the JSON files to CSV using a Dataflow pipeline that flattens all arrays, then load the CSV into BigQuery.

    Why it's wrong here

    Converting to CSV and flattening arrays would lose the nested structure and require additional transformation. The requirement is minimal transformation and the ability to query nested fields with dot notation and unnest arrays. This approach destroys the hierarchy and does not meet the querying needs.

  • ✗

    Load the JSON files into BigQuery using the autodetect schema option, which will automatically flatten all nested structures into separate columns.

    Why it's wrong here

    Autodetect infers a schema but does not flatten nested structures. BigQuery preserves nested and repeated fields as RECORD and REPEATED types when loading JSON. Flattening would require explicit transformation, which contradicts the requirement for minimal transformation and dot notation access.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.