Courseiva

PDE Preparing and Using Data for Analysis Practice Question

You need to load a large CSV file from Cloud Storage into BigQuery. The file has a header row and contains a column with date values in the format 'YYYY-MM-DD'. You want to ensure the date column is correctly recognized as a DATE type. Which method should you use?

⚠ Common exam trap

The trap here is relying on autodetect to infer DATE types, but autodetect can be inconsistent and may default to STRING if the format is not perfectly recognized.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Define an explicit schema in JSON or inline that specifies the date column as DATE, and provide it during the load job.

Defining an explicit schema that specifies the date column as DATE ensures BigQuery correctly interprets the values during load. This method is reliable and avoids the uncertainty of autodetect or post-load casting, making it the best practice for date columns with a known format.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Define an explicit schema in JSON or inline that specifies the date column as DATE, and provide it during the load job.

    Why this is correct

    Providing an explicit schema ensures BigQuery interprets the date column as DATE regardless of the string format. This is the most reliable method, especially when the format matches the expected DATE literal. It avoids ambiguity and guarantees correct type recognition, which is essential for subsequent date functions.

  • ✗

    Use the bq load command with the --autodetect flag and no schema definition.

    Why it's wrong here

    Autodetect can infer DATE types from strings in 'YYYY-MM-DD' format, but it is not guaranteed and may misinterpret if the data is ambiguous. For reliable DATE recognition, explicitly specifying the schema is safer. Autodetect also may not handle all date formats consistently, so it is not the best choice when precision is required.

  • ✗

    Load the file as a string column, then use a SQL query to cast the string to DATE after loading.

    Why it's wrong here

    Loading as string and casting later adds an extra step and may fail if the string format is not perfectly parseable. It also means the column is initially stored as STRING, which is inefficient for storage and query performance. Explicit schema definition during load is simpler and more accurate.

  • ✗

    Convert the CSV to a newline-delimited JSON file with date values as strings, then load with autodetect.

    Why it's wrong here

    Converting to JSON and relying on autodetect does not guarantee DATE type recognition; it may still infer STRING. The conversion adds unnecessary complexity and does not solve the core issue of ensuring the date column is typed as DATE. Explicit schema definition is the direct solution.

About these practice questions

Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Google Cloud exam blueprint

This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.