PDE Preparing and Using Data for Analysis Practice Question
You need to load a large CSV file from Cloud Storage into BigQuery. The file has a header row and contains a column with date values in the format 'YYYY-MM-DD'. You want to ensure the date column is correctly recognized as a DATE type. Which method should you use?
⚠ Common exam trap
The trap here is relying on autodetect to infer DATE types, but autodetect can be inconsistent and may default to STRING if the format is not perfectly recognized.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Define an explicit schema in JSON or inline that specifies the date column as DATE, and provide it during the load job.
Defining an explicit schema that specifies the date column as DATE ensures BigQuery correctly interprets the values during load. This method is reliable and avoids the uncertainty of autodetect or post-load casting, making it the best practice for date columns with a known format.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Define an explicit schema in JSON or inline that specifies the date column as DATE, and provide it during the load job.
Why this is correct
Providing an explicit schema ensures BigQuery interprets the date column as DATE regardless of the string format. This is the most reliable method, especially when the format matches the expected DATE literal. It avoids ambiguity and guarantees correct type recognition, which is essential for subsequent date functions.
- ✗
Use the bq load command with the --autodetect flag and no schema definition.
Why it's wrong here
Autodetect can infer DATE types from strings in 'YYYY-MM-DD' format, but it is not guaranteed and may misinterpret if the data is ambiguous. For reliable DATE recognition, explicitly specifying the schema is safer. Autodetect also may not handle all date formats consistently, so it is not the best choice when precision is required.
- ✗
Load the file as a string column, then use a SQL query to cast the string to DATE after loading.
Why it's wrong here
Loading as string and casting later adds an extra step and may fail if the string format is not perfectly parseable. It also means the column is initially stored as STRING, which is inefficient for storage and query performance. Explicit schema definition during load is simpler and more accurate.
- ✗
Convert the CSV to a newline-delimited JSON file with date values as strings, then load with autodetect.
Why it's wrong here
Converting to JSON and relying on autodetect does not guarantee DATE type recognition; it may still infer STRING. The conversion adds unnecessary complexity and does not solve the core issue of ensuring the date column is typed as DATE. Explicit schema definition is the direct solution.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.