PDE Storing the Data Practice Question
A data engineer is designing a Cloud Storage layout for a data lake that will be queried by BigQuery external tables and by Dataproc jobs. The engineer wants to minimize query cost and improve scan performance across both engines. Which two practices should the engineer follow? (Choose two.)
⚠ Common exam trap
The trap here is focusing on storage-class or versioning settings, which affect durability and retrieval cost, instead of the data layout choices that actually reduce bytes scanned.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Partition the data in Cloud Storage using a Hive-style key prefix such as dt=YYYY-MM-DD.
Columnar formats let both engines read only referenced columns and compress data efficiently, and Hive-style partitioning lets them prune entire prefixes when filters match the partition key. Together these reduce bytes scanned and listing overhead, which lowers query cost and improves scan performance for BigQuery external tables and Dataproc alike.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable object versioning on the bucket to preserve historical query results.
Why it's wrong here
Object versioning protects against accidental overwrites and deletions but does not affect query performance or bytes scanned. It also increases storage cost by retaining noncurrent versions, which is unrelated to the goal of minimizing query cost and improving scan performance.
- ✓
Partition the data in Cloud Storage using a Hive-style key prefix such as dt=YYYY-MM-DD.
Why this is correct
Hive-style partitioning lets BigQuery external tables and Dataproc prune entire prefixes when a query filters on the partition column, so only relevant directories are listed and read. This reduces both listing overhead and bytes scanned, improving performance and lowering cost for both engines.
- ✗
Store each row as a separate object to maximize parallelism during scans.
Why it's wrong here
One object per row creates millions of tiny files, which inflates listing and metadata overhead and prevents efficient block-level reads. Both BigQuery external tables and Dataproc suffer from small-file problems, so this approach increases cost and latency rather than improving scan performance.
- ✓
Store data in columnar formats such as Parquet or ORC instead of CSV.
Why this is correct
Columnar formats allow BigQuery and Dataproc to read only the columns referenced in a query, reducing bytes scanned and I/O. They also compress better than CSV, which lowers storage cost and speeds up scans for both external tables and Spark jobs that use predicate pushdown.
- ✗
Use the Standard storage class for all objects to ensure the lowest possible access latency.
Why it's wrong here
Storage class affects retrieval cost and availability, not the bytes scanned by a query or the efficiency of columnar reads. Choosing Standard for all data may increase storage cost for cold data without improving the scan performance that the engineer is targeting.
Go deeper
Related to this question
About these practice questions
Courseiva writes every PDE question from scratch — 747 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Google Cloud exam blueprint
This PDE practice question is part of Courseiva's free Google Cloud certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the PDE exam.