A data engineer is building a data lake on Google Cloud and needs to separate raw ingested data, curated/cleaned data, and processed/aggregated data. Which Cloud Storage bucket structure is recommended?
Trap 1: Store all data in one bucket and use object labels to distinguish…
Labels are not as effective for lifecycle management and access control as separate paths or buckets.
Trap 2: Store raw data in a different project for security isolation.
Overly complex; typical data lakes use a single project with proper IAM and bucket structure.
Trap 3: Use different storage classes for raw, curated, and processed data…
Storage class is separate from organization; you still need a folder structure to apply rules per zone.
- A
Create three separate folders in a single bucket: raw, curated, processed.
Using prefixes (folders) within a bucket is a standard pattern for organizing data lake zones, allowing different lifecycle rules per prefix.
- B
Store all data in one bucket and use object labels to distinguish raw, curated, and processed.
Why it fails: Labels are not as effective for lifecycle management and access control as separate paths or buckets.
- C
Store raw data in a different project for security isolation.
Why it fails: Overly complex; typical data lakes use a single project with proper IAM and bucket structure.
- D
Use different storage classes for raw, curated, and processed data within the same bucket.
Why it fails: Storage class is separate from organization; you still need a folder structure to apply rules per zone.