Courseiva

PDE · domain

Storing the Data

The Storing the Data domain covers choosing and configuring Google Cloud storage and database services: Cloud SQL, Spanner, BigQuery, Cloud Storage, Bigtable, Firestore, and Memorystore. Questions test matching workload characteristics — scale, consistency, latency, schema, cost — to the right service, plus governance controls like VPC Service Controls, IAM, CMEK, and BigQuery external tables.

109 questions30 easy49 medium30 hard

Focused practice

Practice Storing the Data questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Storing the Data

Match each workload to the correct storage service using consistency, scale, latency, and schema requirements, then apply the right governance controls. The single most important thing: know when Spanner, BigQuery, Cloud Storage, Bigtable, or Firestore is the correct answer, and why the alternatives fail.

Selecting Spanner for globally consistent, horizontally scalable relational workloads with SQL joins and automatic failover.

Using BigQuery external tables or BigLake to query Parquet and other formats in Cloud Storage without loading.

Designing BigQuery schemas with nested and repeated fields to denormalize sessions and avoid joins.

Applying VPC Service Controls perimeter to restrict BigQuery and Cloud Storage access and prevent exfiltration.

Watch out for

Common Storing the Data exam traps

  • ▸Choosing Cloud SQL for global scale or Bigtable for SQL joins instead of Spanner, which supports both strong consistency and relational queries.
  • ▸Assuming BigQuery external tables perform like native tables; querying Cloud Storage directly is slower and lacks clustering and partitioning benefits.
  • ▸Treating VPC Service Controls as IAM only; it enforces network perimeters and does not replace least-privilege IAM roles.

Question index

All Storing the Data questions (109)

Click any question to see the full explanation, or start a practice session above.

1

A media company ingests thousands of small JSON files per hour into a Cloud Storage bucket and wants to analyze them with BigQuery. Analysts frequently filter by event date and by device type, and they want to minimize query cost. The team wants a managed approach that avoids writing custom transformation code. Which BigQuery feature should the engineer use?

Medium
2

A healthcare organization stores patient data in BigQuery. They need to encrypt a specific column (e.g., SSN) using a key they manage, and decrypt it only for authorized queries via a user-defined function. Which approach should they use?

Hard
3

A data engineer needs to create a unified table that combines data from Cloud Storage (Parquet files) and BigQuery native tables, with fine-grained access control and governance. Which three Google Cloud features should they use together? (Choose THREE.)

Hard
4

A media company stores 900 TB of video master files in a Cloud Storage bucket in the US multi-region. Legal requires that the objects be retained for seven years and cannot be deleted or overwritten by any user, including project owners, during that period. The company also wants to minimize storage cost for objects that are rarely accessed after the first 90 days. The data engineer must implement a compliant configuration. (Choose two.)

Hard
5

A data engineer wants to set up automatic deletion of objects from a Cloud Storage bucket after 30 days, and transition objects older than 7 days to Nearline storage. Which THREE steps should they take? (Select three.)

Easy
6

A mobile app needs a real-time NoSQL database that supports offline sync and automatic conflict resolution. Which Google Cloud database is best suited?

Easy
7

A company stores JSON-formatted application logs in Cloud Storage. They need to query these logs with SQL, but they want to avoid the cost and latency of loading them into BigQuery. The logs have a consistent schema, and queries will filter on a timestamp field and a few nested fields. Which approach should they use?

Hard
8

A healthcare company stores patient documents in a Cloud Storage bucket with a retention policy. Auditors require that once an object is written, it cannot be overwritten or deleted by any user, including project owners, for 7 years. The company also wants to minimize storage cost after the first year while preserving immutability. What should the data engineer configure?

Hard
9

A logistics company runs Cloud SQL for MySQL for its order management system. The database is 2 TB and growing 100 GB per month. They need to run analytical queries without impacting transactional performance, and they want to minimize operational overhead. They also require the analytics data to be no more than 15 minutes stale. Which storage approach should they use?

Hard
10

A financial services firm stores trade documents in a Cloud Storage bucket. Auditors require that each object be retained for exactly seven years and that no user, including project owners, be able to delete or overwrite the objects during that period. The firm also needs to prove compliance to auditors. Which combination of controls should the data engineer implement?

Hard
11

A media company stores 50 TB of video metadata in Cloud Storage buckets in the Standard class. Legal requires that records for a subset of titles be retained for exactly seven years and that they cannot be altered or deleted by any user, including project owners, during that period. The company also wants to minimize storage cost for a second bucket holding derived thumbnails that are accessed only a few times per year but must be retrievable within seconds. Which two configurations should the data engineer implement? (Choose two.)

Medium
12

A company has a data lake on Cloud Storage with raw data in the 'raw' bucket, curated data in 'curated', and processed data in 'processed'. They want to implement lifecycle management to reduce costs. Which TWO actions should they take? (Choose 2)

Medium
13

A company is designing a Cloud Bigtable row key for a time-series dataset of device readings. They want to avoid hotspotting (uneven load across tablets). Which TWO row key design patterns are effective? (Choose 2)

Medium
14

A company uses Cloud Spanner and needs to store a parent-child relationship where the child table is frequently queried together with the parent. The parent has millions of rows and the child billions. Which Spanner feature optimizes performance for this pattern?

Medium
15

A data engineer needs to enforce that all datasets in a project expire after 90 days to reduce storage costs. They want to automate this without manual intervention. Which approach should they use?

Medium
16

A logistics company writes shipment telemetry to a Cloud Bigtable table using a row key of shipment_id, where shipment_id values are monotonically increasing and roughly sequential over time. Write throughput has become uneven, with a small number of tablet servers overloaded while others sit idle. The engineer must improve write distribution without changing the schema of the column families. What should the engineer do?

Hard
17

A data engineer needs to restrict access to BigQuery datasets such that only data from approved VPC networks can query them. They also need to audit data access. Which two security controls should they implement? (Choose two.)

Medium
18

A financial services company uses Cloud Bigtable to store trade data. They are experiencing hot-spotting on a single node, causing high latency. The row key format is [trade_id]#[timestamp]. Which row key design change would BEST distribute writes across tablets?

Hard
19

A company has a Cloud SQL for PostgreSQL instance and wants to create a read replica to offload read traffic from the primary. They also need to ensure the replica is in a different region for disaster recovery. Which Cloud SQL feature should they use?

Medium
20

A data engineering team needs to run SQL analytics on a large BigQuery dataset from a Looker Studio dashboard. The dashboard must return results quickly, and the underlying data changes only once per day through a batch load. The team wants to minimize query cost and latency for repeated dashboard queries. What should they do?

Easy
21

An application requires a globally distributed, strongly consistent database with 99.999% availability SLA. The workload is OLTP with high throughput across continents. Which service fits best?

Medium
22

A company wants to run hybrid transactional and analytical workloads on a PostgreSQL-compatible database with high performance. Which service should they choose?

Medium
23

A healthcare analytics team must store patient records in BigQuery and prove to auditors that no one can read or modify the data without an explicit, logged authorization. They want to enforce column-level access so that only the billing team can see insurance_id, while the research team sees only de-identified columns. Which BigQuery feature should they implement?

Easy
24

A company is designing a data lake on Cloud Storage with BigLake tables for unified governance. Which TWO statements about BigLake are correct? (Choose 2.)

Medium
25

A startup runs an operational analytics dashboard on a Cloud SQL for MySQL instance. The dashboard issues many short read queries that are slowing down the write-heavy transactional workload on the primary instance. The team wants to offload reads without changing application write logic and without re-architecting to a different database engine. What should the data engineer configure?

Easy
26

A mobile app needs an offline-first NoSQL database that syncs data across devices when connectivity is available. Which Google Cloud database meets these requirements?

Easy
27

A retail analytics team uses BigQuery to store point-of-sale transactions in a partitioned table. They need to optimize query performance for dashboards that filter on store_id and product_category, and they want to minimize the amount of data scanned. The table is partitioned by transaction_date and has a clustered column of store_id. Queries frequently filter on product_category as well. Which change should the data engineer make to improve performance?

Hard
28

An organization needs to store transactional data for a global e-commerce platform with strong consistency across regions and an SLA of 99.999% availability. The application requires SQL semantics with horizontal scaling. Which Google Cloud database should they choose?

Easy
29

A retail analytics team stores transaction data in a BigQuery table partitioned by day on a DATE column named transaction_date. Analysts repeatedly run queries that filter on transaction_date and also on a high-cardinality STRING column named store_id. These queries scan far more data than expected, and slot usage is high. The team wants to reduce bytes processed without changing query results. What should they do?

Medium
30

A retail company uploads daily sales CSV files to a Cloud Storage bucket. The analytics team wants BigQuery to automatically detect the schema and make new files queryable without running load jobs. The files are added with date-based prefixes, and the team wants to minimize operational overhead. Which BigQuery feature should they use?

Easy
31

A marketing team needs to run ad-hoc SQL queries on terabytes of clickstream data stored in Parquet files in Cloud Storage. They want a serverless solution with no cluster management and the ability to query external data without loading. Which service should they use?

Easy
32

A media company stores final video masters in a Cloud Storage bucket. Regulatory rules require that each object be unalterable for seven years, and the company must be able to prove retention compliance to auditors. Objects are written once and never edited. The data engineer needs the strongest native Cloud Storage control that prevents deletion or overwrite for the required period. What should the engineer configure?

Easy
33

An organization wants to enforce that data in a Cloud Storage bucket cannot be deleted or overwritten for 7 years due to regulatory compliance. Which Cloud Storage feature should they use?

Hard
34

A financial services company must retain trade records for seven years in Cloud Storage. Regulators require that no object can be deleted or overwritten before its retention period expires, even by project owners, and that the policy cannot be removed. The company also needs to prove compliance during audits. Which combination of controls should they implement?

Hard
35

A data engineer is designing a BigQuery table for a fraud detection system. The table will store 500 TB of transaction records and will be queried constantly by filtering on a transaction timestamp column and a customer_id column. The engineer needs to minimize the amount of data scanned by these queries. What should the engineer do?

Medium
36

A data engineer is loading a 2 TB CSV dataset from Cloud Storage into a partitioned BigQuery table. The load job fails with an error indicating too many errors during parsing. The source files have inconsistent column counts and some rows contain embedded newlines. The engineer wants to load the data reliably with minimal preprocessing. What should the engineer do?

Hard
37

A company wants to run complex analytical queries on structured data without managing infrastructure. The data volume is terabytes and queries can take seconds to minutes. Which service is appropriate?

Easy
38

A financial services company stores transaction logs in Cloud Storage. Regulatory requirements mandate that each object be retained for exactly seven years and that no user, including project owners, can delete or modify the objects during that period. The company wants to enforce this with minimal administrative overhead. What should they do?

Hard
39

A data engineer wants to create a data lake on Google Cloud for storing raw streaming data, then transform it into curated and processed zones for analytics. The data is in Avro format and will be queried by BigQuery. Which two services are MOST suitable as the primary storage and query interface?

Medium
40

A data engineer wants to create a BigQuery external table that queries data stored in Parquet format in Cloud Storage without loading the data into BigQuery. Which approach is correct?

Hard
41

A media company stores video files in a Cloud Storage bucket and wants to serve them to users through a custom domain over HTTPS. The files are not sensitive, but the company wants to reduce egress cost and latency for users distributed globally. The company also wants to avoid managing SSL certificates manually. What should the data engineer recommend?

Easy
42

A company has a Cloud SQL for PostgreSQL instance that must be highly available across zones with automatic failover. They also need a read replica for reporting workloads. Which configuration should they use?

Medium
43

A logistics company ingests 5 TB of shipment telemetry per day into a BigQuery table named events_raw. Analysts run exploratory queries that scan the full table, which is now 900 TB, and costs are rising quickly. Most queries filter on the event_date column, which is a DATE field, and on the device_id column, but the table is not partitioned or clustered. The company wants to reduce bytes scanned without changing the query patterns or losing any historical events. Which combination of BigQuery table settings should the data engineer implement?

Medium
44

A team needs to store transactional data for an e-commerce application that requires ACID transactions, automatic backups, and point-in-time recovery. The expected workload is under 10,000 QPS. Which database should they choose?

Easy
45

A company wants to store backups of on-premises databases in Google Cloud for long-term retention. They need WORM (Write Once, Read Many) compliance and object-level retention policies. What should they use?

Medium
46

A company needs to store logs in Cloud Storage for compliance, with a requirement that logs cannot be deleted or overwritten for a period of 7 years. Which Cloud Storage feature should they enable?

Hard
47

Which Google Cloud service is a fully managed relational database for MySQL, PostgreSQL, and SQL Server, offering automatic replication and backups?

Easy
48

A media analytics company ingests 8 TB of new JSON event logs into BigQuery every day and keeps all history for 5 years. Analysts almost never filter on the raw event payload column and only occasionally select it, but they frequently filter on event_date and user_id. Storage cost is the top concern, and query performance on the frequently filtered columns must stay fast. What should the data engineer do?

Medium
49

A company wants to build a data lake on Cloud Storage for raw, curated, and processed data zones. They need to enforce data governance including column-level security and row-level filtering for BigQuery queries. Which solution should they use?

Medium
50

A mobile app needs a NoSQL database that supports offline synchronization when the device goes offline and later reconnects. Which Google Cloud database should be used?

Easy
51

A company uses BigQuery for analytics and needs to ensure that certain columns containing PII are encrypted at query time so that only authorized users can decrypt. What should they use?

Medium
52

A startup is building a mobile app that needs to sync user data across devices in real time. They expect millions of concurrent users and need a NoSQL database with offline support and automatic multi-region replication. Which Google Cloud service meets these requirements?

Easy
53

A team is designing a Spanner database for a global inventory system. They need to optimize query performance for frequently joined tables. Which THREE design decisions help achieve this? (Choose 3.)

Medium
54

A retail analytics team has a 400 TB BigQuery table partitioned by DATE on a transaction_date column. Analysts almost always filter by a single store_id and a date range, and each store has roughly 12 years of history. Queries currently scan the entire partition for the requested dates. The team wants to reduce bytes billed without changing the table name or rewriting the ingestion pipeline. What should they do?

Medium
55

A data engineer is designing a BigQuery table that will store billions of rows of application logs. Queries will almost always filter on a log_date column and then join on a user_id column, and the team wants to minimize both storage cost and bytes scanned. The table is append-only and grows by about 50 GB per day. Which two design choices should the engineer make? (Choose two.)

Hard
56

A company stores highly sensitive financial data in BigQuery. They need to encrypt certain columns (e.g., credit card numbers) with customer-managed encryption keys (CMEK) at the column level. Which BigQuery feature should they use?

Hard
57

A company uses BigQuery for analytics on petabyte-scale data. They want to improve query performance by denormalizing schemas and reducing joins. Which TWO BigQuery features should they use? (Choose 2)

Hard
58

A company is designing a data lake on Cloud Storage with three zones: raw, curated, and processed. They need to enforce data governance by restricting access to each zone using IAM. Which approach should they take?

Medium
59

A company needs to store petabytes of time-series IoT sensor data and query it with single-digit millisecond latency at millions of reads per second. The data has a simple key-value structure with timestamps. Which Google Cloud database is MOST appropriate?

Medium
60

A data engineer needs to design a schema in BigQuery for a dataset that contains customer orders. Each order has a header and multiple line items. Queries frequently need to retrieve the entire order including line items. Which schema design is MOST performant and cost-effective?

Medium
61

A data engineer is designing a BigQuery table for a clickstream dataset with frequent queries aggregating over user sessions. Each user session has multiple events, and the engineer wants to avoid joins for performance. Which schema design pattern should they use?

Medium
62

An application needs to store user profile data in a document database with flexible schema. The data is accessed frequently from a mobile app. Which Google Cloud database is BEST suited?

Easy
63

A mobile app uses Firestore to store user profiles. The app allows offline data creation and syncing when connectivity resumes. Which Firestore feature should the developer enable?

Medium
64

You need to store and query a large dataset of customer profiles. The data is semi-structured and frequently updated. The application requires offline support for mobile users. Which database is MOST appropriate?

Medium
65

A data engineer wants to store archived log files in Cloud Storage with a retention policy that prevents deletion for 5 years. Which feature should they use?

Easy
66

A company stores sensitive data in Cloud Storage and must ensure that data is encrypted at rest with keys that they control and can rotate on demand. They also need to audit key usage and revoke access immediately if a key is compromised. Which Cloud Storage encryption option should they use?

Medium
67

Which Google Cloud database offers global distribution, strong consistency, and a 99.999% SLA?

Easy
68

A data engineer needs to store raw sensor data in Cloud Storage and automatically transition it to a lower-cost storage class after 30 days, then delete it after 365 days. What should they configure?

Medium
69

A healthcare company stores patient imaging metadata in Cloud Firestore in Native mode. The application must run a query that returns all documents where the status field equals 'pending' and the region field equals 'us-east', ordered by created_at descending, and it must paginate through potentially millions of matching documents. Which Firestore capability should the engineer use to satisfy this requirement efficiently?

Hard
70

Which Google Cloud service is a serverless, highly scalable data warehouse for analytical queries, supporting SQL and integration with BI tools?

Easy
71

A company needs to store and analyze large amounts of unstructured data (images, videos) and structured data (CSV logs) in a cost-effective manner. The data should be accessible for analytics with BigQuery. Which two services should they use? (Choose TWO.)

Easy
72

A data engineer wants to automatically move objects from Standard storage class to Nearline after 30 days, and then to Archive after 365 days. Which Cloud Storage feature should they configure?

Easy
73

A data engineer is designing a BigQuery table to store e-commerce order events. Each row contains an order_id (STRING, high cardinality), customer_id (STRING, high cardinality), order_timestamp (TIMESTAMP), and order_amount (NUMERIC). Queries frequently filter on order_id for lookups and also scan by order_timestamp ranges for daily reporting. The engineer wants to minimize bytes scanned by both query patterns. What should the engineer do?

Medium
74

A financial analytics team ingests trade records into BigQuery every minute. Queries filter almost exclusively on trade_date and account_id, and the table grows by roughly 400 GB per day. Analysts report that monthly reports scanning the last 30 days are slow and expensive. The team wants to reduce bytes scanned without changing the ingestion pipeline. What should they do?

Medium
75

A data engineer is designing a Cloud Storage layout for a data lake that will be queried by BigQuery external tables and by Dataproc jobs. The engineer wants to minimize query cost and improve scan performance across both engines. Which two practices should the engineer follow? (Choose two.)

Hard
76

A data engineer needs to choose a storage service for a new application that requires a schemaless document store with automatic multi-region replication, strong consistency for reads, and native mobile SDK support. The application must scale to millions of users without manual sharding. Which Google Cloud service should the engineer select?

Easy
77

A data engineer manages a Cloud SQL for MySQL instance that stores order records. Compliance requires that the data be recoverable to any point in time within the last 30 days, and the team wants the smallest possible recovery window. They also want to avoid managing their own backup rotation scripts. Which configuration should the engineer implement?

Medium
78

A company uses Cloud Storage as a data lake with raw, curated, and processed zones. Data in the raw zone should be automatically moved to a cheaper storage class after 30 days, and deleted after 1 year. What is the most efficient way to implement this?

Medium
79

A team needs to run hybrid transactional/analytical workloads on PostgreSQL-compatible data with low latency. They require high performance on both OLTP and OLAP queries, leveraging a columnar engine. Which Google Cloud service is best suited?

Hard
80

An organization needs to prevent data exfiltration from BigQuery by ensuring all traffic to BigQuery APIs goes through VPC boundaries and is restricted to a specific service perimeter. Which Google Cloud security control should they use?

Medium
81

A company needs to store petabytes of time-series IoT sensor data and query it with single-digit millisecond latency at millions of reads per second. The data has a simple key-value structure with timestamps. Which Google Cloud database is MOST appropriate?

Medium
82

A data team needs to run complex analytical queries on a dataset that is frequently updated with new rows. They want to minimize query costs and avoid scanning old data that is rarely queried. Which BigQuery feature should they use?

Medium
83

A data engineer needs to store quarterly financial data that must remain immutable for 7 years to meet regulatory compliance. The data is accessed infrequently after the first year. Which Cloud Storage feature should be used to enforce immutability?

Medium
84

An e-commerce application uses Cloud SQL (MySQL) for transaction processing. To improve read performance for reporting queries, the team wants to offload read traffic to a separate database instance that stays in sync with the primary. Which Cloud SQL feature should they use?

Medium
85

A healthcare company stores patient records in Cloud SQL for PostgreSQL and needs to run analytical queries on the same data without impacting the production database. The analytics team requires near-real-time replication and wants to use BigQuery for querying. The data must be kept in sync with minimal latency and without writing custom ETL code. Which solution should the data engineer implement?

Medium
86

A team wants to store semi-structured user profile data for a web application. The data is accessed via a REST API and requires security rules to control read/write access. Which database fits best?

Easy
87

A data engineer needs to store JSON documents in a Google Cloud database that must scale horizontally, support automatic multi-region replication, and provide strong consistency for reads. The application performs many small reads and writes keyed by a document ID, and the team wants to avoid managing servers. They also need the ability to run SQL-like queries on the documents. Which Google Cloud service should the engineer choose?

Easy
88

A company needs a fully managed, globally distributed relational database with strong consistency, external consistency, and 99.999% SLA for a financial transaction processing system. Which Google Cloud service should they use?

Easy
89

A data engineer is designing a Cloud Storage bucket for a machine learning training pipeline. The pipeline writes many small Parquet files (about 1 MB each) from a Dataflow job, and a downstream training job reads them sequentially. The engineer wants to minimize the number of storage operations and improve read throughput. The bucket uses Standard storage class and has no lifecycle rules. What should the engineer do?

Hard
90

A data engineer must store JSON documents for a product catalog in Cloud Firestore. Some documents contain nested arrays of product variants that can exceed 1 MiB when serialized. The engineer wants to keep the catalog queryable with strong consistency and low latency. What should the engineer do?

Hard
91

A team wants to use Cloud Storage to build a data lake with separate zones for raw, curated, and processed data. They need to automatically move objects older than 30 days from the raw zone to a cheaper storage class. How can they achieve this?

Medium
92

An organization needs to restrict access to BigQuery and Cloud Storage so that data can only be accessed from within a specific VPC network and cannot be exfiltrated. Which Google Cloud feature should they use?

Medium
93

A Cloud SQL instance for PostgreSQL is experiencing heavy read traffic. The team wants to offload read queries while maintaining data consistency. Which solution meets their needs?

Medium
94

A company is migrating an on-premises PostgreSQL database to Google Cloud. The database runs complex analytical queries mixed with OLTP workloads. They need PostgreSQL compatibility and want to improve analytical query performance without changing the application. Which database should they choose?

Hard
95

A data engineer manages a BigQuery dataset holding a 40 TB partitioned table of point-of-sale transactions. Analysts frequently run queries that filter on a region column and a sale_date column together, but each query scans the full table because the region predicate is not reducing bytes billed. The engineer wants to reduce bytes scanned without changing the analytical queries or the write pipeline. What should the engineer do?

Medium
96

You are designing a row key for Cloud Bigtable to store user activity logs. Each log entry has a timestamp (millisecond precision) and a user ID. There will be millions of writes per second from many users. To avoid hotspotting, which row key design is BEST?

Hard
97

A data engineer is designing a Cloud Bigtable schema for high-volume time-series data. Which TWO practices should they follow to avoid performance issues?

Medium
98

A company wants to use BigQuery for analytics. They need to meet compliance requirements by encrypting data at rest with a key they control. Which TWO actions should they take? (Choose 2.)

Easy
99

A company wants to build a reporting pipeline where data is collected from IoT devices, stored raw in Cloud Storage, and then processed into BigQuery for analytics. They need to ensure data is encrypted at rest using customer-managed keys. Which THREE steps should they take? (Choose 3 correct options)

Medium
100

A healthcare company stores patient records in a Cloud Storage bucket that must remain in a specific region for data residency. The security team requires that all data be encrypted with keys the company controls and can rotate, and that access to the keys be auditable. The company also wants to avoid managing key material on-premises. Which approach should the data engineer choose?

Hard
101

A retail company stores 500 TB of JSON transaction logs in Cloud Storage. Analysts need to run ad hoc SQL over the data, and the company wants to avoid managing a separate cluster while keeping query cost predictable. The logs are already partitioned into date-based prefixes. What should the data engineer do?

Hard
102

Which BigQuery feature allows you to read data directly from Cloud Storage without loading it into BigQuery storage?

Easy
103

A company is migrating an on-premises PostgreSQL database to Google Cloud. They need a fully managed database that is compatible with PostgreSQL and can handle both transactional and analytical workloads with high performance. Which two database services meet these requirements? (Choose TWO.)

Medium
104

A data engineer is designing a BigQuery table to store customer order records. Queries frequently filter by order_date and customer_id, and the dataset grows by about 5 TB per day. The engineer wants to minimize query cost and improve performance for these filtered queries. What should the engineer do?

Medium
105

A company wants to implement a data lake on Google Cloud. They need to store raw, structured data in open formats and allow querying directly from BigQuery without loading. Which TWO services or features should they use? (Choose 2)

Easy
106

A global e-commerce platform requires a relational database that can handle millions of transactions per second across regions with strong consistency and automatic failover. The database must also support SQL joins. Which database should they choose?

Medium
107

A financial services firm stores transaction records in a BigQuery table partitioned by transaction_date. Compliance requires that rows older than seven years be permanently and irreversibly deleted, and the team must prove deletion occurred. The table currently uses the default partition expiration of never. The engineer must implement the retention policy without dropping the whole table. (Choose two.)

Hard
108

A company wants to use BigQuery to query data stored in Parquet files in Cloud Storage without loading the data into BigQuery. Which BigQuery feature should they use?

Easy
109

An e-commerce company uses Cloud Spanner for order processing. They need to query orders by customer ID and retrieve all order items. Which schema design pattern should they use for optimal performance?

Hard

Frequently asked questions

What does the Storing the Data domain cover on the PDE exam?
Match each workload to the correct storage service using consistency, scale, latency, and schema requirements, then apply the right governance controls. The single most important thing: know when Spanner, BigQuery, Cloud Storage, Bigtable, or Firestore is the correct answer, and why the alternatives fail.
How many questions are in this domain?
This page lists all 109 Storing the Data questions in the PDE question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Storing the Data questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
google-pde GOOGLE-PDE pde storing data Practice Questions