Google Cloud · Free Practice Questions · Last reviewed May 2026
30real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
A company needs a fully managed, globally distributed relational database with strong consistency, external consistency, and 99.999% SLA for a financial transaction processing system. Which Google Cloud service should they use?
Firestore
Cloud Spanner
Cloud Spanner is the only Google Cloud relational database offering external consistency globally through TrueTime, satisfying the financial system's strict consistency requirement. Its multi-region configuration delivers the 99.999% SLA, and it is fully managed, unlike self-managed alternatives such as Cloud SQL or Bare Metal Solution.
Bigtable
Cloud SQL
A data engineer needs to store raw sensor data in Cloud Storage and automatically transition it to a lower-cost storage class after 30 days, then delete it after 365 days. What should they configure?
Use Cloud Pub/Sub notifications to trigger a Cloud Function that moves objects.
Use gsutil rewrite command in a cron job.
Configure a lifecycle rule with SetStorageClass to Nearline after 30 days and Delete after 365 days.
A single lifecycle rule with age-based conditions performs both transitions automatically: SetStorageClass to Nearline at 30 days, then Delete at 365 days. This satisfies both the cost-tiering and retention constraints without manual intervention or separate rules.
Set a bucket retention policy with a retention period of 365 days.
An e-commerce company uses Cloud Spanner for order processing. They need to query orders by customer ID and retrieve all order items. Which schema design pattern should they use for optimal performance?
Use interleaved tables where Orders is the parent and OrderItems is an interleaved child table with the same primary key prefix.
Store all data in a single table with nullable columns for order item attributes.
Denormalize by storing order items as a repeated field in the orders table.
Spanner is a relational database; repeated fields are not supported. Denormalization would break relational integrity.
Create two separate tables with a secondary index on customer_id in the orders table and a secondary index on order_id in the order_items table.
A mobile app needs an offline-first NoSQL database that syncs data across devices when connectivity is available. Which Google Cloud database meets these requirements?
Memorystore
Cloud SQL
Bigtable
Firestore
Firestore provides offline persistence through local caching, letting mobile apps read and write without connectivity, then synchronises changes automatically once the device reconnects. Its real-time listeners and multi-device sync satisfy the offline-first NoSQL requirement directly, unlike Cloud SQL or Bigtable, which demand continuous network access.
An organization needs to prevent data exfiltration from BigQuery by ensuring all traffic to BigQuery APIs goes through VPC boundaries and is restricted to a specific service perimeter. Which Google Cloud security control should they use?
Access Transparency
IAM conditions on BigQuery roles
Cloud Armor
VPC Service Controls
VPC Service Controls creates a service perimeter around BigQuery APIs, ensuring traffic stays within VPC boundaries and is restricted to the specified perimeter. This directly satisfies the stem's requirement to prevent exfiltration via API access controls.
A company wants to use BigQuery to query data stored in Parquet files in Cloud Storage without loading the data into BigQuery. Which BigQuery feature should they use?
BigQuery Omni
BigQuery ML
BigQuery external tables
External tables let BigQuery query Parquet files in Cloud Storage directly, reading the data in place without ingestion. This satisfies the no-loading constraint, since no data is copied into BigQuery's native storage and queries run against the external source.
BigQuery BI Engine
Want more Storing the Data practice?
Practice this domain22% of exam · 6 sample questions below
A company uses Dataproc to run daily Spark ML jobs. The jobs run for 2 hours each day. The team wants to reduce costs without changing job characteristics. Which strategy is MOST cost-effective?
Use a single-node cluster to eliminate overhead
Enable high-availability mode to avoid restarts
Use preemptible instances for worker nodes
Preemptible instances cost significantly less than standard Dataproc worker nodes, and Spark can tolerate their loss through retries and task rescheduling. Since the daily jobs are short and their characteristics stay unchanged, using preemptible workers cuts compute spend without altering the job design.
Increase the number of standard workers to finish faster
A data engineer needs to create a BigQuery table that is partitioned by ingestion time and clustered by customer_id and transaction_date. They also want to limit access so that only users from a specific domain can query the table. Which approach should they use?
Create the table with partitioning only, then use a materialized view to restrict access
Create the table without clustering, use row-level security to filter by domain, and grant access to the table
Create the table with partitioning and clustering, then create an authorized view on the table and grant the view access to the domain users
Partitioning by ingestion time and clustering by customer_id and transaction_date optimises the table's physical layout. An authorised view then runs with the owner's permissions, so granting domain users access to the view alone satisfies the domain-restriction requirement without exposing the base table.
Create the table with partitioning and clustering, then grant bigquery.dataViewer to the domain via IAM at the dataset level
A startup needs a fully managed, serverless Spark service to run occasional data processing jobs without managing clusters. They want to pay only for the resources used during job execution. Which Google Cloud service should they use?
Dataproc Serverless
Dataproc Serverless runs Spark workloads without provisioning or managing a cluster, allocating resources only while the job executes and charging for that consumption. This matches the stem's requirements for occasional jobs, no cluster management and pay-per-use billing, unlike a standard Dataproc cluster.
Dataflow
Cloud Data Fusion
Dataproc
A company uses Pub/Sub with push subscriptions to deliver events to a Cloud Run service. Recently, the service has been returning HTTP 429 (Too Many Requests), causing messages to be retried and eventually sent to the dead letter topic. What is the MOST likely cause?
The subscription ackDeadline is set too low, causing messages to be redelivered
The push endpoint is not acknowledging messages quickly enough, causing a backlog
The dead letter topic is misconfigured, causing messages to be sent to it prematurely
The Cloud Run service needs more instances to handle the incoming request rate
Cloud Run scales based on requests; if max instances reached, it returns 429. Increasing instances or adjusting concurrency resolves this.
A data engineer needs to process data in a Dataflow pipeline that reads from a Pub/Sub topic. The pipeline must group events into 5-minute windows and compute the average value per key. Which Beam transform should they use after windowing?
Combine.perKey
Combine.perKey performs per-key aggregation after windowing, computing the average value for each key within each 5-minute window. It satisfies the requirement to group events and average per key, unlike global combines that would merge across keys.
ParDo
GroupByKey
CoGroupByKey
A company uses BigQuery for analytics. They have a table that is queried frequently by date range. To reduce costs, they want to ensure queries only scan the relevant partitions. They also want to improve performance for queries filtering on a specific customer_id. Which table design should they use?
Partition by ingestion time and cluster by customer_id
Use a materialized view that filters by date and customer_id
Cluster by date column and partition by customer_id
Partition by date column and cluster by customer_id
Partitioning on the date column restricts each query to the relevant date range, cutting bytes scanned and cost. Clustering by customer_id then sorts data within each partition, so filters on that column skip irrelevant blocks, satisfying both the cost and performance constraints.
Want more Designing Data Processing Systems practice?
Practice this domain25% of exam · 6 sample questions below
You need to stream real-time user click events from your application into BigQuery for immediate analysis. The events must be available for query within seconds. Which approach is recommended?
Use Pub/Sub to Dataflow to BigQuery with the Storage Write API for high-throughput streaming.
Pub/Sub with Dataflow and the Storage Write API delivers true streaming ingestion, satisfying the seconds-level latency constraint. Dataflow handles windowing and exactly-once processing, while the Storage Write API commits rows directly into BigQuery without the batching delays of load jobs or legacy streaming inserts.
Use Cloud Data Fusion to ingest streaming data from Pub/Sub into BigQuery.
Use Cloud Functions to receive events from Pub/Sub and insert them into BigQuery using the legacy streaming API.
Use Pub/Sub with a BigQuery subscription to directly write events into BigQuery.
Your company is migrating an on-premises Hadoop cluster to Google Cloud. You need to transform large datasets using Spark SQL. Which Google Cloud service should you use?
Dataflow
Dataproc
Dataproc is a managed Spark and Hadoop service, so existing Spark SQL jobs run with minimal refactoring. It satisfies the migration constraint by providing managed cluster provisioning and autoscaling while preserving the Spark SQL transformation logic used on-premises.
BigQuery
Cloud Dataprep
A data engineer needs to transfer 500 TB of on-premises data to Google Cloud Storage. The data is stored on NAS devices and the network bandwidth is limited to 100 Mbps. What is the most cost-effective and timely transfer method?
Use Storage Transfer Service over the internet
Use a VPN connection and rsync
Use gsutil cp in parallel
Use Transfer Appliance
Transfer Appliance bypasses the 100 Mbps network constraint entirely by shipping physical hardware, making 500 TB feasible in days rather than the months a 100 Mbps link would require. For NAS-resident data at this volume, it is also more cost-effective than sustained egress or Interconnect charges.
You are building a Dataflow pipeline in Python that reads messages from Pub/Sub, enriches them with data from a BigQuery table, and writes the results to BigQuery. The enrichment lookup table is large and changes infrequently. Which approach minimizes cost and latency?
Use a CoGroupByKey transform to join the incoming stream with a stream from BigQuery.
Use BigQuery IO to query the table for every incoming message.
Use a side input that reads the BigQuery table periodically and caches it.
A side input reads the slowly changing BigQuery enrichment table periodically and caches it in memory across workers, avoiding a per-element BigQuery lookup. This cuts both query cost and per-message latency, satisfying the large, infrequently changing lookup constraint.
Use a stateful DoFn and store the lookup in state per key.
You are designing a Dataflow pipeline to process streaming data. The pipeline may encounter malformed records. You need to handle these errors without failing the entire pipeline and store the bad records for later analysis. What is the best practice?
Use a dead letter sink to write malformed records to a separate Pub/Sub topic or GCS location.
A dead letter sink isolates malformed records by routing them to a separate Pub/Sub topic or GCS location, so a single bad record cannot fail the whole streaming pipeline. This satisfies the requirement to keep processing valid data while retaining failures for later analysis.
Catch the exception and log it, then continue processing.
Write all records to BigQuery using the Storage Write API and handle errors in the write operation.
Raise an exception in the DoFn to stop the pipeline for manual intervention.
You are migrating an on-premises PostgreSQL database to Cloud SQL. You need to continuously replicate changes to BigQuery for real-time analytics with minimal latency. Which service should you use?
Dataflow with JDBC source
Pub/Sub with a Cloud Function that writes to BigQuery
Storage Transfer Service
Datastream
Datastream provides serverless change data capture, streaming PostgreSQL changes continuously into BigQuery with minimal latency. Unlike batch extracts or scheduled jobs, it replicates ongoing changes in near real time, meeting the continuous replication and low-latency analytics constraint.
Want more Ingesting and Processing the Data practice?
Practice this domain15% of exam · 6 sample questions below
You need to create a Looker model that defines a 'sales' view based on a BigQuery table, with a measure for total revenue. Which LookML object defines the table and dimensions?
explore
view
The view object maps a LookML view to an underlying BigQuery table and declares its dimensions and measures, including total revenue. It satisfies the requirement to define the table structure and the revenue measure in one reusable object.
model
dimension
A company uses Looker Studio to build dashboards from BigQuery data. They notice that queries take several seconds to return. They want to improve performance without changing the schema or adding materialized views. Which option should they use?
Enable BigQuery BI Engine on the relevant project.
BI Engine provides an in-memory analysis layer that caches frequently accessed data, accelerating Looker Studio queries without schema changes or materialised views. It satisfies the stem's constraint of improving dashboard performance while leaving the underlying BigQuery tables untouched, and integrates natively with Looker Studio.
Move the data to Cloud SQL.
Switch to BigQuery Omni for cross-cloud queries.
Use APPROX_COUNT_DISTINCT to speed up distinct counts.
A data scientist is training a binary classification model on an imbalanced dataset (95% negative, 5% positive) using AutoML Tables. Which strategy should they use to handle the class imbalance?
Set the budget to a higher value to allow more training on minority class.
Use SMOTE in a Dataflow pipeline before importing the data to AutoML Tables.
Specify a weight column with higher weights for positive examples in the dataset.
A weight column lets AutoML Tables apply per-row loss multipliers, so positive examples at 5% prevalence can contribute proportionally more during training. This directly addresses the stem's imbalance constraint without resampling, preserving all 95% negative rows while preventing the model from defaulting to the majority class.
Create duplicate copies of the positive class rows to balance the dataset.
You need to split a time-series dataset into training and evaluation sets for a forecasting model. The data is ordered by timestamp. Which splitting technique should you use?
Sequential split where training data precedes evaluation data in time.
A sequential split keeps all training observations earlier in time than evaluation observations, preserving temporal order. This satisfies the forecasting constraint, since random splitting would leak future information into training and produce misleadingly optimistic evaluation metrics.
Use k-fold cross-validation with random folds.
Stratified split based on the target variable.
Random split with 80% training, 20% evaluation.
Which BigQuery SQL function returns the rank of a row within a window, with gaps in the ranking for ties?
RANK()
RANK() assigns positions ordered by the window's ORDER BY, and tied values receive the same rank, leaving gaps before the next rank. It satisfies the explicit gap requirement, unlike DENSE_RANK(), which numbers ties consecutively without gaps.
NTILE()
DENSE_RANK()
ROW_NUMBER()
A company uses Dataplex to manage data quality across multiple BigQuery datasets. They want to define a data quality rule that checks if a column 'email' contains a valid email format. Which Dataplex feature should they use?
Use Cloud DLP to classify and validate emails.
Use the built-in 'email' rule type in Dataplex.
Create a custom Data Quality rule using the 'regex' type.
Dataplex data quality rules support a regex rule type, letting you define a pattern that each value in the email column must match. This satisfies the requirement to validate email format without writing custom SQL assertions.
Create a Dataflow pipeline to validate emails and write results to a separate table.
Want more Preparing and Using Data for Analysis practice?
Practice this domain18% of exam · 6 sample questions below
A data engineer uses Cloud Composer to orchestrate a daily batch pipeline. A downstream task should only start after an upstream BigQuery load job finishes successfully and a specific file appears in Cloud Storage. Which combination of operators should the engineer use in the Airflow DAG?
BigQueryInsertJobOperator with wait_for_downstream=True
BigQueryInsertJobOperator and GCSObjectExistenceSensor with upstream dependency
BigQueryInsertJobOperator runs the load job, while GCSObjectExistenceSensor pokes Cloud Storage until the file appears. Setting the sensor as an upstream dependency forces the downstream task to wait for both the successful load and the file's arrival, satisfying the dual trigger condition.
DataflowPythonOperator and GCSObjectExistenceSensor
BigQueryOperator and FileSensor with downstream dependency
An engineer needs to create a reusable Dataflow pipeline that can be executed with different parameters without modifying code. Which Dataflow feature should they use?
Dataflow Shuffle
Dataflow Flex Templates
Flex Templates package the pipeline as a Docker image with a metadata specification, so the same template runs repeatedly with different runtime parameters and no code changes. This satisfies the stem's reusability and parameterisation constraint.
Dataflow SQL
Dataflow Classic Templates
A data engineer needs to alert when Pub/Sub subscription has messages older than 1 hour. Which Cloud Monitoring metric and filter should they use?
Metric: topic/send_message_operation_count; filter: topic_id
Metric: subscription/ack_message_count; filter: subscription_id
Metric: subscription/num_undelivered_messages; filter: subscription_id
Metric: subscription/oldest_unacked_message_age; filter: subscription_id
The `subscription/oldest_unacked_message_age` metric reports the age of the oldest unacknowledged message, directly satisfying the one-hour staleness threshold. Filtering by `subscription_id` scopes the alert to the specific Pub/Sub subscription, so Cloud Monitoring evaluates only that subscription's backlog age rather than aggregating across the project.
A team wants to enforce data quality rules on BigQuery tables using Dataplex. They need to run column-level checks for null values and row-level checks for value ranges on a schedule. Which Dataplex feature should they use?
Dataplex Data Profiling
BigQuery stored procedures with scheduled queries
Dataplex Data Quality Tasks
Dataplex Data Quality Tasks run scheduled, rule-based checks directly against BigQuery tables, supporting both column-level null validation and row-level range conditions. This satisfies the stem's requirement for automated, recurring enforcement of data quality rules without external tooling, unlike profiling or discovery features that only observe metadata.
Cloud DLP inspection jobs
An organization uses BigQuery on-demand pricing. To control costs, they want to estimate the bytes processed by a query before running it. Which command or method should they use?
Use the bq query --dry_run command
The `bq query --dry_run` flag validates a query and returns the bytes it would process without executing it, so no on-demand charges are incurred. This directly satisfies the requirement to estimate bytes processed beforehand, letting the organisation predict cost before committing to the query.
Use bq ls to list table sizes
Use BigQuery reservations to get cost estimate
Use INFORMATION_SCHEMA.JOBS_BY_PROJECT to view past costs
A streaming Dataflow pipeline needs to be updated without draining the existing pipeline. Which update strategy should be used?
Drain the pipeline first, then start a new one
Replace the job with a new job using the same pipeline name
A replacement job cannot be updated in place, so the pipeline is relaunched under the same name, letting the new job start while the old one is cancelled. This avoids the drain step that would stop ingestion, satisfying the no-drain constraint.
Use a different pipeline name and cancel the old one
Stop the job, update the code, and restart
Want more Maintaining and Automating Data Workloads practice?
Practice this domainThe PDE exam has 60 questions and must be completed in 120 minutes. The passing score is 720/1000.
Scenario-based questions covering exam objectives with detailed answer explanations.
The exam covers 5 domains: Storing the Data, Designing Data Processing Systems, Ingesting and Processing the Data, Preparing and Using Data for Analysis, Maintaining and Automating Data Workloads. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official Google Cloud PDE exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.