Courseiva

PDE · domain

Designing Data Processing Systems

This domain covers designing data processing systems on Google Cloud, including ingestion, transformation, storage, and serving. It tests selecting appropriate services like Pub/Sub, Dataflow, Dataproc, and BigQuery, and configuring them for scalability, cost, and reliability. Questions present scenarios requiring architectural decisions based on requirements.

100 questions29 easy45 medium26 hard

Focused practice

Practice Designing Data Processing Systems questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Designing Data Processing Systems

Select appropriate Google Cloud services for data ingestion, processing, and storage based on requirements. The most important thing is to match service capabilities to the use case, ensuring scalability, cost-efficiency, and reliability.

Choosing between Pub/Sub, Dataflow, Dataproc, and BigQuery for ingestion and processing.

Configuring Dataproc clusters with preemptible VMs and handling failures.

Optimizing BigQuery tables with partitioning and clustering for performance and cost.

Applying IAM roles to control access to BigQuery datasets and other resources.

Watch out for

Common Designing Data Processing Systems exam traps

  • ▸Assuming Dataflow is always best for batch; Dataproc may be more cost-effective for existing Spark/Hadoop workloads.
  • ▸Forgetting that preemptible VMs can be reclaimed, so jobs must tolerate failures or use regular VMs for critical tasks.
  • ▸Overlooking that BigQuery partitioning and clustering are key for query performance and cost reduction on large tables.

Question index

All Designing Data Processing Systems questions (100)

Click any question to see the full explanation, or start a practice session above.

1

A logistics company wants to optimize their delivery routes using historical GPS data. The data is stored in BigQuery and is updated daily. They need to run a complex machine learning model that requires iterative processing over the entire dataset using Apache Spark. The model training takes several hours and must be run weekly. They want to minimize cost and operational overhead. Which approach should they take?

Hard
2

A company uses Pub/Sub with push subscriptions to deliver events to a Cloud Run service. Recently, the service has been returning HTTP 429 (Too Many Requests), causing messages to be retried and eventually sent to the dead letter topic. What is the MOST likely cause?

Hard
3

An online gaming company runs a Dataflow streaming pipeline that aggregates player actions into per-session metrics. Sessions are defined by a gap duration of 30 minutes of inactivity, and the pipeline must emit final session results even when a session's events arrive out of order by up to 10 minutes. Late data beyond that window can be dropped. Which combination of Beam concepts should the pipeline use?

Medium
4

An organization is using BigQuery for analytics. They have a table that is 500 GB and is frequently queried by 'date' and 'region'. They want to optimize query performance and reduce costs. Which TWO actions should they take?

Medium
5

A company runs Apache Spark jobs on Dataproc. They want to reduce costs by using preemptible instances for worker nodes. The jobs are fault-tolerant and can handle occasional node loss. However, the cluster must remain available for interactive querying during business hours. Which Dataproc cluster configuration meets these requirements?

Medium
6

A company uses Pub/Sub to ingest events from multiple sources. They need to ensure that messages from a specific source are processed in order (per source partition). They also need to deduplicate messages. Which TWO features should they use?

Medium
7

A retail analytics team must move 30 TB of Parquet files from an on-premises Hadoop cluster into BigQuery once, then run standard SQL dashboards. The transfer window is 48 hours and the source cluster has limited outbound bandwidth. Which approach should the data engineer choose?

Easy
8

A healthcare analytics group must build a pipeline that ingests HL7 messages from an on-premises interface engine, must retain raw messages for seven years for compliance, and must expose de-identified aggregates to analysts. The security team requires that protected health information never be written to a dataset analysts can query, and that all data be encrypted with keys the organization manages and can revoke. Which two design choices satisfy these requirements? (Choose two.)

Hard
9

A media company processes video metadata using a Dataflow pipeline. They need to join two streaming sources: user activity (Pub/Sub) and video catalog updates (Pub/Sub). Which THREE transforms should be used in the pipeline?

Hard
10

You are designing a BigQuery data lake for a healthcare organization. The data includes patient records that must be access-controlled at the row level. Which TWO features should you use to meet this requirement?

Medium
11

You have a BigQuery table that is partitioned by ingestion time and clustered on user_id. The table stores event logs and is queried frequently by user_id to analyze user behavior over the last 30 days. Queries are still scanning too many partitions. Which optimization should you apply first?

Medium
12

A company uses Cloud Pub/Sub to ingest events from multiple sources. They need to guarantee that each event is processed exactly once by downstream consumers. However, Pub/Sub guarantees at-least-once delivery. Which additional steps should they implement to achieve exactly-once processing?

Hard
13

A media company streams playback telemetry through Pub/Sub into a Dataflow pipeline that writes to BigQuery. During prime-time peaks, the pipeline's BigQuery write step shows growing latency and the job repeatedly reports that it is backing off on insert retries. The team wants to reduce write pressure without changing the downstream table schema or losing exactly-once semantics. What should they do?

Hard
14

A data engineer is designing a pipeline that reads from Cloud Pub/Sub, aggregates events into 5-minute windows, and writes the results to BigQuery. The engineer wants to ensure that late-arriving data (up to 2 minutes late) is included in the correct window. Which Dataflow feature should they configure?

Medium
15

You are designing a data pipeline that processes streaming events with late-arriving data (up to 2 hours late). The pipeline must compute hourly aggregations and emit results as soon as possible, but must also accurately update results when late data arrives. You want to minimize overall processing cost. Which Dataflow windowing and trigger configuration should you use?

Hard
16

Your company is building a real-time anomaly detection system for financial transactions. The system must process streams of transactions and flag anomalies within seconds. The volume is moderate (5000 transactions per second). You want a fully managed solution that integrates with BigQuery for historical analysis. Which service should you use for stream processing?

Easy
17

A financial services firm is designing a Dataflow pipeline that reads from Pub/Sub and writes enriched transactions to BigQuery. The pipeline must guarantee exactly-once processing semantics for the BigQuery sink, even during pipeline updates and worker restarts. The team plans to use the Apache Beam Java SDK with the BigQueryIO connector. Which combination of configurations should they use?

Hard
18

A media company is designing a data processing system on Google Cloud to analyze video streaming logs. The logs are generated continuously and stored in Cloud Storage. The company wants to use Cloud Dataflow to process these logs, but they need to ensure the pipeline can handle late-arriving data and provide accurate results for both real-time dashboards and historical analysis. Which two features of Cloud Dataflow should they use? (Choose two.)

Medium
19

A healthcare analytics team needs to run a series of SQL transformations on data stored in BigQuery. The transformations must run on a schedule, and the team wants to minimize operational overhead by using a fully managed service that integrates with BigQuery and Cloud Logging. They also need to parameterize the SQL queries with runtime values such as the current date. Which Google Cloud service should they use?

Easy
20

A team wants to use Cloud Pub/Sub Lite for a high-throughput, low-cost messaging system. They need exactly-once delivery to subscribers. What should they know about Pub/Sub Lite's delivery guarantees?

Medium
21

A company uses Dataproc to run daily Spark ML jobs. The jobs run for 2 hours each day. The team wants to reduce costs without changing job characteristics. Which strategy is MOST cost-effective?

Medium
22

Your company uses Pub/Sub to ingest clickstream data. Messages must be processed in order for the same user_id. How should you configure the Pub/Sub subscription to guarantee ordering?

Medium
23

A financial services company must analyze transaction data that includes customers' full names, account numbers, and home addresses. Regulations require that this personally identifiable information (PII) never be stored in raw form in their BigQuery analytics warehouse. The data engineering team plans to use Cloud Dataflow to read from a Pub/Sub topic and write to BigQuery. Which approach best satisfies the regulatory requirement while keeping the pipeline simple?

Medium
24

A company needs to process streaming sensor data from millions of devices with sub-second latency, apply transformations, and write results to BigQuery for real-time dashboards. The data volume varies, and they want to avoid managing servers. Which service should they use?

Medium
25

A company is using Pub/Sub to ingest clickstream events. They need to ensure that events are delivered to a subscriber at least once, but duplicates can be tolerated. They also need to filter events by type before processing. Which subscription configuration should be used?

Medium
26

A company has a BigQuery dataset containing sensitive customer data. They want to share a subset of this data with external partners, ensuring that partners can only see specific columns and rows. Which BigQuery feature should they use?

Easy
27

An organization is implementing a data lake on Google Cloud using Cloud Storage. They need to process both batch and streaming data with a unified pipeline. The team has experience with Apache Beam. Which architecture should they use to minimize operational overhead?

Hard
28

A data engineer needs to create a BigQuery table that is partitioned by ingestion time and clustered by customer_id and transaction_date. They also want to limit access so that only users from a specific domain can query the table. Which approach should they use?

Medium
29

A company needs to process high-throughput streaming data with low latency. They are considering Cloud Pub/Sub for ingestion and Cloud Dataflow for processing. However, they are concerned about cost. Which alternative to Cloud Pub/Sub would reduce costs while still meeting the throughput requirements?

Medium
30

A financial services firm runs a Dataflow batch pipeline that joins a 2 TB transaction dataset with a 40 GB customer reference dataset. The reference data changes only once per day and is currently read from a BigQuery table with a side-input transform on every element. Job cost is dominated by repeated BigQuery reads, and the pipeline occasionally hits quota errors. The team wants to minimize cost and quota pressure while keeping the daily refresh. What should the data engineer change?

Hard
31

A data pipeline uses Cloud Data Fusion to perform ETL jobs. The pipeline reads from BigQuery, transforms data using Wrangler, and writes to Cloud Storage. The team notices that the pipeline runs slower than expected. They suspect the Data Fusion instance is under-provisioned. Which action should be taken to improve performance?

Hard
32

You need to store petabytes of data in a data warehouse that supports ANSI SQL, automatic scaling, and real-time analytics. The data is primarily used for ad-hoc queries and business intelligence. Which Google Cloud service should you use?

Easy
33

You are designing a BigQuery data warehouse for a retail company. Queries frequently filter on order_date and customer_id. To optimize query performance and cost, which table design should you use?

Medium
34

A small startup wants to run a nightly batch job that transforms a 2 GB CSV file in Cloud Storage and writes the result back as Parquet. The team has no cluster administration experience, wants per-job pricing rather than an always-on cluster, and needs the transformation to run in under thirty minutes. Which Google Cloud approach should they choose?

Easy
35

You are building a real-time fraud detection system using Dataflow. Events from Pub/Sub need to be grouped by user_id within a 5-minute window to detect suspicious patterns. Some events may be delayed by up to 2 minutes. How should you configure the window and trigger to balance accuracy and latency?

Hard
36

A data pipeline ingests streaming events into Pub/Sub and needs to join them with a slowly updating reference table (few thousand rows) from a Cloud Storage CSV file. The pipeline runs on Dataflow with Apache Beam. Which approach is most cost-effective and operationally simple?

Medium
37

A startup is building a data lake on Google Cloud. They need to store raw JSON, CSV, and Parquet files from various sources. The files will be accessed by multiple analytics tools, including BigQuery and Dataproc. The startup wants a cost-effective, durable, and highly available storage solution that integrates natively with these services. Which Google Cloud service should they use?

Easy
38

A logistics company uses Cloud Dataflow to process a continuous stream of GPS events from delivery trucks. The pipeline must compute the distance traveled per truck per hour and write the results to BigQuery. The events are keyed by truck ID, and the pipeline uses windowing with a one-hour fixed window. The team notices that some trucks report events with timestamps that are several minutes late due to network delays. They want to ensure that late events are still included in the correct window and that the results are emitted only after a reasonable wait. What should they configure?

Hard
39

A retail company runs a batch pipeline in Cloud Dataflow that reads from Cloud Storage and writes to BigQuery. The pipeline uses a GroupByKey operation on a key that is heavily skewed: one customer ID accounts for 40% of all transactions. This causes a single worker to process a massive amount of data, and the job takes hours longer than expected. The team wants to reduce the skew without changing the pipeline's logic or output. What should they do?

Hard
40

A financial services firm ingests trade events into Cloud Storage and must load them into BigQuery. Compliance requires that each event be processed exactly once and that the load be idempotent across retries, even if a Dataflow job restarts mid-batch. The destination table must also be queryable immediately after each successful load. Which loading approach best satisfies these requirements?

Hard
41

A retail company ingests point-of-sale clickstream events into Cloud Pub/Sub at roughly 200,000 messages per second during flash sales. Analysts need near-real-time dashboards that aggregate revenue by product category over sliding 5-minute windows, with results visible in BigQuery within 30 seconds of the event. The pipeline must handle occasional bursts up to 3x the normal rate without dropping messages, and the team wants to minimize operational overhead. Which design should the data engineer use?

Hard
42

A startup needs a fully managed, serverless Spark service to run occasional data processing jobs without managing clusters. They want to pay only for the resources used during job execution. Which Google Cloud service should they use?

Easy
43

A financial services firm is designing a data processing system on Google Cloud that must ingest change data capture (CDC) streams from an on-premises PostgreSQL database into BigQuery with sub-minute latency, preserve the ordering of changes per primary key, and apply updates and deletes so that BigQuery reflects the current state of each row. The source database cannot be modified to add triggers. Which two design elements should you include? (Choose two.)

Hard
44

A retail company is designing a Dataflow pipeline to process point-of-sale transactions from Cloud Pub/Sub and write to BigQuery. The pipeline must handle late-arriving data up to 24 hours and ensure that all data is written to BigQuery exactly once, even in the event of worker failures. Which two features should the engineer implement to meet these requirements? (Choose two.)

Medium
45

Your company is migrating an on-premises Apache Hadoop cluster to Google Cloud. The cluster runs Hive for SQL-like queries and stores data in HDFS. You want a managed service that minimizes operational overhead while supporting existing Hive scripts. Which Google Cloud service should you choose?

Medium
46

A Dataproc cluster uses preemptible worker nodes to reduce costs. The cluster runs a long-running Spark job that occasionally experiences worker failures. How should the job be configured to handle preemptible worker failures gracefully?

Hard
47

You are designing a Dataflow pipeline for processing real-time clickstream data. The pipeline must group events into 30-second windows and handle late data up to 5 minutes. You want to output partial results every 10 seconds for low-latency monitoring. Which THREE configurations should you use? (Choose three.)

Medium
48

A company needs a messaging service for event-driven applications that require low cost for high-throughput, but can tolerate occasional message loss. Which Pub/Sub product should they choose?

Easy
49

A company wants to use Dataproc Metastore to manage metadata for their Spark jobs. Which TWO benefits does Dataproc Metastore provide?

Medium
50

A data engineer needs to design a streaming pipeline that ingests events from multiple sources, enriches them with a lookup table stored in BigQuery (updated every hour), and writes the results to a BigQuery table for real-time dashboards. The pipeline must handle late-arriving data up to 1 hour. Which Dataflow feature should be configured to manage late data?

Medium
51

An engineer needs to create a Pub/Sub subscription that sends messages to an HTTPS endpoint. The endpoint must be able to acknowledge messages individually. Which type of subscription should they use?

Easy
52

You are designing a data processing architecture on Google Cloud. You need to ingest data from multiple sources, including streaming events and batch files, and process them to produce a unified dataset for analytics. The solution must support both real-time and historical processing with the same codebase, and be able to handle late-arriving data. Which two Google Cloud services should you use together to achieve this? (Choose two.)

Hard
53

A logistics company uses Cloud Pub/Sub to ingest shipment tracking events. They want to archive all events to Cloud Storage for long-term retention and also process them in real time with Dataflow. The events are published to a single topic. Which design should the data engineer use to ensure both archiving and real-time processing without data loss?

Medium
54

A company uses Cloud Dataproc to run Spark ML training jobs. They want to persist the trained models and metadata in a Hive-compatible metastore. Which Dataproc feature should they use?

Medium
55

A company is using Cloud Storage to store raw logs. They want to use Cloud Data Fusion to transform and load the data into BigQuery on a daily schedule. The transformations are complex and involve joining multiple datasets. What is the most efficient way to run these pipelines?

Medium
56

A data engineer needs to process data in a Dataflow pipeline that reads from a Pub/Sub topic. The pipeline must group events into 5-minute windows and compute the average value per key. Which Beam transform should they use after windowing?

Easy
57

A startup is building a data lake on Google Cloud. They need to store raw JSON logs in a cost-effective manner and later query them using SQL with minimal transformation. The logs are infrequently accessed but must be retained for 7 years for compliance. Which storage solution should they use?

Easy
58

Which Google Cloud service provides a visual interface for building ETL pipelines using a drag-and-drop design and includes pre-built transforms from a marketplace?

Easy
59

A retail company runs a nightly batch pipeline that loads point-of-sale transactions into BigQuery. The pipeline uses a Cloud Composer (Apache Airflow) DAG with a BigQueryInsertJobOperator task that runs a SQL MERGE statement to upsert yesterday's sales into a large fact_sales table partitioned by transaction_date. The data engineering team notices that the MERGE task occasionally fails with a 'Resources exceeded during query execution' error when the source staging table contains more than 50 million rows. They need to redesign the DAG to reliably load large daily volumes without changing the final table schema or downstream dashboards. What should they do?

Medium
60

A financial services company needs to process credit card transactions in real time to detect fraudulent patterns. The pipeline must handle late-arriving data (up to 2 hours) and produce accurate results. They want to use a unified programming model that works for both batch and streaming. Which Google Cloud service should they use?

Hard
61

You need to allow a data analyst to run queries on a BigQuery dataset but prevent them from modifying the data or deleting the dataset. Which IAM role should you grant?

Easy
62

You are designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline must handle late-arriving data (up to 1 hour) and group events into 10-minute windows. Which configuration is correct?

Hard
63

A company wants to use Cloud Data Fusion for ETL pipelines. They need to integrate with custom transformations not available in the marketplace. What should they do?

Medium
64

A data engineer needs to process streaming data from thousands of IoT devices and generate real-time dashboards. The data volume is low but requires exactly-once processing semantics. Which Google Cloud service combination should they use?

Easy
65

Your company runs a Dataflow streaming pipeline that processes user activity from Pub/Sub and writes aggregated results to BigQuery. Lately, the pipeline is experiencing high latency and backlog growth during peak hours. You need to troubleshoot and improve performance. Which THREE actions should you take? (Choose 3.)

Hard
66

You need to create a BigQuery table that stores customer transaction data. The table will be queried frequently by a customer_id column to retrieve recent transactions (last 30 days). Which table design optimizes query performance and cost?

Medium
67

You are designing a streaming pipeline that needs to handle sudden spikes in traffic without losing data. The pipeline uses Pub/Sub and Dataflow. Which configuration ensures data is not lost if Dataflow falls behind?

Medium
68

A financial services firm needs to design a batch processing system on Google Cloud to analyze large volumes of historical transaction data stored in Cloud Storage. The data is in Parquet format and must be processed using Apache Spark. The firm wants to minimize operational overhead and only pay for the resources used during job execution. Which Google Cloud service should they use?

Easy
69

Your team is migrating an on-premises Apache Hadoop cluster to Google Cloud. The cluster runs MapReduce jobs that read and write data to HDFS. You want to minimize code changes and operational overhead. Which Google Cloud service should you use to run these jobs?

Medium
70

A healthcare company is designing a system to ingest HL7 messages from multiple hospitals into Google Cloud. The messages must be processed in near real-time to extract patient vitals and trigger alerts if thresholds are exceeded. The system must guarantee that no messages are lost and that processing is exactly-once. Which combination of Google Cloud services should they use?

Hard
71

Your data engineering team needs to process a continuous stream of clickstream events from a website and update a real-time dashboard showing user activity over the last hour. The pipeline should have minimal operational overhead and support exactly-once processing semantics. Which Google Cloud service should you use?

Easy
72

Which Google Cloud service provides a serverless Spark environment where you can run Spark jobs without provisioning or managing a cluster?

Easy
73

A company is evaluating BigQuery for a data warehouse migration. They have a mix of reporting queries and ad-hoc analytical queries. They want to control query costs and prevent runaway queries. Which THREE strategies should they implement?

Medium
74

A company wants to use BigQuery materialized views to accelerate queries on a table that is updated every hour. Which statement about materialized views is true?

Medium
75

You need to choose a messaging service for a real-time streaming application that requires low cost and can tolerate occasional message loss. Which service is MOST suitable?

Easy
76

You are migrating on-premises Hadoop jobs to Google Cloud. The existing jobs use Spark for ETL and Hive for querying. You want to minimize changes to the existing code and maintain the ability to use Hive queries with the same metastore across multiple clusters. Which service combination should you use?

Easy
77

You need to design a data processing system that ingests streaming data from thousands of IoT devices. The data must be processed in real-time to calculate average temperature per device over 1-minute intervals, and the results should be stored in BigQuery for analysis. You want a serverless solution with minimal management. Which combination of Google Cloud services should you use?

Easy
78

A company uses Cloud Pub/Sub for event ingestion. They want to ensure that if a subscriber fails to process a message after 5 attempts, the message is sent to a dead letter topic for analysis. Which TWO configurations are needed?

Medium
79

An organization runs periodic Apache Spark jobs on Dataproc to process data from Cloud Storage. They want to reduce costs by using preemptible instances for worker nodes. What is a key consideration when using preemptible instances in Dataproc?

Medium
80

You need to analyse streaming data from thousands of IoT devices, each sending temperature readings every second. You want to calculate the average temperature per device over the last 5 minutes, updating every minute. Which windowing strategy should you use in Dataflow?

Medium
81

A financial services firm runs a batch risk calculation on Dataproc. The job reads from Cloud Storage, processes data in memory, and writes results to BigQuery. The job must complete within a 2-hour window each night, and the cluster must be shut down automatically after completion to minimize cost. You want to orchestrate this with minimal operational overhead. What should you do?

Hard
82

A company uses Dataproc Serverless for Spark batch jobs. They notice that some jobs are failing due to out-of-memory (OOM) errors. Which configuration parameter should they adjust to allocate more memory per executor?

Medium
83

A startup wants to analyze user clickstream data stored in Cloud Storage in Parquet format. They need to run ad-hoc SQL queries without managing any servers and want to pay only for the queries they run. Which Google Cloud service should they use?

Easy
84

A developer wants to create a BigQuery table that automatically expires data older than 30 days to reduce storage costs. Which table design feature should be used?

Easy
85

A company uses BigQuery for analytics. They have a table that is queried frequently by date range. To reduce costs, they want to ensure queries only scan the relevant partitions. They also want to improve performance for queries filtering on a specific customer_id. Which table design should they use?

Medium
86

A retail company ingests point-of-sale events from thousands of stores into Cloud Pub/Sub. They need to process these events in a streaming Dataflow pipeline that enriches each event with store metadata from a slowly changing BigQuery table. The enrichment table is updated only once per day. The pipeline must minimize latency and avoid querying BigQuery for every event. Which approach should they use?

Medium
87

A data team is building a near-real-time dashboard that displays aggregated metrics from Kafka topics. They want to use Pub/Sub as a managed messaging service and Dataflow for stream processing. They need to ingest data from Kafka into Pub/Sub with minimal custom code. Which THREE Google Cloud services should they use together? (Choose three.)

Medium
88

A startup wants to build a data lake on Google Cloud to store raw JSON, CSV, and Parquet files from various sources. They need a storage solution that is highly durable, globally accessible, and integrates natively with BigQuery and Dataproc. They want to minimize management overhead. Which Google Cloud service should they use?

Easy
89

Which Google Cloud service provides a fully managed, serverless Spark environment without requiring cluster provisioning?

Easy
90

You need to process a large Spark ML training job on a Dataproc cluster. The job is fault-tolerant and can handle occasional node failures. To reduce costs, which type of worker nodes should you use?

Easy
91

You are designing a Dataflow pipeline that joins two unbounded PCollections from different sources. Which transform should you use?

Medium
92

Which BigQuery feature allows you to share query results with specific users without giving them direct access to the underlying tables?

Easy
93

A logistics company collects GPS pings from delivery vehicles into Pub/Sub and needs to compute the distance traveled per vehicle per hour. The data volume is high and bursty, and the company wants a managed service that automatically scales the number of workers based on load while allowing custom windowing and stateful processing. Which service should they use?

Medium
94

A financial services firm must process payment events in strict order per account and cannot tolerate duplicates. The events arrive in Pub/Sub and must be written to BigQuery. The engineering team is designing the pipeline and wants to guarantee that each account's events are applied in the order they were published. Which approach should they take?

Hard
95

A media company needs to process a large number of small JSON files stored in Cloud Storage. They want to use a serverless, SQL-based approach to transform and aggregate the data without managing infrastructure. Which Google Cloud service should they use?

Easy
96

Your company ingests millions of events per second into a Pub/Sub topic. The downstream consumer must process events with minimal latency and high throughput. However, the consumer occasionally falls behind during traffic spikes, and you need to ensure no data loss while minimizing costs. Which subscription type and configuration should you choose?

Medium
97

Your team is using Cloud Dataprep to clean and transform a dataset. Which TWO features of Cloud Dataprep help you understand data quality issues before running the pipeline? (Choose 2.)

Easy
98

Your team is designing a data processing system that ingests JSON messages from millions of IoT devices. The ingestion rate is highly variable, with spikes up to 500,000 messages per second. You need a fully managed, serverless messaging service that can buffer messages and decouple producers from consumers. Which Google Cloud service should you choose?

Medium
99

A data engineer needs to run an existing Spark job on Google Cloud with minimal code changes. The job requires Hive metastore access. Which Dataproc feature should they use to provide a managed Hive metastore?

Medium
100

A logistics company ingests GPS telemetry from delivery vehicles into Pub/Sub. They need to process the stream in Dataflow to calculate real-time estimated arrival times (ETAs). The pipeline must handle late-arriving data up to 2 hours and must emit results every 5 minutes. The team wants to use Apache Beam's windowing and triggering. Which windowing strategy and trigger should they use to meet these requirements?

Hard

Frequently asked questions

What does the Designing Data Processing Systems domain cover on the PDE exam?
Select appropriate Google Cloud services for data ingestion, processing, and storage based on requirements. The most important thing is to match service capabilities to the use case, ensuring scalability, cost-efficiency, and reliability.
How many questions are in this domain?
This page lists all 100 Designing Data Processing Systems questions in the PDE question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Designing Data Processing Systems questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
google-pde GOOGLE-PDE pde designing data systems Practice Questions