Courseiva

CCNA Designing Data Processing Systems Questions

75 of 100 questions · Page 1/2 · Designing Data Processing Systems · Answers revealed

1
MCQhard

A logistics company wants to optimize their delivery routes using historical GPS data. The data is stored in BigQuery and is updated daily. They need to run a complex machine learning model that requires iterative processing over the entire dataset using Apache Spark. The model training takes several hours and must be run weekly. They want to minimize cost and operational overhead. Which approach should they take?

A.Run the Spark job on a long-running Dataproc cluster that is always available.
B.Use Dataproc Serverless for Spark with the BigQuery connector to read data directly from BigQuery.
C.Use BigQuery ML to train the model directly in BigQuery without Spark.
D.Export the data to Cloud Storage and run a Dataproc cluster with autoscaling, then delete the cluster after training.
AnswerB

Dataproc Serverless for Spark can read data directly from BigQuery using the BigQuery connector, eliminating the need to export data. It is a serverless solution, so it automatically provisions and scales resources, and charges only for the duration of the job. This minimizes both cost and operational overhead, as there is no cluster to manage. It supports complex Spark MLlib models and iterative processing, making it ideal for this scenario.

Why this answer

Dataproc Serverless for Spark with the BigQuery connector allows the company to run Spark jobs directly on BigQuery data without exporting it. It is serverless, so it minimizes operational overhead and costs by charging only for job execution. It supports complex Spark MLlib models and iterative processing, making it the best fit for weekly training on historical GPS data.

This approach aligns with the goals of minimizing cost and operational overhead while leveraging Spark's capabilities.

Exam trap

The trap here is assuming that data must be exported from BigQuery to Cloud Storage for Spark processing, overlooking the direct BigQuery connector available in Dataproc Serverless.

2
MCQhard

A company uses Pub/Sub with push subscriptions to deliver events to a Cloud Run service. Recently, the service has been returning HTTP 429 (Too Many Requests), causing messages to be retried and eventually sent to the dead letter topic. What is the MOST likely cause?

A.The subscription ackDeadline is set too low, causing messages to be redelivered
B.The push endpoint is not acknowledging messages quickly enough, causing a backlog
C.The dead letter topic is misconfigured, causing messages to be sent to it prematurely
D.The Cloud Run service needs more instances to handle the incoming request rate
AnswerD

Cloud Run scales based on requests; if max instances reached, it returns 429. Increasing instances or adjusting concurrency resolves this.

Why this answer

Push subscriptions can be rate limited by the receiving service. Increasing ackDeadline gives more time but doesn't reduce rate. Using pull subscriptions shifts the rate control to the subscriber.

Adjusting max delivery attempts only affects how many retries before dead letter, not rate limiting.

3
MCQmedium

An online gaming company runs a Dataflow streaming pipeline that aggregates player actions into per-session metrics. Sessions are defined by a gap duration of 30 minutes of inactivity, and the pipeline must emit final session results even when a session's events arrive out of order by up to 10 minutes. Late data beyond that window can be dropped. Which combination of Beam concepts should the pipeline use?

A.Fixed windows of 30 minutes with the default trigger and allowed lateness of zero.
B.Sliding windows of 30 minutes with a one-minute period and an early trigger firing every minute.
C.Session windows with a 30-minute gap, an allowed lateness of 10 minutes, and a trigger that fires after the watermark passes the end of the window.
D.Global window with a repeated element-count trigger and a discard accumulation mode.
AnswerC

Session windows merge based on inactivity gaps, so a 30-minute gap duration models the session definition exactly. Setting allowed lateness to 10 minutes keeps window state alive so out-of-order events within that bound are still incorporated, and firing on the watermark produces final results once the window closes, which matches the emit-final-results requirement.

Why this answer

Session windows are the only window type that models inactivity gaps, and combining a 30-minute gap with 10 minutes of allowed lateness keeps state open long enough for out-of-order events while dropping anything later. A watermark-based trigger emits the final result once the window is considered complete.

Exam trap

The trap here is treating the 30-minute gap as a fixed window length, when in session windows the gap is an inactivity threshold that merges windows rather than a fixed time bucket.

4
Multi-Selectmedium

An organization is using BigQuery for analytics. They have a table that is 500 GB and is frequently queried by 'date' and 'region'. They want to optimize query performance and reduce costs. Which TWO actions should they take?

Select 2 answers
A.Use an authorized view
B.Use a wildcard table
C.Use materialized views
D.Cluster the table by region
E.Partition the table by date
AnswersD, E

Clustering by region sorts data within partitions by that column, so filters on region read only relevant blocks rather than scanning everything. This satisfies the stem's region-filtering access pattern, reducing bytes billed and improving performance.

Why this answer

Option E is correct because partitioning the table by date means BigQuery only scans the partitions that match the query's date filter, drastically reducing bytes processed and cost for a 500 GB table frequently filtered by date. Option D is correct because clustering the table by region physically sorts data within each partition by region, so queries filtering on region benefit from block pruning and scan less data. Together, partitioning by date and clustering by region match the two most common query predicates and are the standard BigQuery optimization pattern.

Option A (authorized view) only controls access to data and does not improve query performance or reduce scan costs. Option B (wildcard table) is used to query multiple similarly named tables with a UNION-like syntax and does not optimize a single table's scans. Option C (materialized views) can help some workloads but is not the primary optimization for a frequently filtered base table and adds storage and refresh overhead, so it is not one of the two best actions here.

Exam trap

PDE often tests the confusion between access control features (authorized views) and performance features (partitioning, clustering); candidates pick materialized views or wildcard tables without recognizing that partitioning and clustering directly address the query patterns.

5
MCQmedium

A company runs Apache Spark jobs on Dataproc. They want to reduce costs by using preemptible instances for worker nodes. The jobs are fault-tolerant and can handle occasional node loss. However, the cluster must remain available for interactive querying during business hours. Which Dataproc cluster configuration meets these requirements?

A.Use a single-node cluster that automatically scales with preemptible instances
B.Use a standard cluster with preemptible instances as secondary workers
C.Use standard cluster with master and worker nodes as preemptible instances
D.Use a high-availability cluster with preemptible instances for primary workers
AnswerB

Secondary workers (preemptible workers) are ideal for fault-tolerant batch jobs. They do not store HDFS data, so losing them does not affect data durability. The cluster remains available because primary workers and master nodes are regular instances.

Why this answer

Dataproc supports preemptible VMs as secondary workers, which are added to a cluster alongside non-preemptible primary workers. Secondary workers handle the bulk of shuffle and task execution, so if a preemptible node is reclaimed, the job can recover while the primary workers and master keep the cluster alive and available for interactive queries. This gives the cost benefit of preemptible pricing without risking cluster or master availability.

Exam trap

The trap here is conflating 'fault-tolerant job' with 'fault-tolerant cluster' — candidates assume preemptible nodes can be used anywhere because the job can retry, forgetting that master and primary worker roles must remain non-preemptible to preserve cluster availability.

How to eliminate wrong answers

Option A is wrong because a single-node cluster has no separate master/worker separation and cannot use preemptible secondary workers — losing the node would take down the entire cluster, breaking interactive availability. Option C is wrong because making master and worker nodes preemptible risks losing the master (which terminates the cluster) and the primary workers, violating the availability requirement. Option D is wrong because high-availability mode protects the master via multiple masters, but using preemptible instances for primary workers still risks losing the workers that must remain stable for interactive querying.

6
Multi-Selectmedium

A company uses Pub/Sub to ingest events from multiple sources. They need to ensure that messages from a specific source are processed in order (per source partition). They also need to deduplicate messages. Which TWO features should they use?

Select 2 answers
A.Set a message schema to enforce ordering
B.Use a dead letter topic to handle out-of-order messages
C.Enable exactly-once delivery on the subscription
D.Use a pull subscription with a large ack deadline
E.Enable message ordering by setting an ordering key
AnswersC, E

Exactly-once delivery on the subscription removes duplicate messages by tracking acknowledgement state, satisfying the deduplication requirement. However, it does not guarantee ordering; that needs the ordering key or message ordering property set on the publisher side, so this feature alone covers only one of the two stated constraints.

Why this answer

Option E is correct because Pub/Sub message ordering is enabled by setting an ordering key on messages; messages sharing the same ordering key are delivered to subscribers in the order they were published, which satisfies the per-source-partition ordering requirement. Option C is correct because enabling exactly-once delivery on the subscription ensures that messages are not redelivered after successful acknowledgment, providing the required deduplication behavior. Option A is incorrect because schemas validate message format and structure, not ordering.

Option B is incorrect because a dead letter topic only captures messages that cannot be successfully processed after retries, and does not reorder or deduplicate messages. Option D is incorrect because a large ack deadline only extends the processing window to reduce premature redelivery; it does not guarantee ordering or exactly-once deduplication.

Exam trap

The trap is confusing schema enforcement or dead letter topics with ordering and deduplication — candidates pick schema (A) thinking it 'enforces order' or DLQ (B) thinking it 'handles out-of-order messages,' when only ordering keys and exactly-once delivery address the actual requirements.

7
MCQeasy

A retail analytics team must move 30 TB of Parquet files from an on-premises Hadoop cluster into BigQuery once, then run standard SQL dashboards. The transfer window is 48 hours and the source cluster has limited outbound bandwidth. Which approach should the data engineer choose?

A.Export the Parquet files to encrypted portable drives and use the Transfer Appliance to ship them to a Google ingest location, then load the objects into BigQuery.
B.Use the Storage Transfer Service with a transfer job configured against the on-premises source, then load the resulting Cloud Storage objects into BigQuery using a load job.
C.Create a Dataproc cluster in the same region and use a Spark job to read from the on-premises Hadoop cluster over the network and write directly into BigQuery.
D.Run gcloud storage cp from a Compute Engine VM to upload the files into a BigQuery-managed bucket, then query them with an external table.
AnswerA

Transfer Appliance is purpose-built for bulk offline migration when network bandwidth makes online transfer impractical. Shipping encrypted appliances avoids the limited outbound link entirely and comfortably moves 30 TB within the window. After Google uploads the appliance contents to Cloud Storage, a BigQuery load job reads the Parquet objects into native tables, after which dashboards query managed storage with full columnar performance.

Why this answer

When the data volume is large and the source's outbound bandwidth is the constraint, an offline transfer avoids the network bottleneck. Transfer Appliance ships encrypted storage to Google, where the data lands in Cloud Storage, and a BigQuery load job then ingests the Parquet files into native managed tables. Dashboards subsequently query columnar managed storage rather than external objects, so both the migration deadline and query performance requirements are met.

Exam trap

The trap here is defaulting to a network-based copy or a Spark job when the stated constraint is limited outbound bandwidth within a fixed time window, which points to offline transfer.

8
Multi-Selecthard

A healthcare analytics group must build a pipeline that ingests HL7 messages from an on-premises interface engine, must retain raw messages for seven years for compliance, and must expose de-identified aggregates to analysts. The security team requires that protected health information never be written to a dataset analysts can query, and that all data be encrypted with keys the organization manages and can revoke. Which two design choices satisfy these requirements? (Choose two.)

Select 2 answers
A.Load the raw messages directly into a BigQuery table that analysts can query, and rely on column-level security to mask the PHI columns.
B.De-identify the messages in Dataflow, write only the de-identified aggregates into a separate BigQuery dataset secured with CMEK and policy tags, and grant analysts access only to that dataset.
C.Land raw HL7 messages in a Cloud Storage bucket configured with a customer-managed encryption key (CMEK) and a retention lock, and restrict access with IAM to the ingestion service account only.
D.Store the raw messages in BigQuery and use authorized views so analysts query a view that filters out PHI columns instead of the base table.
E.Encrypt the raw messages with a customer-supplied key in the on-premises interface engine and upload only the ciphertext to Cloud Storage, decrypting in Dataflow for de-identification and discarding the key afterward.
AnswersB, C

Performing de-identification in Dataflow before the data reaches BigQuery ensures PHI never enters an analyst-queryable dataset, satisfying the hard boundary. Writing aggregates to a distinct dataset protected by CMEK gives the organization revocable key control, and policy tags add fine-grained access restrictions. Analysts are granted access only to this curated dataset, keeping the raw zone entirely separate.

Why this answer

The requirements demand a hard separation between raw PHI and anything analysts can query, plus organization-managed revocable encryption and seven-year retention. Landing raw messages in Cloud Storage with CMEK and a retention lock, accessible only to the ingestion service account, satisfies retention and key control. De-identifying in Dataflow and publishing only aggregates to a separate CMEK-protected BigQuery dataset keeps PHI out of analyst-facing storage entirely, so both boundaries hold independently.

Exam trap

The trap here is treating masking or authorized views over a raw BigQuery table as equivalent to never writing PHI into an analyst-queryable dataset, when the requirement demands physical separation, not query-time filtering.

9
Multi-Selecthard

A media company processes video metadata using a Dataflow pipeline. They need to join two streaming sources: user activity (Pub/Sub) and video catalog updates (Pub/Sub). Which THREE transforms should be used in the pipeline?

Select 3 answers
A.Flatten to combine the two PCollections
B.ParDo to process each element individually
C.CoGroupByKey to join the two PCollections on a common key (e.g., video_id)
D.Window both PCollections into a common window (e.g., fixed 1-minute)
E.GroupByKey on each PCollection separately before joining
AnswersB, C, D

ParDo is needed to process each element and extract the common key (e.g., video_id) before joining.

Why this answer

To join two streaming Pub/Sub sources in Dataflow, you need to: (1) Use ParDo to extract the key from each element (e.g., video_id), (2) Window both PCollections into a common window (e.g., fixed 1-minute) to align the data, and (3) Use CoGroupByKey to join on the common key. GroupByKey separately is not required because CoGroupByKey internally groups by key. Flatten is used to combine PCollections of the same type, which is not applicable here.

Exam trap

Candidates often think GroupByKey is needed before CoGroupByKey, but CoGroupByKey does the grouping internally.

10
Multi-Selectmedium

You are designing a BigQuery data lake for a healthcare organization. The data includes patient records that must be access-controlled at the row level. Which TWO features should you use to meet this requirement?

Select 2 answers
A.Row-level security using row access policies
B.Authorised views with row filters
C.Dataset-level IAM roles
D.Clustering on patient_id
E.Materialised views
AnswersA, B

Row access policies enforce row-level filtering directly in BigQuery by attaching a filter expression to a table, so each user sees only permitted patient records. This satisfies the stem's row-level access-control constraint natively, without duplicating data or maintaining separate views per role.

Why this answer

Row-level security using row access policies (A) is correct because BigQuery row access policies let you define a filter predicate (e.g., based on a patient_id or facility column) that restricts which rows each user or group can see, enforcing row-level access control directly on the table. Authorised views with row filters (B) are also correct because you can create a view that applies a WHERE clause to expose only permitted rows, then authorise that view on the dataset so users query the view without direct table access, achieving row-level control. Dataset-level IAM roles (C) only grant access at the dataset, table, or view level, not per row, so they cannot satisfy row-level requirements.

Clustering on patient_id (D) is a physical storage optimization that improves query performance and reduces scanned bytes but provides no access control. Materialised views (E) cache precomputed query results for performance and do not enforce row-level security.

Exam trap

PDE often tests the distinction between row-level and dataset-level access controls, causing candidates to overlook authorised views as a valid row-level security method.

11
MCQmedium

You have a BigQuery table that is partitioned by ingestion time and clustered on user_id. The table stores event logs and is queried frequently by user_id to analyze user behavior over the last 30 days. Queries are still scanning too many partitions. Which optimization should you apply first?

A.Create a materialized view that pre-aggregates data by user_id and date
B.Remove partitioning and rely solely on clustering
C.Change the partition column to a DATE column based on event_timestamp and keep clustering on user_id
D.Add clustering on a second column like event_type
AnswerC

Ingestion-time partitioning groups rows by load time, so a 30-day event query scans every partition loaded in that period regardless of event dates. Partitioning on a DATE derived from event_timestamp enables partition pruning to only the relevant event dates, while clustering on user_id still accelerates per-user filtering.

Why this answer

Ingestion-time partitioning creates partitions based on when data arrives, not the event timestamp, so a query filtering on the last 30 days of events may still scan partitions that contain older events ingested recently. Changing the partition column to a DATE derived from event_timestamp aligns partition pruning with the query filter, and keeping clustering on user_id preserves the benefit for user-based filtering. This is the highest-impact first optimization.

Exam trap

PDE often tests whether candidates recognize that ingestion-time partitioning does not align with event-time filters — many assume any partitioning enables pruning, missing that the partition column must match the query predicate.

How to eliminate wrong answers

Option A is wrong because a materialized view adds cost and complexity and does not fix the root cause — the partition scheme still mismatches the query predicate, so the view may still scan too much. Option B is wrong because removing partitioning eliminates partition pruning entirely, which would make scans worse, not better. Option D is wrong because adding a second clustering column can help some queries but does not address the fundamental mismatch between ingestion-time partitions and event-time filters, so it is not the first fix.

12
MCQhard

A company uses Cloud Pub/Sub to ingest events from multiple sources. They need to guarantee that each event is processed exactly once by downstream consumers. However, Pub/Sub guarantees at-least-once delivery. Which additional steps should they implement to achieve exactly-once processing?

A.Set the subscription's acknowledgment deadline to 0.
B.Enable message deduplication on the subscription.
C.Store each message's unique ID in a database and ignore duplicates.
D.Use a dead letter topic to capture duplicates.
AnswerC

Pub/Sub delivers at-least-once, so the same message can arrive repeatedly. Persisting each message's unique ID and discarding already-seen IDs gives idempotent deduplication at the consumer, converting at-least-once delivery into effectively exactly-once processing as the stem requires.

Why this answer

Pub/Sub's at-least-once delivery means a message can be redelivered if the ack isn't received in time or if the subscriber crashes after processing but before acking. To achieve exactly-once processing, the consumer must implement idempotency by tracking processed message IDs and discarding duplicates. Storing each message's unique ID in a database and ignoring duplicates ensures that even if the same message is delivered multiple times, it is only processed once.

This is the standard pattern for exactly-once semantics on top of at-least-once infrastructure.

Exam trap

PDE often tests the misconception that enabling a Pub/Sub feature (like deduplication or dead letter topics) can achieve exactly-once processing, when in fact exactly-once requires application-level idempotency.

How to eliminate wrong answers

Option A is wrong because setting the acknowledgment deadline to 0 would cause messages to be redelivered immediately, making duplicates more likely, not less. Option B is wrong because Pub/Sub does not have a subscription-level deduplication feature; deduplication is only available on the publish side when using ordering keys or message IDs, and it does not guarantee exactly-once processing on the subscriber side. Option D is wrong because a dead letter topic captures messages that cannot be processed after a certain number of attempts; it does not prevent duplicate processing and is unrelated to exactly-once semantics.

13
MCQhard

A media company streams playback telemetry through Pub/Sub into a Dataflow pipeline that writes to BigQuery. During prime-time peaks, the pipeline's BigQuery write step shows growing latency and the job repeatedly reports that it is backing off on insert retries. The team wants to reduce write pressure without changing the downstream table schema or losing exactly-once semantics. What should they do?

A.Enable Dataflow Shuffle and set the pipeline's disk size larger so the write stage can buffer more rows before retrying.
B.Change the pipeline to use a BigQuery load job triggered every 15 minutes by writing micro-batches to Cloud Storage first.
C.Switch the sink to the BigQuery Storage Write API with the exactly-once stream type and group records into batches before writing.
D.Increase the number of BigQuery Streaming Inserts API calls by adding more Dataflow workers so each record is inserted individually.
AnswerC

The Storage Write API supports an exactly-once stream type that uses offsets to deduplicate on retry, preserving exactly-once semantics while allowing high-throughput batched appends. Batching amortizes request overhead and dramatically lowers the per-row cost and quota pressure that caused the backoff. Because it writes to the same table schema, no downstream change is required, and it is the recommended high-volume ingestion path for Dataflow into BigQuery.

Why this answer

Backoff on insert retries signals that the write path is saturating BigQuery's streaming ingestion limits. The Storage Write API's exactly-once stream type preserves exactly-once semantics through offset-based deduplication, and batching rows into larger appends cuts request volume and quota pressure. This keeps the existing table schema and the streaming model intact while removing the write bottleneck that was throttling the pipeline during peak load.

Exam trap

The trap here is scaling out workers or tuning Dataflow shuffle and disk settings to fix a problem that actually originates in the BigQuery ingestion API's quotas and retry behavior.

14
MCQmedium

A data engineer is designing a pipeline that reads from Cloud Pub/Sub, aggregates events into 5-minute windows, and writes the results to BigQuery. The engineer wants to ensure that late-arriving data (up to 2 minutes late) is included in the correct window. Which Dataflow feature should they configure?

A.Use a sliding window of 5 minutes with 2-minute slide
B.Set the window duration to 7 minutes to account for lateness
C.Set the allowed lateness to 2 minutes with a trigger that fires on late data
D.Use a global window and watermark
AnswerC

Allowed lateness keeps each 5-minute window's state alive for 2 minutes past the watermark, and a late-data trigger emits updated results when those events arrive. This ensures events up to 2 minutes late land in the correct window rather than being dropped.

Why this answer

Dataflow's allowed lateness feature (set to 2 minutes) ensures that late-arriving data within that threshold is still assigned to the correct 5-minute window. Combined with a trigger that fires on late data, the pipeline can emit updated results for the window after the watermark passes, which is exactly what the engineer needs to handle late-arriving events up to 2 minutes late.

Exam trap

The trap here is that candidates confuse window duration adjustments (Option B) or sliding windows (Option A) with the proper late-data handling mechanism, not realizing that allowed lateness and triggers are the correct Dataflow primitives for including late-arriving data in the correct event-time window.

How to eliminate wrong answers

Option A is wrong because a sliding window of 5 minutes with a 2-minute slide creates overlapping windows that emit results every 2 minutes, not a single 5-minute window with late data handling; it would double-count events and not solve the late-arrival problem. Option B is wrong because setting the window duration to 7 minutes does not account for lateness—it simply shifts the window boundaries, causing data to be assigned to a different time range, which is incorrect for the intended 5-minute aggregation. Option D is wrong because a global window and watermark would aggregate all data into a single unbounded window, losing the per-5-minute grouping required by the pipeline.

15
MCQhard

You are designing a data pipeline that processes streaming events with late-arriving data (up to 2 hours late). The pipeline must compute hourly aggregations and emit results as soon as possible, but must also accurately update results when late data arrives. You want to minimize overall processing cost. Which Dataflow windowing and trigger configuration should you use?

A.Fixed windows of 1 hour with allowed lateness of 2 hours and trigger every 5 minutes (early) and on watermark (late) with accumulating fired panes
B.Global window with triggers every 5 minutes
C.Sliding windows of 1 hour with 30-minute offset
D.Session windows with 10-minute gap duration
AnswerA

Fixed one-hour windows with two hours of allowed lateness retain late events, while early periodic triggers plus a watermark trigger emit results promptly and accumulating panes revise prior output. This satisfies both low-latency emission and accurate late-data correction without over-provisioning resources.

Why this answer

Fixed windows of 1 hour with allowed lateness of 2 hours and triggers every 5 minutes (early) and on watermark (late) with accumulating fired panes is the correct configuration. This setup computes hourly aggregations, emits early results every 5 minutes, and updates results when late data arrives within the 2-hour allowed lateness. Accumulating panes ensure that late data updates the previous results.

This minimizes cost by using fixed windows and appropriate triggers.

Exam trap

The trap is misunderstanding triggers and allowed lateness; candidates may choose global windows or sliding windows without considering the need for hourly aggregations and late data handling.

How to eliminate wrong answers

Option B is wrong because a global window with triggers every 5 minutes does not provide hourly aggregations; it would aggregate all data into one window, which is not the requirement. Option C is wrong because sliding windows of 1 hour with a 30-minute offset would produce overlapping windows and may not align with hourly aggregations; also, it does not specify allowed lateness or triggers for late data. Option D is wrong because session windows with a 10-minute gap are for grouping events based on activity gaps, not for fixed hourly aggregations, and they do not handle late data as specified.

16
MCQeasy

Your company is building a real-time anomaly detection system for financial transactions. The system must process streams of transactions and flag anomalies within seconds. The volume is moderate (5000 transactions per second). You want a fully managed solution that integrates with BigQuery for historical analysis. Which service should you use for stream processing?

A.Cloud Dataflow
B.Cloud Pub/Sub with push subscriptions
C.Cloud Dataproc with Spark Streaming
D.Cloud Data Fusion
AnswerA

Cloud Dataflow provides fully managed, autoscaling stream processing with exactly-once semantics and low-latency windowing, satisfying the sub-second anomaly flagging requirement. Its native BigQuery sink connector streams results directly into BigQuery for historical analysis, meeting the integration constraint without custom code or cluster management.

Why this answer

Cloud Dataflow is a fully managed, serverless stream and batch processing service based on Apache Beam. It natively supports real-time streaming with low latency (sub-second to seconds), integrates seamlessly with Pub/Sub for ingestion and BigQuery for output, and auto-scales to handle moderate throughput like 5000 transactions per second. Its managed nature eliminates cluster operations, making it ideal for this use case.

Exam trap

PDE often tests the distinction between messaging (Pub/Sub) and stream processing (Dataflow), and candidates may confuse fully managed services with self-managed ones like Dataproc.

How to eliminate wrong answers

Option B is wrong because Cloud Pub/Sub with push subscriptions is a messaging service, not a stream processing engine; it can deliver messages but cannot perform anomaly detection logic. Option C is wrong because Cloud Dataproc with Spark Streaming requires managing a cluster and is not fully managed, adding operational overhead. Option D is wrong because Cloud Data Fusion is a managed ETL tool for batch and micro-batch data integration, not designed for low-latency stream processing.

17
MCQhard

A financial services firm is designing a Dataflow pipeline that reads from Pub/Sub and writes enriched transactions to BigQuery. The pipeline must guarantee exactly-once processing semantics for the BigQuery sink, even during pipeline updates and worker restarts. The team plans to use the Apache Beam Java SDK with the BigQueryIO connector. Which combination of configurations should they use?

A.Use BigQueryIO.write() with withMethod(STREAMING_INSERTS) and set withFailedInsertRetryPolicy to RETRY_NEVER. This ensures that failed inserts are not retried, preventing duplicates.
B.Use BigQueryIO.write() with withMethod(FILE_LOADS), withTriggeringFrequency set to a fixed interval, and enable withAutoSharding. Rely on the default insertion semantics.
C.Use BigQueryIO.write() with withMethod(STREAMING_INSERTS) and provide a deterministic insertId derived from the transaction ID. Set withFailedInsertRetryPolicy to RETRY_ALWAYS.
D.Use BigQueryIO.write() with withMethod(FILE_LOADS) and set withWriteDisposition(WRITE_TRUNCATE) for each window. This ensures that each window's data replaces the previous contents of the table.
AnswerC

STREAMING_INSERTS with a deterministic insertId enables BigQuery to deduplicate retried inserts, providing exactly-once semantics. RETRY_ALWAYS ensures that transient failures are retried, and because the insertId is stable, duplicates are suppressed. This combination is the recommended way to achieve exactly-once writes to BigQuery from a streaming Dataflow pipeline, even across worker restarts.

Why this answer

Exactly-once semantics for BigQuery writes in Dataflow require deduplication on the BigQuery side. With STREAMING_INSERTS, BigQuery can use the insertId to discard duplicate rows within a short time window. A deterministic insertId derived from the transaction ID ensures that retries after worker restarts or pipeline updates do not create duplicates.

RETRY_ALWAYS preserves data by retrying transient errors, while deduplication prevents duplicates.

Exam trap

The trap here is believing that FILE_LOADS or RETRY_NEVER alone can guarantee exactly-once semantics, when the real mechanism is a deterministic insertId with STREAMING_INSERTS and retries enabled.

18
Multi-Selectmedium

A media company is designing a data processing system on Google Cloud to analyze video streaming logs. The logs are generated continuously and stored in Cloud Storage. The company wants to use Cloud Dataflow to process these logs, but they need to ensure the pipeline can handle late-arriving data and provide accurate results for both real-time dashboards and historical analysis. Which two features of Cloud Dataflow should they use? (Choose two.)

Select 2 answers
A.Side inputs for enrichment
B.Triggers to emit early results
C.Exactly-once processing
D.Session windows
E.Windowing with allowed lateness
AnswersB, E

Triggers in Dataflow control when to emit aggregated results for a window. Using triggers, such as early triggers or late triggers, allows the pipeline to produce speculative results before the window closes and then update them as late data arrives. This enables real-time dashboards to show near-instantaneous insights while still refining results for historical accuracy. Triggers are a key mechanism for balancing latency and completeness in streaming pipelines.

Why this answer

To handle late-arriving data and provide accurate results for both real-time dashboards and historical analysis, the pipeline should use windowing with allowed lateness and triggers to emit early results. Windowing with allowed lateness ensures that late data is incorporated into the correct window, while triggers allow early results to be emitted for real-time dashboards and updated as more data arrives. Together, they balance latency and completeness.

Exam trap

The trap here is focusing on data integrity features like exactly-once processing instead of the timing features that actually manage late data and result emission.

19
MCQeasy

A healthcare analytics team needs to run a series of SQL transformations on data stored in BigQuery. The transformations must run on a schedule, and the team wants to minimize operational overhead by using a fully managed service that integrates with BigQuery and Cloud Logging. They also need to parameterize the SQL queries with runtime values such as the current date. Which Google Cloud service should they use?

A.Cloud Composer with a DAG that uses BigQueryInsertJobOperator to run the SQL queries.
B.BigQuery scheduled queries, using the @run_date parameter for runtime values.
C.Dataflow with a pipeline that reads from BigQuery, applies SQL transformations using Beam SQL, and writes back to BigQuery.
D.Cloud Scheduler triggering a Cloud Function that calls the BigQuery API to execute the SQL.
AnswerB

BigQuery scheduled queries are a fully managed feature that runs SQL on a schedule, supports parameterization with @run_date and other system variables, and integrates with Cloud Logging for monitoring. It requires no infrastructure management, making it ideal for scheduled SQL transformations. This directly meets the team's requirements with minimal operational overhead.

Why this answer

BigQuery scheduled queries are a native, fully managed feature that executes SQL on a defined schedule. They support parameterization with system variables like @run_date, which allows dynamic date-based filtering. Integration with Cloud Logging provides visibility.

This eliminates the need to manage infrastructure or write code, perfectly matching the requirement for a low-overhead, scheduled SQL transformation solution.

Exam trap

The trap here is overcomplicating a simple scheduled SQL task by choosing a general-purpose orchestrator or compute service instead of the native BigQuery scheduling feature.

20
MCQmedium

A team wants to use Cloud Pub/Sub Lite for a high-throughput, low-cost messaging system. They need exactly-once delivery to subscribers. What should they know about Pub/Sub Lite's delivery guarantees?

A.Pub/Sub Lite provides at-least-once delivery, same as standard Pub/Sub.
B.Pub/Sub Lite provides exactly-once delivery when using push subscriptions.
C.Pub/Sub Lite provides exactly-once delivery when using pull subscriptions.
D.Pub/Sub Lite supports exactly-once delivery by default.
AnswerA

Pub/Sub Lite guarantees at-least-once delivery, so duplicate messages can reach subscribers; exactly-once is unavailable. This satisfies the stem's constraint by clarifying that the team cannot rely on deduplication, and must implement idempotent processing or their own deduplication logic to achieve exactly-once semantics.

Why this answer

Pub/Sub Lite offers at-least-once delivery like standard Pub/Sub; exactly-once is not guaranteed.

21
MCQmedium

A company uses Dataproc to run daily Spark ML jobs. The jobs run for 2 hours each day. The team wants to reduce costs without changing job characteristics. Which strategy is MOST cost-effective?

A.Use a single-node cluster to eliminate overhead
B.Enable high-availability mode to avoid restarts
C.Use preemptible instances for worker nodes
D.Increase the number of standard workers to finish faster
AnswerC

Preemptible instances cost significantly less than standard Dataproc worker nodes, and Spark can tolerate their loss through retries and task rescheduling. Since the daily jobs are short and their characteristics stay unchanged, using preemptible workers cuts compute spend without altering the job design.

Why this answer

Preemptible (Spot) VMs cost up to 80% less than standard VMs and are ideal for fault-tolerant, batch-oriented workloads like Spark ML jobs that can tolerate occasional preemption. Since the jobs run only 2 hours daily and Dataproc automatically handles node replacement when preemptible instances are reclaimed, using preemptible workers delivers the largest cost reduction without changing job characteristics.

Exam trap

The trap is assuming 'reduce cost' means 'reduce cluster size' — candidates pick single-node or fewer workers, but the exam expects you to recognize that preemptible/Spot instances reduce cost per unit while preserving job characteristics.

How to eliminate wrong answers

Option A is wrong because a single-node cluster eliminates worker parallelism entirely, which would drastically slow or break distributed Spark ML jobs — it reduces cost by removing capacity, not by optimizing it. Option B is wrong because high-availability mode adds a second master node, increasing cost rather than reducing it, and it addresses master reliability, not worker cost. Option D is wrong because adding more standard workers increases cost linearly and only helps if the job is worker-bound; it does not reduce per-unit cost.

22
MCQmedium

Your company uses Pub/Sub to ingest clickstream data. Messages must be processed in order for the same user_id. How should you configure the Pub/Sub subscription to guarantee ordering?

A.Use a pull subscription with enable_message_ordering=true
B.Use a pull subscription with exactly-once delivery enabled
C.Use a push subscription with acknowledgement deadline set to 600 seconds
D.Use a push subscription with a dead letter topic
AnswerA

Ordering in Pub/Sub is enforced per ordering key, so setting enable_message_ordering=true on a pull subscription makes the service deliver messages sharing the same user_id sequentially. Without this flag, Pub/Sub provides at-least-once delivery with no ordering guarantee, so clickstream events for one user could arrive out of sequence.

Why this answer

Pub/Sub message ordering is enabled per-subscription by setting enable_message_ordering=true, and ordering keys (such as user_id) must be set on published messages. With ordering enabled, Pub/Sub delivers messages with the same ordering key in publish order to a single subscriber, satisfying the per-user_id ordering requirement.

Exam trap

The trap is confusing exactly-once delivery with ordering — candidates see 'guarantee' and pick exactly-once, but ordering and exactly-once are orthogonal Pub/Sub features that must be enabled separately.

How to eliminate wrong answers

Option B is wrong because exactly-once delivery guarantees no duplicate delivery but does not guarantee ordering — these are independent features, and exactly-once alone will not preserve per-key order. Option C is wrong because extending the ack deadline to 600 seconds only affects redelivery timing; it has no bearing on message ordering. Option D is wrong because a dead letter topic captures messages that fail processing after retries, which is a reliability pattern, not an ordering mechanism.

23
MCQmedium

A financial services company must analyze transaction data that includes customers' full names, account numbers, and home addresses. Regulations require that this personally identifiable information (PII) never be stored in raw form in their BigQuery analytics warehouse. The data engineering team plans to use Cloud Dataflow to read from a Pub/Sub topic and write to BigQuery. Which approach best satisfies the regulatory requirement while keeping the pipeline simple?

A.Apply a ParDo transform in Dataflow that uses the Cloud Data Loss Prevention (DLP) API to de-identify sensitive fields before writing to BigQuery.
B.Use BigQuery column-level security to restrict access to the PII columns, and grant permissions only to authorized analysts.
C.Enable BigQuery default encryption with a customer-managed encryption key (CMEK) so that the PII columns are encrypted at rest.
D.Load the raw data into a BigQuery table, then run a scheduled query that replaces PII columns with hashed values.
AnswerA

The Cloud DLP API is purpose-built for discovering and de-identifying sensitive data such as names, account numbers, and addresses. Calling it from a Dataflow ParDo transform lets you tokenize, mask, or encrypt PII in-flight, so only de-identified records reach BigQuery. This directly satisfies the regulation without adding storage layers or external systems.

Why this answer

De-identifying data in-flight with Cloud DLP inside a Dataflow pipeline prevents raw PII from ever reaching BigQuery storage. The DLP API can inspect and transform fields like names, account numbers, and addresses using techniques such as masking or tokenization. This meets the regulatory requirement at the earliest point, while keeping the pipeline serverless and managed.

Exam trap

The trap here is assuming that access controls or encryption at rest satisfy a requirement that PII must never be stored in raw form, when only de-identification before persistence actually removes the raw values.

24
MCQmedium

A company needs to process streaming sensor data from millions of devices with sub-second latency, apply transformations, and write results to BigQuery for real-time dashboards. The data volume varies, and they want to avoid managing servers. Which service should they use?

A.Cloud Data Fusion
B.Dataflow
C.Dataproc
D.Dataprep
AnswerB

Dataflow provides serverless, autoscaling stream processing with exactly-once semantics, satisfying the sub-second latency and variable-volume constraints. Its Apache Beam pipeline can transform sensor data and write directly into BigQuery via the built-in BigQueryIO connector, with no servers to manage.

Why this answer

Dataflow is Google Cloud's fully managed, serverless data processing service built on Apache Beam, designed for both batch and streaming pipelines with exactly-once processing semantics. It natively supports streaming ingestion from Pub/Sub or Kafka, applies transformations via Beam's windowing/triggers, and can write directly to BigQuery using the BigQueryIO connector with streaming inserts or Storage Write API. Its autoscaling (Horizontal Autoscaling and Streaming Engine) handles variable data volumes without server management, making it the only option that meets sub-second latency, serverless, and BigQuery streaming requirements simultaneously.

Exam trap

PDE often tests the distinction between serverless, streaming-native Dataflow and cluster-based or batch-oriented tools like Dataproc, Data Fusion, and Dataprep, causing candidates to pick Dataproc for 'streaming' because Spark Streaming sounds similar, or Data Fusion because it's an ETL tool.

How to eliminate wrong answers

Option A is wrong because Cloud Data Fusion is a GUI-based, code-free ETL orchestration tool built on CDAP; it is primarily batch-oriented, runs on a Dataproc cluster (so you still manage underlying infrastructure), and is not designed for sub-second streaming latency. Option C is wrong because Dataproc is a managed Hadoop/Spark cluster service where you still provision and size clusters (even with autoscaling, it's not serverless in the same sense) and Spark Streaming micro-batches typically have higher latency than Dataflow's per-record streaming. Option D is wrong because Dataprep by Trifacta is a data wrangling tool for interactive, visual preparation of batch data (often used with Dataflow under the hood for execution), not a streaming pipeline engine, and it lacks native sub-second streaming semantics.

25
MCQmedium

A company is using Pub/Sub to ingest clickstream events. They need to ensure that events are delivered to a subscriber at least once, but duplicates can be tolerated. They also need to filter events by type before processing. Which subscription configuration should be used?

A.Pull subscription with exactly-once delivery enabled
B.Push subscription with no filter
C.Pull subscription with a filter on event type attribute
D.Push subscription with a dead letter topic
AnswerC

A pull subscription supports attribute filters, so Pub/Sub discards non-matching messages server-side before delivery, satisfying the filtering requirement. Pull delivery retains at-least-once semantics, meaning occasional duplicates may reach the subscriber, which the stem explicitly tolerates.

Why this answer

A pull subscription with a filter on the event type attribute satisfies both requirements: Pub/Sub's at-least-once delivery is the default for pull subscriptions, so duplicates are tolerated, and message filtering via subscription filter expressions (on attributes like event type) prevents unwanted events from being delivered. This is the only option that combines at-least-once semantics with server-side filtering.

Exam trap

PDE often tests the confusion between at-least-once and exactly-once delivery semantics, tricking candidates into selecting exactly-once when the scenario explicitly says duplicates are acceptable.

How to eliminate wrong answers

Option A is wrong because exactly-once delivery is the opposite of at-least-once — it eliminates duplicates rather than tolerating them, and it is not the requirement here. Option B is wrong because a push subscription with no filter provides no filtering capability, so all events would be delivered regardless of type. Option D is wrong because a dead letter topic handles undeliverable messages, not event-type filtering, and push subscriptions do not address the filtering requirement.

26
MCQeasy

A company has a BigQuery dataset containing sensitive customer data. They want to share a subset of this data with external partners, ensuring that partners can only see specific columns and rows. Which BigQuery feature should they use?

A.Materialized views
B.Authorized views
C.Clustered tables
D.Dataset-level access controls
AnswerB

Authorized views grant external partners query access to a view while restricting them to the columns and rows the view exposes, without granting underlying table access. This directly satisfies the requirement to share only a subset of sensitive customer data.

Why this answer

Authorized views in BigQuery allow a view to be created in one dataset that reads from a source dataset, and then access to the view is granted to external partners without granting access to the underlying source data. This enables column- and row-level filtering (via the view's SQL) so partners see only the permitted subset. This is the standard BigQuery pattern for sharing restricted data.

Exam trap

PDE often tests the misconception that dataset-level permissions or materialized views provide fine-grained sharing — candidates pick dataset access controls when the requirement is column/row filtering, which only authorized views deliver.

How to eliminate wrong answers

Option A is wrong because materialized views are for performance optimization (precomputed results) and do not by themselves provide a secure sharing boundary; they still require access to underlying data. Option C is wrong because clustered tables improve query performance by physically sorting data, not access control. Option D is wrong because dataset-level access controls grant access to the entire dataset, not a filtered subset, so partners would see all columns and rows.

27
MCQhard

An organization is implementing a data lake on Google Cloud using Cloud Storage. They need to process both batch and streaming data with a unified pipeline. The team has experience with Apache Beam. Which architecture should they use to minimize operational overhead?

A.Kappa architecture with Cloud Dataflow using the same pipeline for batch and streaming
B.Use Cloud Dataproc for batch and Cloud Dataflow for streaming
C.Lambda architecture with Cloud Dataflow for batch and Cloud Pub/Sub for streaming
D.Use Cloud Data Fusion for both batch and streaming
AnswerA

Cloud Dataflow runs the same Apache Beam pipeline for both bounded and unbounded sources, so one codebase serves batch and streaming. This satisfies the unified-pipeline requirement while remaining fully managed, minimising operational overhead for the Beam-experienced team.

Why this answer

Kappa architecture uses a single streaming pipeline for both batch and streaming, simplifying operations. Dataflow implements Beam and supports both modes.

28
MCQmedium

A data engineer needs to create a BigQuery table that is partitioned by ingestion time and clustered by customer_id and transaction_date. They also want to limit access so that only users from a specific domain can query the table. Which approach should they use?

A.Create the table with partitioning only, then use a materialized view to restrict access
B.Create the table without clustering, use row-level security to filter by domain, and grant access to the table
C.Create the table with partitioning and clustering, then create an authorized view on the table and grant the view access to the domain users
D.Create the table with partitioning and clustering, then grant bigquery.dataViewer to the domain via IAM at the dataset level
AnswerC

Partitioning by ingestion time and clustering by customer_id and transaction_date optimises the table's physical layout. An authorised view then runs with the owner's permissions, so granting domain users access to the view alone satisfies the domain-restriction requirement without exposing the base table.

Why this answer

Authorized views allow sharing query results with specific users/groups without giving direct table access. Clustering and partitioning are defined at table creation. IAM roles at dataset level are too broad.

Row-level security filters rows but doesn't restrict domain.

29
MCQmedium

A company needs to process high-throughput streaming data with low latency. They are considering Cloud Pub/Sub for ingestion and Cloud Dataflow for processing. However, they are concerned about cost. Which alternative to Cloud Pub/Sub would reduce costs while still meeting the throughput requirements?

A.Cloud Pub/Sub with pull subscriptions
B.Cloud Tasks
C.Cloud Pub/Sub Lite
D.Cloud Pub/Sub with push subscriptions
AnswerC

Pub/Sub Lite provides the same high-throughput, low-latency ingestion as Pub/Sub but at a substantially lower cost, because it trades away automatic multi-zone replication and requires you to provision capacity in a specific zone. That trade-off satisfies the throughput requirement while directly addressing the stated cost concern.

Why this answer

Cloud Pub/Sub Lite is a zonal, cost-optimized messaging service designed for high-throughput streaming with predictable, low latency. Unlike standard Pub/Sub, which replicates data across zones and charges for data transfer and storage, Pub/Sub Lite lets you choose a capacity (in MiB/s) and charges a flat hourly rate, significantly reducing costs for steady, high-volume workloads. It integrates with Cloud Dataflow via the Pub/Sub Lite I/O connector, so the processing pipeline remains viable.

Thus, it meets the throughput and latency requirements while addressing cost concerns.

Exam trap

PDE often tests the misconception that pull or push subscriptions are cost-saving alternatives, when they are merely delivery methods for standard Pub/Sub, and confuses Cloud Tasks with a streaming ingestion service.

How to eliminate wrong answers

Option A is wrong because pull subscriptions are a consumption method for standard Pub/Sub, not a separate service; they do not inherently reduce costs compared to push subscriptions and still incur standard Pub/Sub pricing for storage and data transfer. Option B is wrong because Cloud Tasks is a task queue for asynchronous, single-consumer task execution, not a high-throughput streaming ingestion service; it lacks the scalability and ordering guarantees needed for streaming data and would not meet throughput requirements. Option D is wrong because push subscriptions are another consumption method for standard Pub/Sub, not a cost-reducing alternative; they can actually increase costs due to HTTP endpoint overhead and retries, and do not change the underlying Pub/Sub pricing model.

30
MCQhard

A financial services firm runs a Dataflow batch pipeline that joins a 2 TB transaction dataset with a 40 GB customer reference dataset. The reference data changes only once per day and is currently read from a BigQuery table with a side-input transform on every element. Job cost is dominated by repeated BigQuery reads, and the pipeline occasionally hits quota errors. The team wants to minimize cost and quota pressure while keeping the daily refresh. What should the data engineer change?

A.Load the reference dataset into a BigQuery table partitioned by ingestion date and read only the latest partition as the side input.
B.Enable BigQuery BI Engine on the reference table so side-input queries are served from memory instead of disk.
C.Replace the side input with a CoGroupByKey transform that groups transactions and customer records by customer ID before the join.
D.Stage the reference dataset as an Avro file in Cloud Storage, load it into the pipeline with a file-based source, and pass it as a side input refreshed by a scheduled daily export from BigQuery.
AnswerD

Exporting the slowly changing 40 GB reference data once per day to Cloud Storage removes per-element BigQuery reads entirely, and Avro gives compact, schema-carrying storage that Dataflow reads efficiently. The single daily export keeps the reference fresh while eliminating the quota pressure and repeated scan cost that the BigQuery side input caused on every pipeline run.

Why this answer

Moving the slowly changing reference dataset out of BigQuery into a compact Cloud Storage format and refreshing it once daily eliminates read amplification. A single scheduled export replaces thousands of per-element reads, cutting both cost and quota pressure while preserving the daily update cadence. CoGroupByKey, partition pruning on a side input, and BI Engine all leave the repeated BigQuery access pattern intact.

Exam trap

The trap here is believing that a side input is always cheap, when a large side input is re-materialized per worker and re-read from its source.

31
MCQhard

A data pipeline uses Cloud Data Fusion to perform ETL jobs. The pipeline reads from BigQuery, transforms data using Wrangler, and writes to Cloud Storage. The team notices that the pipeline runs slower than expected. They suspect the Data Fusion instance is under-provisioned. Which action should be taken to improve performance?

A.Add more Dataproc Metastore instances
B.Change the Data Fusion instance type from Basic to Enterprise
C.Enable Data Fusion accelerator for BigQuery
D.Rewrite the pipeline using Cloud Dataprep instead
AnswerB

Data Fusion instance type determines the available compute and memory for pipeline execution. Basic instances are limited to a small, fixed profile, so upgrading to Enterprise provides greater resources and the ability to scale executors, directly addressing the suspected under-provisioning causing slow ETL runs.

Why this answer

Cloud Data Fusion instance types (Basic, Enterprise, Developer) determine the compute resources available to run pipelines. Upgrading from Basic to Enterprise increases the number of Dataproc worker nodes and resources available to the CDAP runtime, improving pipeline throughput and reducing runtime. This directly addresses the under-provisioned instance suspicion.

Exam trap

PDE often tests the confusion that adding metadata or auxiliary services (like Dataproc Metastore) improves ETL performance, when the real lever is the Data Fusion instance type and Dataproc cluster sizing.

How to eliminate wrong answers

Option A is wrong because Dataproc Metastore is a metadata service for Hive/Spark catalogs and does not provide compute capacity for Data Fusion pipelines. Option C is wrong because there is no 'Data Fusion accelerator for BigQuery' feature; Data Fusion uses plugins/connectors, not accelerators. Option D is wrong because rewriting in Cloud Dataprep changes the tooling but does not address the under-provisioned Data Fusion instance and Dataprep is more of a data preparation UI than a full ETL runtime.

32
MCQeasy

You need to store petabytes of data in a data warehouse that supports ANSI SQL, automatic scaling, and real-time analytics. The data is primarily used for ad-hoc queries and business intelligence. Which Google Cloud service should you use?

A.Cloud SQL
B.Cloud Bigtable
C.Cloud Spanner
D.BigQuery
AnswerD

BigQuery is a fully managed, petabyte-scale data warehouse that supports ANSI SQL and provides automatic scaling and real-time analytics. It is designed for ad-hoc queries and business intelligence workloads, with separation of storage and compute, and integrates with BI tools. It is the ideal choice for this scenario.

Why this answer

BigQuery is the correct choice because it is a serverless, petabyte-scale data warehouse that supports ANSI SQL, automatically scales, and is optimized for ad-hoc queries and business intelligence. Its separation of storage and compute allows independent scaling and cost-effective analytics on large datasets.

Exam trap

The trap here is confusing a transactional database like Cloud Spanner with an analytical data warehouse, overlooking that BigQuery is purpose-built for large-scale SQL analytics.

33
MCQmedium

You are designing a BigQuery data warehouse for a retail company. Queries frequently filter on order_date and customer_id. To optimize query performance and cost, which table design should you use?

A.Cluster by order_date and partition by customer_id
B.Partition by ingestion_time and cluster by order_date
C.Use a clustered table without partitioning
D.Partition by order_date and cluster by customer_id
AnswerD

Partitioning by order_date prunes scanned data when queries filter on that column, while clustering by customer_id co-locates rows sharing that value within each partition. This combination satisfies both frequent filter predicates, reducing bytes scanned and query cost compared with clustering or partitioning alone.

Why this answer

Partitioning by order_date and clustering by customer_id aligns the physical storage with the query access pattern: partition pruning eliminates irrelevant date ranges, and clustering sorts data within each partition by customer_id so filter and aggregation operations scan fewer blocks. This combination minimizes bytes scanned, which directly reduces both query latency and cost in BigQuery's on-demand pricing model.

Exam trap

PDE often tests the partition-then-cluster ordering — candidates reverse the two or choose ingestion_time partitioning, not realizing that the partition column must match the dominant filter predicate to enable pruning.

How to eliminate wrong answers

Option A is wrong because it reverses the roles — clustering by order_date and partitioning by customer_id would create a partition per customer, which is a poor cardinality choice and prevents effective date-range pruning. Option B is wrong because partitioning by ingestion_time does not match the order_date filter, so queries filtering on order_date cannot prune partitions and must scan all ingestion-time partitions. Option C is wrong because a clustered table without partitioning forgoes partition pruning entirely, so date-range filters scan the full table even though clustering helps with customer_id.

34
MCQeasy

A small startup wants to run a nightly batch job that transforms a 2 GB CSV file in Cloud Storage and writes the result back as Parquet. The team has no cluster administration experience, wants per-job pricing rather than an always-on cluster, and needs the transformation to run in under thirty minutes. Which Google Cloud approach should they choose?

A.Use Cloud Data Fusion to build a visual pipeline and run it on a dedicated ephemeral Dataproc cluster each night.
B.Create a persistent Dataproc cluster with a fixed number of worker nodes and submit the job to it every night.
C.Use a serverless Spark job on Dataproc Serverless, submitting the transformation and letting the platform provision and tear down resources per run.
D.Run the transformation on a Compute Engine VM with a cron job that invokes a local Spark installation each night.
AnswerC

Dataproc Serverless runs Spark workloads without provisioning or managing a cluster, and billing is based on the resources the workload consumes during execution. For a nightly 2 GB transformation, the job completes quickly and the team pays only for that run. This matches the requirements for no administration and per-job cost.

Why this answer

Dataproc Serverless executes Spark workloads without any cluster provisioning, and you pay for the resources consumed while the workload runs. A nightly 2 GB CSV-to-Parquet transformation finishes well within the thirty-minute window, so the team gets per-job pricing and avoids cluster administration entirely.

Exam trap

The trap here is equating Dataproc with clusters you must manage, when Dataproc Serverless runs Spark workloads on demand with no cluster to provision or tear down.

35
MCQhard

You are building a real-time fraud detection system using Dataflow. Events from Pub/Sub need to be grouped by user_id within a 5-minute window to detect suspicious patterns. Some events may be delayed by up to 2 minutes. How should you configure the window and trigger to balance accuracy and latency?

A.Sliding window of 5 minutes with a 1-minute period and no allowed lateness
B.Session window with a gap duration of 5 minutes
C.Fixed window of 5 minutes with no allowed lateness and default trigger
D.Fixed window of 5 minutes with allowed lateness of 2 minutes and early trigger every 1 minute
AnswerD

Fixed windows align to epoch boundaries, so a 5-minute window groups events by user_id without overlap. Allowed lateness of 2 minutes retains late-arriving events, satisfying the stated delay constraint. Early triggers every minute emit speculative results, balancing latency against the accuracy gained from late data.

Why this answer

Option D is correct because it directly addresses both the accuracy requirement (handling up to 2 minutes of late data) and the latency requirement (emitting early results every 1 minute). A fixed 5-minute window groups events into non-overlapping intervals, which is appropriate for per-user fraud pattern detection. Setting allowed lateness to 2 minutes ensures that events delayed by up to 2 minutes are still included in the correct window, and an early trigger (e.g., repeating every 1 minute) provides low-latency partial results before the window closes.

This combination balances completeness and timeliness.

Exam trap

PDE often tests the misconception that a sliding window is needed for overlapping patterns, but the question specifies grouping by user_id within a 5-minute window, which implies non-overlapping fixed windows; also, candidates may overlook the need for allowed lateness to handle delayed events, opting for no lateness and thus sacrificing accuracy.

How to eliminate wrong answers

Option A is wrong because a sliding window with a 1-minute period creates overlapping windows, which would cause the same event to be counted in multiple windows, leading to duplicate fraud alerts and incorrect aggregation; also, no allowed lateness means delayed events are dropped, reducing accuracy. Option B is wrong because session windows group events based on gaps of inactivity, not fixed time intervals; a 5-minute gap duration would create windows that vary in length and do not align with the requirement to group events within a 5-minute window, making it unsuitable for detecting patterns in fixed time frames. Option C is wrong because a fixed window with no allowed lateness and the default trigger will drop any event arriving after the window closes, so the 2-minute delayed events would be lost, compromising accuracy; the default trigger also emits results only at the end of the window, increasing latency.

36
MCQmedium

A data pipeline ingests streaming events into Pub/Sub and needs to join them with a slowly updating reference table (few thousand rows) from a Cloud Storage CSV file. The pipeline runs on Dataflow with Apache Beam. Which approach is most cost-effective and operationally simple?

A.Read the CSV in a DoFn and perform a BigQuery query each time an event is processed
B.Use a side input that reads the CSV once and broadcasts it to all workers
C.Implement a custom sink that writes events to Cloud SQL and performs a SQL JOIN there
D.Use CoGroupByKey to join the stream and batch PCollections by a common key after reading the CSV into a batch PCollection each window
AnswerB

A side input reads the CSV once and broadcasts the small reference table to every worker, so each streaming element joins locally without per-element Cloud Storage reads or external lookups. This satisfies the cost-effectiveness and operational-simplicity constraints for a few-thousand-row table, avoiding the latency and expense of querying an external store per event.

Why this answer

Option B is correct because a Beam side input reads the small Cloud Storage CSV once and broadcasts the reference table to all workers, letting each streaming element be enriched in memory without per-event external calls, which is both cheap and operationally simple. Since the table has only a few thousand rows, it easily fits in worker memory as a side input. Option A is wrong because issuing a BigQuery query per event is slow, expensive, and adds an external dependency.

Option C is wrong because routing events through Cloud SQL and doing SQL JOINs adds infrastructure and cost. Option D is wrong because CoGroupByKey requires reading the CSV into a batch PCollection per window and shuffling both sides, which is more complex and less efficient than a side input for a small, slowly changing table.

37
MCQeasy

A startup is building a data lake on Google Cloud. They need to store raw JSON, CSV, and Parquet files from various sources. The files will be accessed by multiple analytics tools, including BigQuery and Dataproc. The startup wants a cost-effective, durable, and highly available storage solution that integrates natively with these services. Which Google Cloud service should they use?

A.Cloud SQL for MySQL with a large SSD and binary large object (BLOB) storage.
B.Cloud Bigtable with a column family for each data format.
C.Persistent Disk attached to a Compute Engine instance, shared via NFS.
D.Cloud Storage with a Standard storage class.
AnswerD

Cloud Storage is the foundational object store for data lakes on Google Cloud. It provides durable, highly available, and cost-effective storage with native integration into BigQuery, Dataproc, and other analytics services. The Standard class is suitable for frequently accessed data, and lifecycle policies can transition older data to colder classes to optimize cost.

Why this answer

Cloud Storage is the correct choice for a data lake because it is a durable, highly available, and cost-effective object store that integrates natively with BigQuery, Dataproc, and other Google Cloud analytics services. It supports all file formats and can be organized into buckets with lifecycle policies for cost optimization. The Standard storage class is appropriate for frequently accessed data.

Exam trap

The trap here is considering Bigtable or Cloud SQL because they are managed services, but they are not designed for storing raw files in a data lake architecture and lack native integration with analytics tools for file-based access.

38
MCQhard

A logistics company uses Cloud Dataflow to process a continuous stream of GPS events from delivery trucks. The pipeline must compute the distance traveled per truck per hour and write the results to BigQuery. The events are keyed by truck ID, and the pipeline uses windowing with a one-hour fixed window. The team notices that some trucks report events with timestamps that are several minutes late due to network delays. They want to ensure that late events are still included in the correct window and that the results are emitted only after a reasonable wait. What should they configure?

A.Use a session window with a gap duration equal to the expected network delay.
B.Set a trigger that fires when the watermark passes the end of the window, and set allowed lateness to zero.
C.Set a trigger that fires when the watermark passes the end of the window, and set allowed lateness to a duration longer than the expected delay.
D.Set a trigger that fires every minute, and set allowed lateness to zero.
AnswerC

Configuring allowed lateness to a duration greater than the maximum expected delay ensures that late events are still accepted and processed into the correct window. The trigger firing on the watermark passing the window end provides timely results, and the allowed lateness extends the window's lifetime to accommodate late data. This balances completeness and latency.

Why this answer

To include late events in the correct window, you must set allowed lateness to a value greater than the expected delay. The trigger on the watermark passing the window end ensures results are emitted after the window closes, and the allowed lateness keeps the window state alive to accept late data. This configuration balances timely results with completeness.

Exam trap

The trap here is focusing only on triggers and forgetting that allowed lateness controls whether late data is accepted; without it, late events are discarded regardless of the trigger.

39
MCQhard

A retail company runs a batch pipeline in Cloud Dataflow that reads from Cloud Storage and writes to BigQuery. The pipeline uses a GroupByKey operation on a key that is heavily skewed: one customer ID accounts for 40% of all transactions. This causes a single worker to process a massive amount of data, and the job takes hours longer than expected. The team wants to reduce the skew without changing the pipeline's logic or output. What should they do?

A.Increase the number of workers and use a larger machine type for the Dataflow job.
B.Replace GroupByKey with CombinePerKey using an associative and commutative CombineFn.
C.Apply a hot key fanout by adding a random secondary key, perform the GroupByKey, then remove the secondary key and re-group.
D.Increase the disk size of the worker VMs to allow more data to be spilled to disk.
AnswerC

Hot key fanout splits a single hot key into multiple keys by appending a random suffix, so the GroupByKey distributes the load across many workers. After the first grouping, a second grouping removes the suffix and combines the partial results. This reduces skew without altering the pipeline's logic or output, and it is a standard Dataflow pattern for hot keys.

Why this answer

Hot key fanout is the correct technique for mitigating skew in a GroupByKey. By adding a random secondary key, the hot key's records are spread across multiple workers, and a subsequent grouping removes the secondary key to produce the final result. This preserves the pipeline's semantics while dramatically improving parallelism for the skewed key.

Exam trap

The trap here is thinking that adding more workers or larger machines can solve a hot key bottleneck, when the issue is that a single key must be processed by one worker regardless of cluster size.

40
MCQhard

A financial services firm ingests trade events into Cloud Storage and must load them into BigQuery. Compliance requires that each event be processed exactly once and that the load be idempotent across retries, even if a Dataflow job restarts mid-batch. The destination table must also be queryable immediately after each successful load. Which loading approach best satisfies these requirements?

A.Use BigQuery's LOAD DATA statement triggered by Cloud Storage object notifications, appending each new file into the destination table.
B.Use a Dataflow pipeline with the BigQueryIO Storage Write API in exactly-once mode, writing to the table with deterministic record identifiers.
C.Use a Dataproc Spark job that writes Parquet files to Cloud Storage, then run a scheduled bq load command every fifteen minutes to append into BigQuery.
D.Use a Dataflow pipeline that writes to BigQuery with STREAMING_INSERTS and relies on the pipeline's exactly-once semantics to deduplicate on retry.
AnswerB

The Storage Write API with exactly-once semantics assigns stream offsets and only commits data when checkpoints succeed, so retries do not duplicate rows. Combined with deterministic record identifiers and Dataflow checkpointing, a restarted job resumes without re-inserting committed records. Data written through committed streams becomes immediately queryable, satisfying the requirement that the table be readable right after each successful load.

Why this answer

Exactly-once loading into BigQuery is achieved with the Storage Write API's exactly-once streams, where offsets are committed only on checkpoint and deterministic record IDs prevent duplicates on restart. This gives idempotent loads across retries and makes committed rows immediately queryable. Streaming inserts and periodic batch appends both risk duplicates and cannot meet the compliance constraint.

Exam trap

The trap here is believing that Dataflow's exactly-once processing guarantee automatically extends to any BigQuery sink, when the sink itself must support offset-based commits to avoid duplicates.

41
MCQhard

A retail company ingests point-of-sale clickstream events into Cloud Pub/Sub at roughly 200,000 messages per second during flash sales. Analysts need near-real-time dashboards that aggregate revenue by product category over sliding 5-minute windows, with results visible in BigQuery within 30 seconds of the event. The pipeline must handle occasional bursts up to 3x the normal rate without dropping messages, and the team wants to minimize operational overhead. Which design should the data engineer use?

A.Use Cloud Pub/Sub push subscriptions delivering events directly to a Cloud Run service that aggregates in memory and calls the BigQuery streaming insert API once per minute.
B.Deploy a Dataproc cluster with Spark Structured Streaming reading from Pub/Sub and writing micro-batches to BigQuery every 10 seconds, sizing the cluster for the 3x burst rate.
C.Create a Cloud Dataflow streaming pipeline in the Apache Beam Java SDK using Pub/Sub as the source, apply windowing with a 5-minute sliding window and a 30-second allowed lateness, and write aggregated results to BigQuery using the Storage Write API.
D.Load raw events into Cloud Storage using a Pub/Sub subscription with a Cloud Storage sink, then run a BigQuery scheduled query every 5 minutes that reads the new files and computes the sliding window aggregates.
AnswerC

Dataflow provides exactly-once streaming semantics, autoscaling that absorbs the 3x burst, and native Pub/Sub and BigQuery connectors. Sliding windows with allowed lateness handle out-of-order events, and the Storage Write API gives low-latency, high-throughput BigQuery inserts, satisfying the 30-second visibility requirement without managing clusters.

Why this answer

The requirement for sliding 5-minute windows with sub-30-second BigQuery visibility under bursty load points to a managed stream processor with native windowing and autoscaling. Dataflow's sliding windows, allowed lateness, and BigQuery Storage Write API together meet latency, correctness, and operational-overhead constraints. Cluster-based and file-based approaches add latency or management burden that the scenario rules out.

Exam trap

The trap here is assuming that any Pub/Sub consumer with a BigQuery sink is equivalent, when only a stateful stream processor with true windowing semantics can meet the latency and correctness requirements.

42
MCQeasy

A startup needs a fully managed, serverless Spark service to run occasional data processing jobs without managing clusters. They want to pay only for the resources used during job execution. Which Google Cloud service should they use?

A.Dataproc Serverless
B.Dataflow
C.Cloud Data Fusion
D.Dataproc
AnswerA

Dataproc Serverless runs Spark workloads without provisioning or managing a cluster, allocating resources only while the job executes and charging for that consumption. This matches the stem's requirements for occasional jobs, no cluster management and pay-per-use billing, unlike a standard Dataproc cluster.

Why this answer

Dataproc Serverless provides a serverless Spark environment where you pay per job execution. Cloud Data Fusion is for visual ETL. Dataproc is managed but not serverless.

Dataflow is serverless for Beam, not Spark.

43
Multi-Selecthard

A financial services firm is designing a data processing system on Google Cloud that must ingest change data capture (CDC) streams from an on-premises PostgreSQL database into BigQuery with sub-minute latency, preserve the ordering of changes per primary key, and apply updates and deletes so that BigQuery reflects the current state of each row. The source database cannot be modified to add triggers. Which two design elements should you include? (Choose two.)

Select 2 answers
A.Add database triggers to the PostgreSQL source that publish change events to a Pub/Sub topic.
B.Use Datastream to read the PostgreSQL write-ahead log and stream changes into Cloud Storage or BigQuery.
C.Use a federated BigQuery external table that queries the on-premises PostgreSQL instance directly at read time.
D.Enable BigQuery streaming inserts with insertId deduplication to apply deletes directly to the target table.
E.Configure the BigQuery destination to use the staging and merge pattern with a primary key so deletes and updates are applied.
AnswersB, E

Datastream performs log-based CDC by reading the PostgreSQL write-ahead log, so no triggers or schema changes are needed on the source, satisfying the constraint that the database cannot be modified. It delivers low-latency change records that can land in Cloud Storage or be written directly into BigQuery, providing the ingestion mechanism the design requires.

Why this answer

Log-based CDC through Datastream reads the PostgreSQL write-ahead log without touching the source schema, and the staging-plus-merge destination pattern keyed on the primary key is what turns a change stream into correct current-state rows in BigQuery, including deletes. Together they meet the latency, ordering, and mutation requirements.

Exam trap

The trap here is assuming that streaming inserts into BigQuery can apply updates and deletes, when the streaming API is append-only and mutations require a merge pattern.

44
Multi-Selectmedium

A retail company is designing a Dataflow pipeline to process point-of-sale transactions from Cloud Pub/Sub and write to BigQuery. The pipeline must handle late-arriving data up to 24 hours and ensure that all data is written to BigQuery exactly once, even in the event of worker failures. Which two features should the engineer implement to meet these requirements? (Choose two.)

Select 2 answers
A.Use the Dataflow ExactlyOnce processing mode and enable the BigQueryIO.Write with FILE_LOADS and a trigger that fires after 24 hours.
B.Set the allowed lateness to 24 hours on the windowing strategy.
C.Enable the Dataflow ExactlyOnce processing mode and use the BigQueryIO.Write with STORAGE_WRITE_API and a deterministic deduplication key.
D.Use BigQueryIO.Write with the UseBeamSchema option and set the write disposition to WRITE_APPEND.
E.Configure the pipeline to use the BigQueryIO.Write with STREAMING_INSERTS and set the insertId based on a unique transaction identifier.
AnswersB, C

Setting allowed lateness to 24 hours allows the pipeline to accept and process data that arrives up to 24 hours after the window ends. This directly addresses the late-arriving data requirement. Combined with a proper trigger, it ensures that late data is included in the results and written to BigQuery.

Why this answer

To handle late-arriving data up to 24 hours, the pipeline must set allowed lateness to 24 hours on the windowing strategy. To achieve exactly-once writes to BigQuery, the engineer should use the Storage Write API with a deterministic deduplication key and enable Dataflow's ExactlyOnce mode. Together, these features satisfy both requirements.

Exam trap

The trap here is assuming that streaming inserts with insertId provide exactly-once semantics, but they only offer best-effort deduplication and can still produce duplicates.

45
MCQmedium

Your company is migrating an on-premises Apache Hadoop cluster to Google Cloud. The cluster runs Hive for SQL-like queries and stores data in HDFS. You want a managed service that minimizes operational overhead while supporting existing Hive scripts. Which Google Cloud service should you choose?

A.Cloud Dataflow
B.Cloud Bigtable
C.Cloud Dataproc
D.BigQuery
AnswerC

Cloud Dataproc is a managed Spark and Hadoop service that supports Hive, HDFS, and other Hadoop ecosystem tools. It allows you to run existing Hive scripts with minimal changes and provides cluster management, autoscaling, and integration with Cloud Storage. It is the ideal choice for migrating on-premises Hadoop workloads.

Why this answer

Cloud Dataproc is the correct choice because it is a managed Hadoop and Spark service that supports Hive, HDFS, and other ecosystem tools, allowing you to run existing Hive scripts with minimal changes. It reduces operational overhead by handling cluster provisioning, configuration, and scaling.

Exam trap

The trap here is assuming that a serverless data warehouse like BigQuery can run existing Hive scripts, overlooking that Dataproc is designed for Hadoop compatibility.

46
MCQhard

A Dataproc cluster uses preemptible worker nodes to reduce costs. The cluster runs a long-running Spark job that occasionally experiences worker failures. How should the job be configured to handle preemptible worker failures gracefully?

A.Set spark.task.maxFailures to a high number to allow retries.
B.Disable preemptible workers for the job.
C.Use persistent disks for preemptible workers.
D.Enable automatic restart of the Spark driver on failure.
AnswerA

Raising `spark.task.maxFailures` lets individual tasks retry after a preemptible worker is reclaimed, so the job survives transient node loss without failing the stage. However, this only addresses task-level retries; Dataproc's enhanced flexibility mode, which pairs a standard primary pool with preemptible secondary workers, is what actually enables graceful recovery.

Why this answer

Preemptible VMs on Dataproc can be reclaimed at any time, causing executor loss mid-task. Setting spark.task.maxFailures to a higher value (default is 4) allows Spark to retry failed tasks on remaining executors instead of failing the entire job. This is the standard, low-cost way to tolerate preemptible worker churn without abandoning the cost savings.

Exam trap

PDE often tests whether candidates confuse driver-level resilience (automatic restart) with executor/task-level resilience (maxFailures), or mistakenly believe persistent disks prevent preemption.

How to eliminate wrong answers

Option B is wrong because disabling preemptible workers eliminates the cost benefit and defeats the purpose of the scenario; the question asks how to handle failures gracefully, not avoid them. Option C is wrong because persistent disks do not prevent preemption — preemptible VMs are still reclaimed, and persistent disks only preserve data, not running executors. Option D is wrong because restarting the Spark driver addresses driver failure, not executor/task failure caused by preemptible worker loss; the driver is typically on a non-preemptible node.

47
Multi-Selectmedium

You are designing a Dataflow pipeline for processing real-time clickstream data. The pipeline must group events into 30-second windows and handle late data up to 5 minutes. You want to output partial results every 10 seconds for low-latency monitoring. Which THREE configurations should you use? (Choose three.)

Select 3 answers
A.Use sliding windows of 30 seconds with a 10-second period
B.Use a trigger that fires after the end of the window
C.Use fixed windows of 30 seconds
D.Set allowed lateness to 5 minutes
E.Use a trigger with early firings every 10 seconds
AnswersC, D, E

Fixed 30-second windows partition the clickstream into non-overlapping intervals matching the grouping requirement. Combined with allowed lateness and early triggering, they satisfy the windowing constraint while the other settings handle late data and partial results.

Why this answer

Option C is correct because fixed (tumbling) 30-second windows partition the clickstream into non-overlapping 30-second intervals, which is the required grouping for this scenario. Option D is correct because setting allowed lateness to 5 minutes lets the pipeline retain window state and accept events that arrive up to 5 minutes after the window closes, matching the late-data requirement. Option E is correct because a trigger with early firings every 10 seconds emits speculative partial results before the window closes, providing the requested low-latency monitoring output.

Option A is not appropriate because sliding windows with a 10-second period would create overlapping 30-second windows and duplicate events across windows, which is not the specified grouping. Option B is not appropriate because a trigger that fires only after the end of the window would produce results only at window closure and would not deliver the required 10-second partial outputs.

Exam trap

Candidates might focus only on windowing and lateness, forgetting that early firing triggers are essential for periodic partial results as specified.

48
MCQeasy

A company needs a messaging service for event-driven applications that require low cost for high-throughput, but can tolerate occasional message loss. Which Pub/Sub product should they choose?

A.Pub/Sub with pull subscriptions
B.Pub/Sub with dead letter topics
C.Pub/Sub with push subscriptions
D.Pub/Sub Lite
AnswerD

Pub/Sub Lite provisions capacity in zonal or regional reservations, giving a much lower throughput price than standard Pub/Sub. Its at-least-once delivery offers no replication guarantee across zones, matching the stated tolerance for occasional message loss at high volume.

Why this answer

Pub/Sub Lite is designed for cost-sensitive workloads with relaxed durability. Standard Pub/Sub offers at-least-once delivery and high durability. Push vs pull is irrelevant to cost.

49
Multi-Selectmedium

A company wants to use Dataproc Metastore to manage metadata for their Spark jobs. Which TWO benefits does Dataproc Metastore provide?

Select 2 answers
A.Automatic scaling of compute resources
B.High availability with automatic failover
C.Fully managed Hive metastore service
D.Integration with BigQuery
E.Built-in data lineage tracking
AnswersB, C

Dataproc Metastore is a fully managed, highly available service that replicates metadata across zones and performs automatic failover without administrative intervention. This satisfies the stem's requirement for a managed metadata service for Spark jobs, eliminating the single point of failure inherent in a self-managed Hive Metastore deployment.

Why this answer

Option B is correct because Dataproc Metastore is a highly available, fully managed service that provides automatic failover across zones, ensuring metadata remains accessible even if a zone fails. Option C is correct because Dataproc Metastore is essentially a fully managed, serverless implementation of the Hive Metastore (HMS), which Spark and other engines use to store and retrieve metadata such as table schemas and partitions. Option A is incorrect because automatic scaling of compute resources is a feature of Dataproc clusters or autoscaling policies, not of the metadata service itself.

Option D is incorrect because Dataproc Metastore does not provide native BigQuery integration; it serves Hive/Spark-style metadata via the Hive Metastore Thrift protocol. Option E is incorrect because data lineage tracking is provided by tools like Data Catalog or Dataplex, not by Dataproc Metastore.

Exam trap

PDE often tests the distinction between Dataproc (compute) features like autoscaling and Dataproc Metastore (metadata) features like HA and managed Hive metastore, causing candidates to pick compute-related options.

50
MCQmedium

A data engineer needs to design a streaming pipeline that ingests events from multiple sources, enriches them with a lookup table stored in BigQuery (updated every hour), and writes the results to a BigQuery table for real-time dashboards. The pipeline must handle late-arriving data up to 1 hour. Which Dataflow feature should be configured to manage late data?

A.A custom watermark estimation function
B.Using side inputs with a periodic refresh
C.A trigger that fires on every late element
D.Allowed lateness on the window
AnswerD

Setting allowed lateness on the window lets Dataflow retain window state and emit updated results for up to one hour after the watermark passes, directly satisfying the stem's late-arriving data constraint. Unlike discarding mode, it triggers late firings so enriched events still reach the BigQuery dashboard table.

Why this answer

Allowed lateness on the window is the Dataflow feature that explicitly defines how long the pipeline will retain window state and accept late-arriving elements after the watermark has passed the end of the window. Setting allowed lateness to 1 hour lets Dataflow keep the window's state for that duration, so events arriving up to 60 minutes late are still processed and emitted with correct results. This directly satisfies the requirement to handle late data up to 1 hour without dropping or misassigning events.

Exam trap

PDE often tests the confusion between watermark estimation (which controls when the pipeline thinks data is complete) and allowed lateness (which controls how long late data is actually accepted), causing candidates to pick the watermark option.

How to eliminate wrong answers

Option A is wrong because a custom watermark estimation function only affects how the watermark advances (i.e., when the pipeline believes all data for a window has arrived); it does not retain window state or accept late elements after the watermark passes. Option B is wrong because side inputs with periodic refresh are used to enrich streaming elements with slowly changing reference data (like the BigQuery lookup table), not to handle late-arriving events. Option C is wrong because a trigger that fires on every late element only controls when results are emitted; it does not extend the window's lifetime, so late elements beyond the default allowed lateness would still be dropped.

51
MCQeasy

An engineer needs to create a Pub/Sub subscription that sends messages to an HTTPS endpoint. The endpoint must be able to acknowledge messages individually. Which type of subscription should they use?

A.Pull subscription
B.Push subscription
C.BigQuery subscription
D.Cloud Storage subscription
AnswerB

Push subscriptions deliver messages to an HTTPS endpoint and let the endpoint return an HTTP success code per message, which acknowledges that individual message. This satisfies the requirement for individual acknowledgement without the subscriber pulling.

Why this answer

Push subscriptions deliver messages to a configured HTTPS endpoint. The endpoint can acknowledge by returning a 200 status.

52
Multi-Selecthard

You are designing a data processing architecture on Google Cloud. You need to ingest data from multiple sources, including streaming events and batch files, and process them to produce a unified dataset for analytics. The solution must support both real-time and historical processing with the same codebase, and be able to handle late-arriving data. Which two Google Cloud services should you use together to achieve this? (Choose two.)

Select 2 answers
A.Cloud Composer
B.Cloud Pub/Sub
C.Cloud Dataflow
D.Cloud Dataproc
E.BigQuery
AnswersB, C

Cloud Pub/Sub is a scalable, serverless messaging service for ingesting streaming data. It can handle high-throughput event streams and integrate seamlessly with Dataflow. Pub/Sub ensures reliable delivery and can buffer messages, which is essential for ingesting streaming events into a unified pipeline that also processes batch data.

Why this answer

To achieve unified batch and stream processing with the same codebase and handle late data, you need a processing engine that supports Apache Beam, such as Cloud Dataflow, and a scalable ingestion service for streaming data, such as Cloud Pub/Sub. Dataflow's Beam model allows you to write one pipeline that works for both batch and streaming, with built-in support for event-time windowing and late data. Pub/Sub provides reliable, scalable ingestion for streaming events.

Exam trap

The trap here is assuming that any data processing service can handle both batch and streaming with the same code, or that orchestration services like Cloud Composer can replace a stream processing engine.

53
MCQmedium

A logistics company uses Cloud Pub/Sub to ingest shipment tracking events. They want to archive all events to Cloud Storage for long-term retention and also process them in real time with Dataflow. The events are published to a single topic. Which design should the data engineer use to ensure both archiving and real-time processing without data loss?

A.Create two subscriptions on the topic: one for the Dataflow pipeline and one for a Cloud Function that writes to Cloud Storage.
B.Configure the Pub/Sub topic to write to Cloud Storage directly using a Pub/Sub to Cloud Storage subscription.
C.Use a single subscription for both the Dataflow pipeline and a Cloud Function that writes to Cloud Storage.
D.Use a single subscription for the Dataflow pipeline, and have the pipeline write to both Cloud Storage and its real-time processing output.
AnswerA

Pub/Sub allows multiple subscriptions on a single topic, each receiving a copy of every message. By creating one subscription for Dataflow and another for a Cloud Function that archives to Cloud Storage, both consumers independently receive all events. This ensures no data loss and allows independent processing. This is the standard fan-out pattern.

Why this answer

The correct design is to create two separate subscriptions on the Pub/Sub topic: one for Dataflow and one for a Cloud Function that archives to Cloud Storage. This fan-out pattern ensures that each event is delivered to both consumers independently, preventing data loss and allowing each to process at its own pace. It decouples archiving from real-time processing.

Exam trap

The trap here is assuming that a single subscription can be shared by multiple consumers to receive all messages, but Pub/Sub delivers each message to only one consumer per subscription.

54
MCQmedium

A company uses Cloud Dataproc to run Spark ML training jobs. They want to persist the trained models and metadata in a Hive-compatible metastore. Which Dataproc feature should they use?

A.Cloud Hive Metastore (self-managed)
B.Cloud Bigtable
C.Dataproc Metastore
D.Cloud Data Catalog
AnswerC

Dataproc Metastore provides a fully managed, Hive-compatible metastore service that persists table metadata and model artefacts independently of cluster lifecycle. This satisfies the requirement to retain trained models and metadata in a Hive-compatible store across ephemeral Dataproc clusters.

Why this answer

Dataproc Metastore is a fully managed, Hive-compatible metastore service on Google Cloud that integrates natively with Dataproc clusters, allowing Spark and Hive jobs to persist and share table metadata and schemas. It provides a serverless, scalable alternative to running a self-managed Hive metastore on a cluster, and it supports the Hive Metastore Thrift API so existing Spark ML workflows can store model metadata without code changes. This directly meets the requirement for a Hive-compatible metastore for trained models and metadata.

Exam trap

PDE often tests the confusion between metadata storage services (Dataproc Metastore vs. Data Catalog) and database services (Bigtable), causing candidates to choose a non-Hive-compatible option for metastore requirements.

How to eliminate wrong answers

Option A is wrong because a self-managed Cloud Hive Metastore requires provisioning and operating a Hive metastore on a VM or cluster, adding operational overhead and not being a managed Dataproc feature. Option B is wrong because Cloud Bigtable is a NoSQL wide-column database for low-latency workloads, not a Hive-compatible metadata store. Option D is wrong because Cloud Data Catalog is a metadata management and discovery service, not a Hive Metastore replacement that Spark/Hive can read and write table metadata to.

55
MCQmedium

A company is using Cloud Storage to store raw logs. They want to use Cloud Data Fusion to transform and load the data into BigQuery on a daily schedule. The transformations are complex and involve joining multiple datasets. What is the most efficient way to run these pipelines?

A.Use Cloud Composer to orchestrate Dataproc jobs that run the transformations
B.Use Cloud Functions to trigger a Dataflow job that does the transformations
C.Use Cloud Data Fusion to design the pipeline and schedule it to run on a Dataproc cluster
D.Use Cloud Dataprep to design the transformation and export to BigQuery
AnswerC

Cloud Data Fusion pipelines execute on ephemeral Dataproc clusters, which provide the distributed compute needed for complex joins across multiple datasets. Scheduling daily runs on Dataproc satisfies the stem's transformation complexity while remaining cost-efficient between executions.

Why this answer

Cloud Data Fusion is a fully managed, code-free ETL/ELT service built on CDAP that provides a visual pipeline designer with a rich set of plugins (including BigQuery sinks) and can execute pipelines on ephemeral Dataproc clusters. For complex transformations involving joins across multiple datasets, Data Fusion's Wrangler and Joiner transforms plus its scheduling capability make it the most efficient native choice. Scheduling the pipeline to run on a Dataproc cluster gives the compute needed for joins while keeping orchestration managed.

Exam trap

PDE often tests whether candidates confuse Cloud Data Fusion (visual ETL with scheduling) with Cloud Dataprep (data preparation only) or Cloud Composer (general orchestration) — the key differentiator is that Data Fusion is the managed ETL tool that natively schedules and runs on Dataproc.

How to eliminate wrong answers

Option A is wrong because Cloud Composer orchestrating Dataproc jobs requires writing Spark code and DAGs manually — it is not the most efficient way when Data Fusion already provides the visual transform and scheduling layer. Option B is wrong because Cloud Functions is serverless and short-lived, unsuitable for orchestrating complex multi-dataset joins, and Dataflow would require custom Apache Beam code rather than the requested Data Fusion approach. Option D is wrong because Cloud Dataprep is a data-preparation tool for profiling and cleaning, not a full ETL orchestrator with scheduling and complex join semantics at scale.

56
MCQeasy

A data engineer needs to process data in a Dataflow pipeline that reads from a Pub/Sub topic. The pipeline must group events into 5-minute windows and compute the average value per key. Which Beam transform should they use after windowing?

A.Combine.perKey
B.ParDo
C.GroupByKey
D.CoGroupByKey
AnswerA

Combine.perKey performs per-key aggregation after windowing, computing the average value for each key within each 5-minute window. It satisfies the requirement to group events and average per key, unlike global combines that would merge across keys.

Why this answer

Combine.perKey is the correct transform because it performs a per-key aggregation (here, computing the average value) after windowing, combining elements within each key and window efficiently. It is a fused Combine operation that reduces data before shuffling, making it more efficient than GroupByKey followed by a separate aggregation. Since the requirement is to compute an average per key within 5-minute windows, Combine.perKey directly expresses that intent.

Exam trap

The trap is choosing GroupByKey because it 'groups by key,' but the question asks for an aggregation — Combine.perKey is the efficient, purpose-built transform for per-key aggregation and avoids the full shuffle penalty of GroupByKey.

How to eliminate wrong answers

Option B is wrong because ParDo is a general-purpose element-wise transform for mapping/filtering individual elements; it does not perform per-key aggregation across a window on its own. Option C is wrong because GroupByKey groups all values for a key into an iterable but does not compute an aggregate — you would still need a subsequent transform, and it shuffles all values without pre-aggregation, which is less efficient. Option D is wrong because CoGroupByKey joins multiple PCollections by key, which is unnecessary here since only one source (Pub/Sub) is being aggregated.

57
MCQeasy

A startup is building a data lake on Google Cloud. They need to store raw JSON logs in a cost-effective manner and later query them using SQL with minimal transformation. The logs are infrequently accessed but must be retained for 7 years for compliance. Which storage solution should they use?

A.Cloud Bigtable with a column family for JSON logs and a HBase client for querying.
B.Cloud SQL for PostgreSQL with JSONB columns and scheduled exports to Cloud Storage.
C.Cloud Storage with Archive storage class and BigQuery external tables.
D.Cloud Storage with Nearline storage class and BigQuery external tables.
AnswerC

Archive storage is the lowest-cost option for long-term retention and is ideal for data accessed less than once a year. BigQuery external tables allow querying JSON data directly from Cloud Storage without loading, satisfying the SQL query requirement with minimal transformation. This combination is both cost-effective and functional for infrequent access over 7 years.

Why this answer

Archive storage in Cloud Storage is the most cost-effective for long-term retention of infrequently accessed data. BigQuery external tables enable SQL querying of JSON logs directly from Cloud Storage without loading or transformation. This meets the cost, retention, and query requirements.

The other options use higher-cost storage or databases not optimized for this use case.

Exam trap

The trap here is choosing Nearline or Coldline for 7-year retention when Archive is specifically designed for the lowest-cost, long-term storage.

58
MCQeasy

Which Google Cloud service provides a visual interface for building ETL pipelines using a drag-and-drop design and includes pre-built transforms from a marketplace?

A.Dataproc
B.Cloud Data Fusion
C.Dataprep
D.BigQuery
AnswerB

Cloud Data Fusion provides a graphical drag-and-drop interface for building ETL pipelines, with a reusable plugin marketplace of pre-built transforms and connectors. This directly satisfies the stem's requirement for visual pipeline design plus marketplace transforms, unlike code-first services such as Dataflow or Dataproc.

Why this answer

Cloud Data Fusion is Google Cloud's fully managed, code-free ETL/ELT service built on the open-source CDAP project, offering a graphical drag-and-drop pipeline designer and a marketplace of pre-built plugins and transformations. It lets data engineers build batch and streaming pipelines visually and deploy them to ephemeral Dataproc clusters. This matches the requirement for a visual interface with marketplace transforms.

Exam trap

The trap is confusing Dataprep with Data Fusion — both are visual, but Dataprep is for interactive data preparation while Data Fusion is the full ETL pipeline builder with a transform marketplace.

How to eliminate wrong answers

Option A is wrong because Dataproc is a managed Hadoop/Spark service that requires you to write code (Spark, Hive, Pig) — it provides infrastructure, not a drag-and-drop ETL designer. Option C is wrong because Dataprep (by Trifacta) is a visual data-wrangling tool for exploring and preparing data, but it is oriented toward interactive data preparation rather than full ETL pipeline orchestration with a transform marketplace. Option D is wrong because BigQuery is a serverless data warehouse for storing and querying data with SQL; it is a destination/analytics engine, not a visual ETL pipeline builder.

59
MCQmedium

A retail company runs a nightly batch pipeline that loads point-of-sale transactions into BigQuery. The pipeline uses a Cloud Composer (Apache Airflow) DAG with a BigQueryInsertJobOperator task that runs a SQL MERGE statement to upsert yesterday's sales into a large fact_sales table partitioned by transaction_date. The data engineering team notices that the MERGE task occasionally fails with a 'Resources exceeded during query execution' error when the source staging table contains more than 50 million rows. They need to redesign the DAG to reliably load large daily volumes without changing the final table schema or downstream dashboards. What should they do?

A.Replace the single MERGE statement with a multi-statement script that first deletes rows from the target partition for the affected dates and then inserts the staged rows using a BigQueryInsertJobOperator task with writeDisposition set to WRITE_APPEND.
B.Schedule the MERGE task to run during off-peak hours and increase the BigQuery slot reservation for the project so that the MERGE has more compute resources available.
C.Increase the number of Dataflow workers in the DAG by adding a DataflowOperator task that reads from the staging table and writes to the fact_sales table using a BigQueryIO write with STREAMING_INSERTS.
D.Convert the fact_sales table to a clustered table on transaction_date and re-run the original MERGE statement, relying on clustering to reduce the bytes scanned and avoid the resource error.
AnswerA

This approach avoids the memory-intensive shuffle of a single large MERGE by splitting work into a bounded DELETE on a partition and a streaming-friendly INSERT. BigQuery can execute each statement with less resource contention, and WRITE_APPEND avoids rewriting the entire target table. It preserves the schema and downstream behavior while scaling to large daily volumes.

Why this answer

A single large MERGE in BigQuery can hit internal resource limits because it must shuffle and join all source and target rows. Splitting the operation into a partition-scoped DELETE followed by an INSERT with WRITE_APPEND reduces the working set per statement and avoids the expensive shuffle. This pattern is a common, reliable way to upsert large daily batches while keeping the target schema and downstream queries unchanged.

Exam trap

The trap here is assuming that adding slots or clustering will fix a query that fails due to internal shuffle/memory limits, when the real fix is to decompose the MERGE into smaller, partition-scoped operations.

60
MCQhard

A financial services company needs to process credit card transactions in real time to detect fraudulent patterns. The pipeline must handle late-arriving data (up to 2 hours) and produce accurate results. They want to use a unified programming model that works for both batch and streaming. Which Google Cloud service should they use?

A.Cloud Pub/Sub with Cloud Functions
B.BigQuery with scheduled queries
C.Cloud Dataproc with Spark Streaming
D.Cloud Dataflow with Apache Beam
AnswerD

Cloud Dataflow, based on Apache Beam, provides a unified programming model for batch and streaming. It supports event-time processing, watermarks, and triggers to handle late-arriving data accurately. With windowing and allowed lateness, it can produce correct results even when data arrives up to 2 hours late, making it ideal for real-time fraud detection with late data.

Why this answer

Cloud Dataflow with Apache Beam is the correct choice because it offers a unified batch and streaming programming model, supports event-time processing with watermarks and triggers, and can handle late-arriving data accurately through configurable allowed lateness. This makes it well-suited for real-time fraud detection with data arriving up to 2 hours late.

Exam trap

The trap here is assuming that any streaming service can handle late data, but only Dataflow with Beam provides built-in event-time semantics and allowed lateness for accurate results.

61
MCQeasy

You need to allow a data analyst to run queries on a BigQuery dataset but prevent them from modifying the data or deleting the dataset. Which IAM role should you grant?

A.roles/bigquery.dataOwner
B.roles/bigquery.dataViewer
C.roles/bigquery.jobUser
D.roles/bigquery.dataEditor
AnswerB

roles/bigquery.dataViewer grants read access to datasets, tables and views, permitting queries while denying write operations such as inserts, updates, deletes or dataset removal. It therefore satisfies both constraints: the analyst can run queries but cannot modify data or delete the dataset.

Why this answer

roles/bigquery.dataViewer grants read-only access to BigQuery datasets, tables, and views, allowing the analyst to run queries and view metadata but not modify data or delete the dataset. It is the least-privilege role that satisfies the requirement of querying without write or delete permissions. This role is typically paired with roles/bigquery.jobUser at the project level so the user can actually run jobs.

Exam trap

The trap is picking roles/bigquery.jobUser alone because it 'lets you run queries,' but jobUser does not grant data access — you need dataViewer (or higher) on the dataset as well.

How to eliminate wrong answers

Option A is wrong because roles/bigquery.dataOwner grants full control over the dataset, including the ability to modify data, delete tables, and even delete the dataset — far more than the read-only access required. Option C is wrong because roles/bigquery.jobUser only allows the user to run jobs (queries, loads, exports) in a project; it does not by itself grant access to read the data in a dataset, so the analyst could not query the data. Option D is wrong because roles/bigquery.dataEditor allows the user to modify and delete data and tables, which violates the requirement to prevent modifications.

62
MCQhard

You are designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline must handle late-arriving data (up to 1 hour) and group events into 10-minute windows. Which configuration is correct?

A.Use global windows with a trigger that fires every 10 minutes
B.Use sliding windows of 10 minutes with a 5-minute period and allowed lateness of 1 hour
C.Use fixed windows of 10 minutes with allowed lateness of 0 seconds
D.Use fixed windows of 10 minutes with allowed lateness of 1 hour and a trigger that fires after watermark plus early firings
AnswerD

Fixed 10-minute windows satisfy the grouping requirement, while one hour of allowed lateness lets late events still update their window's results. The trigger firing on watermark plus early firings emits speculative results promptly, then corrects them once the watermark passes, so no late data within the hour is dropped.

Why this answer

Fixed windows of 10 minutes match the required grouping interval, allowed lateness of 1 hour accommodates late-arriving data up to one hour, and a trigger that fires after the watermark plus early firings ensures results are emitted promptly while still accepting late data. This combination is the canonical Apache Beam pattern for windowed aggregation with late data on a streaming pipeline.

Exam trap

PDE often tests whether candidates conflate windowing (grouping) with triggering (emission timing) — the trap is choosing a trigger-only answer (global windows with a periodic trigger) when the requirement is per-window aggregation with late data.

How to eliminate wrong answers

Option A is wrong because global windows do not partition events into 10-minute groups — a trigger firing every 10 minutes emits cumulative results over the entire stream, not per-window aggregates. Option B is wrong because sliding windows of 10 minutes with a 5-minute period produce overlapping windows, which is not what 'group events into 10-minute windows' requires and would duplicate aggregates. Option C is wrong because allowed lateness of 0 seconds discards all late data, directly violating the 1-hour late-arrival requirement.

63
MCQmedium

A company wants to use Cloud Data Fusion for ETL pipelines. They need to integrate with custom transformations not available in the marketplace. What should they do?

A.Switch to Dataproc and write a Spark job.
B.Use the Data Fusion Hub to download a custom plugin.
C.Use Dataprep to create the transformation.
D.Write a custom plugin using the CDAP SDK and deploy it.
AnswerD

The CDAP SDK lets developers author custom transformations as plugins, package them, and deploy into Cloud Data Fusion, extending the pipeline beyond marketplace offerings. This satisfies the requirement for transformations unavailable in the marketplace, which built-in operators and existing plugins cannot supply.

Why this answer

Cloud Data Fusion allows extending its capabilities by writing custom plugins using the CDAP SDK, which can then be deployed to the Data Fusion instance. This enables integration of transformations not available in the marketplace, providing full flexibility for custom ETL logic.

Exam trap

PDE often tests the extensibility of Cloud Data Fusion, and candidates might incorrectly assume that the Data Fusion Hub provides custom plugins or that other GCP services like Dataprep can be used interchangeably.

How to eliminate wrong answers

Option A is wrong because switching to Dataproc and writing a Spark job would abandon Cloud Data Fusion entirely, which is not necessary and would require re-architecting the pipeline. Option B is wrong because the Data Fusion Hub is a marketplace for pre-built plugins; it does not offer custom plugin downloads. Option C is wrong because Dataprep is a separate data preparation tool that does not integrate custom transformations into Data Fusion pipelines.

64
MCQeasy

A data engineer needs to process streaming data from thousands of IoT devices and generate real-time dashboards. The data volume is low but requires exactly-once processing semantics. Which Google Cloud service combination should they use?

A.Cloud Pub/Sub + Cloud Data Fusion
B.Cloud Pub/Sub + Cloud Dataproc
C.Cloud Pub/Sub + Cloud Dataflow
D.Cloud Pub/Sub + Cloud Dataprep
AnswerC

Cloud Pub/Sub ingests the device streams, while Dataflow provides exactly-once processing through its streaming engine's state and deduplication. This satisfies the stem's low-volume, real-time dashboard requirement, since Dataflow's windowing and triggers emit results continuously rather than only on batch completion.

Why this answer

Dataflow supports exactly-once processing via its streaming engine and checkpointing. Pub/Sub is the ingest service. Together they provide the required semantics.

65
Multi-Selecthard

Your company runs a Dataflow streaming pipeline that processes user activity from Pub/Sub and writes aggregated results to BigQuery. Lately, the pipeline is experiencing high latency and backlog growth during peak hours. You need to troubleshoot and improve performance. Which THREE actions should you take? (Choose 3.)

Select 3 answers
A.Change the worker machine type to a higher CPU/memory configuration
B.Decrease the window duration to reduce data per window
C.Enable Dataflow Streaming Engine
D.Increase the number of workers in the pipeline
E.Add additional Pub/Sub subscriptions to the same topic
AnswersA, C, D

Scaling worker machine type raises per-worker CPU and memory, directly relieving the compute saturation causing backlog growth at peak. Dataflow autoscaling adds workers, but each remains bound by its machine's vCPU and memory limits, so upgrading the type increases throughput per worker for CPU-bound aggregation stages.

Why this answer

Option A is correct because moving to a higher CPU/memory worker machine type gives each worker more compute and memory to process elements, which directly addresses CPU or memory saturation that causes high latency and backlog growth during peak hours. Option C is correct because enabling Dataflow Streaming Engine moves pipeline state and windowing execution off the worker VMs to the Dataflow service, reducing worker CPU/memory pressure and improving streaming performance and autoscaling responsiveness. Option D is correct because increasing the number of workers adds horizontal processing capacity, allowing the pipeline to drain the Pub/Sub backlog faster and keep up with peak throughput.

Option B is not appropriate because shortening the window duration does not reduce the total volume of data processed and can actually increase per-window overhead and output frequency, potentially worsening latency. Option E is not appropriate because adding more subscriptions to the same topic does not increase the throughput of the existing pipeline; each subscription would need its own reader/pipeline, and extra subscriptions only add fan-out consumers rather than improving the current pipeline's performance.

Exam trap

The trap is assuming that more Pub/Sub subscriptions or smaller windows improve throughput — both actually increase overhead or duplicate data rather than scaling the pipeline.

66
MCQmedium

You need to create a BigQuery table that stores customer transaction data. The table will be queried frequently by a customer_id column to retrieve recent transactions (last 30 days). Which table design optimizes query performance and cost?

A.Partition by customer_id and cluster by transaction_date
B.Partition by ingestion_time and cluster by customer_id
C.Cluster by transaction_date and customer_id without partitioning
D.Partition by transaction_date and cluster by customer_id
AnswerD

Partitioning on transaction_date lets queries prune to the last 30 days, and clustering on customer_id sorts data within each partition so lookups scan far fewer bytes. Together they cut both query latency and bytes-billed cost for the stated access pattern.

Why this answer

Partitioning by transaction_date aligns with the query filter (last 30 days), so BigQuery prunes all partitions outside that window, scanning only the relevant date range. Clustering by customer_id then sorts data within each partition, so the customer_id filter benefits from block-level pruning. This combination minimizes both bytes scanned (cost) and query latency.

Exam trap

PDE often tests the misconception that the most selective column (customer_id) should be the partition key, when the correct rule is that the partition column must match the time-based filter predicate to enable pruning.

How to eliminate wrong answers

Option A is wrong because partitioning by customer_id creates one partition per customer, which explodes partition count (BigQuery's practical limit is ~4,000 partitions per table) and does not help filter by date range. Option B is wrong because partitioning by ingestion_time does not match the transaction_date filter, so no partition pruning occurs for 'last 30 days' queries. Option C is wrong because clustering alone provides no partition pruning — the entire table is scanned and only block-level pruning applies, which is far less efficient than partitioning plus clustering.

67
MCQmedium

You are designing a streaming pipeline that needs to handle sudden spikes in traffic without losing data. The pipeline uses Pub/Sub and Dataflow. Which configuration ensures data is not lost if Dataflow falls behind?

A.Use Pub/Sub with a pull subscription and set the message retention duration to 7 days
B.Use Cloud Pub/Sub Lite with a smaller retention period
C.Use Pub/Sub with a push subscription and increase the acknowledgment deadline
D.Use Pub/Sub with exactly-once delivery and Dataflow with at-least-once processing
AnswerA

A pull subscription with seven-day retention keeps unacknowledged messages durably stored while Dataflow's backlog grows during spikes, so no data is dropped. This satisfies the stem's no-loss constraint when the pipeline falls behind, since push subscriptions cannot replay expired messages.

Why this answer

A Pub/Sub pull subscription with a 7-day message retention duration ensures that if Dataflow falls behind, messages are retained in the subscription backlog for up to 7 days, preventing data loss. Dataflow pulls messages and acknowledges them only after processing, so unacknowledged messages remain available for redelivery. This configuration decouples ingestion from processing and provides a buffer for spikes.

Exam trap

The trap is assuming that exactly-once delivery or push subscriptions prevent data loss, when actually retention and pull-based backpressure are what protect against Dataflow lag.

How to eliminate wrong answers

Option B is wrong because Pub/Sub Lite has a smaller retention period (up to 7 days but typically less) and is not designed for the same durability guarantees; it also requires manual capacity management. Option C is wrong because push subscriptions push messages to an endpoint, and increasing the acknowledgment deadline only extends the time before redelivery, but if Dataflow is overwhelmed, messages may still be lost or the endpoint may fail. Option D is wrong because exactly-once delivery in Pub/Sub does not guarantee no data loss if Dataflow falls behind; it only prevents duplicates, and at-least-once processing in Dataflow may still drop messages if the pipeline is overwhelmed.

68
MCQeasy

A financial services firm needs to design a batch processing system on Google Cloud to analyze large volumes of historical transaction data stored in Cloud Storage. The data is in Parquet format and must be processed using Apache Spark. The firm wants to minimize operational overhead and only pay for the resources used during job execution. Which Google Cloud service should they use?

A.Cloud Dataflow with a custom Apache Spark runner
B.Cloud Dataproc Serverless for Spark
C.BigQuery with external tables over Cloud Storage
D.Cloud Dataproc with a long-running cluster
AnswerB

Cloud Dataproc Serverless for Spark is a fully managed, serverless Spark environment that eliminates the need to provision or manage clusters. It automatically scales resources based on workload and charges only for the resources consumed during job execution. It natively supports Spark jobs and can read Parquet data from Cloud Storage, making it the ideal choice for minimizing operational overhead and paying only for what is used.

Why this answer

Cloud Dataproc Serverless for Spark provides a fully managed Spark environment without cluster provisioning. It charges only for the resources used during job execution, aligning with the cost and operational requirements. It supports reading Parquet from Cloud Storage and is designed for batch processing workloads, making it the best fit for analyzing historical transaction data with Spark while minimizing overhead.

Exam trap

The trap here is confusing Dataflow with a Spark execution service or assuming that Dataproc requires a persistent cluster, when serverless Spark is available.

69
MCQmedium

Your team is migrating an on-premises Apache Hadoop cluster to Google Cloud. The cluster runs MapReduce jobs that read and write data to HDFS. You want to minimize code changes and operational overhead. Which Google Cloud service should you use to run these jobs?

A.BigQuery
B.Cloud Composer
C.Cloud Dataproc
D.Cloud Dataflow
AnswerC

Cloud Dataproc is a fully managed service for running Apache Hadoop and Spark clusters. It supports running existing MapReduce jobs with minimal changes, as it provides HDFS-compatible storage via the cluster's local HDFS or Cloud Storage connector. Dataproc handles cluster provisioning, scaling, and management, reducing operational overhead while allowing you to lift and shift your Hadoop workloads.

Why this answer

Cloud Dataproc is designed for running Hadoop and Spark workloads in Google Cloud. It provides a managed cluster with Hadoop, Spark, and Hive, and can use Cloud Storage as a replacement for HDFS via the Cloud Storage connector. Existing MapReduce jobs can often run with little modification, and Dataproc handles cluster management, making it the best choice for migrating on-premises Hadoop workloads with minimal code changes and operational overhead.

Exam trap

The trap here is assuming that any data processing service can run Hadoop jobs, or that Dataflow can execute MapReduce code without rewriting.

70
MCQhard

A healthcare company is designing a system to ingest HL7 messages from multiple hospitals into Google Cloud. The messages must be processed in near real-time to extract patient vitals and trigger alerts if thresholds are exceeded. The system must guarantee that no messages are lost and that processing is exactly-once. Which combination of Google Cloud services should they use?

A.Cloud Pub/Sub with a pull subscription and a custom application on Compute Engine that uses the Pub/Sub client library
B.Cloud Pub/Sub with a pull subscription and Cloud Dataflow with exactly-once processing
C.Cloud Pub/Sub with a push subscription to a Cloud Function that writes to Cloud Firestore
D.Cloud Pub/Sub with a push subscription to Cloud Run, which then publishes to a second Pub/Sub topic for Dataflow
AnswerB

Cloud Pub/Sub provides reliable, scalable message ingestion with at-least-once delivery. When combined with Cloud Dataflow, which supports exactly-once processing semantics for streaming pipelines, this ensures that each message is processed exactly once. Dataflow can handle near real-time processing, extract vitals, and trigger alerts. This combination meets the requirements for no message loss and exactly-once processing in a near real-time system.

Why this answer

Cloud Pub/Sub with a pull subscription ensures reliable message delivery, and Cloud Dataflow supports exactly-once processing for streaming pipelines. Together, they provide a managed solution that guarantees no message loss and exactly-once semantics. Dataflow can process HL7 messages in near real-time, extract vitals, and trigger alerts.

This combination is the most robust and operationally efficient way to meet the strict processing guarantees required by the healthcare scenario.

Exam trap

The trap here is assuming that any message processing service can achieve exactly-once semantics without considering the specific guarantees of the processing engine.

71
MCQeasy

Your data engineering team needs to process a continuous stream of clickstream events from a website and update a real-time dashboard showing user activity over the last hour. The pipeline should have minimal operational overhead and support exactly-once processing semantics. Which Google Cloud service should you use?

A.Cloud Dataproc with Apache Spark Streaming
B.Cloud Data Fusion with batch pipelines
C.Cloud Dataflow with Apache Beam
D.Cloud Pub/Sub Lite with push subscriptions
AnswerC

Cloud Dataflow with Apache Beam provides serverless, autoscaling stream processing with exactly-once semantics via its streaming engine, satisfying the minimal operational overhead and exactly-once constraints. Beam's windowing and triggers handle the rolling one-hour dashboard aggregation over continuous clickstream events.

Why this answer

Cloud Dataflow with Apache Beam is Google's fully managed, serverless stream and batch processing service, and Beam provides built-in exactly-once semantics for streaming pipelines. It integrates natively with Pub/Sub for ingestion and supports windowing, triggers, and state for real-time aggregations like a rolling 1-hour user activity view. Because Dataflow is fully managed, it meets the 'minimal operational overhead' requirement without cluster management.

Exam trap

The trap is equating 'streaming' with any managed service — candidates pick Dataproc or Pub/Sub Lite because they sound real-time, but only Dataflow provides managed exactly-once stream processing with windowing.

How to eliminate wrong answers

Option A is wrong because Dataproc with Spark Streaming requires managing a cluster (sizing, scaling, patching) and does not provide exactly-once semantics out of the box — it typically offers at-least-once and needs idempotent sinks. Option B is wrong because Cloud Data Fusion batch pipelines are designed for batch ETL, not continuous real-time streaming dashboards. Option D is wrong because Pub/Sub Lite is a messaging service, not a processing engine — it can deliver messages but cannot compute windowed aggregations or maintain exactly-once processing state.

72
MCQeasy

Which Google Cloud service provides a serverless Spark environment where you can run Spark jobs without provisioning or managing a cluster?

A.Dataflow
B.Dataproc Serverless
C.Dataprep
D.Cloud Data Fusion
AnswerB

Dataproc Serverless runs Spark workloads on managed, ephemeral infrastructure, so no cluster provisioning, sizing or teardown is required. This satisfies the serverless constraint directly, unlike Dataproc on Compute Engine, which requires you to create and manage clusters yourself.

Why this answer

Dataproc Serverless is a Google Cloud service that allows you to run Spark jobs without provisioning or managing a cluster. It automatically scales resources and charges only for the duration of the job, making it ideal for serverless Spark workloads. This matches the requirement exactly.

Exam trap

PDE often tests the distinction between serverless Spark (Dataproc Serverless) and other serverless data services like Dataflow, causing candidates to confuse the underlying processing engines.

How to eliminate wrong answers

Option A is wrong because Dataflow is a serverless service for Apache Beam, not Spark; it is used for stream and batch data processing but does not run Spark jobs. Option C is wrong because Dataprep is a data preparation tool for visual exploration and transformation, not a Spark execution environment. Option D is wrong because Cloud Data Fusion is a fully managed data integration service for building ETL pipelines, but it does not provide a serverless Spark environment for running arbitrary Spark jobs.

73
Multi-Selectmedium

A company is evaluating BigQuery for a data warehouse migration. They have a mix of reporting queries and ad-hoc analytical queries. They want to control query costs and prevent runaway queries. Which THREE strategies should they implement?

Select 3 answers
A.Grant authorized view access to limit data visibility
B.Set a custom quota for concurrent queries
C.Partition and cluster tables to reduce bytes processed
D.Create materialized views for all reporting queries
E.Use BigQuery reservations (flex slots) for predictable workloads
AnswersB, C, E

A custom quota capping concurrent queries prevents runaway ad-hoc queries from monopolising slots and exhausting capacity. It satisfies the cost-control requirement by throttling query concurrency, so a single user or workload cannot overwhelm the project's shared BigQuery resources.

Why this answer

Option B is correct because setting a custom quota for concurrent queries (via BigQuery custom quotas in Cloud Console/IAM) caps the number of simultaneously running queries per project or user, preventing a flood of ad-hoc queries from exhausting slots and causing runaway costs. Option C is correct because partitioning (e.g., by ingestion/date column) and clustering (e.g., by frequently filtered columns) prune the data scanned, directly reducing bytes processed and therefore on-demand query cost. Option E is correct because BigQuery reservations with flex slots (or committed slots) let predictable reporting workloads run on dedicated capacity with fixed pricing, isolating them from unpredictable ad-hoc on-demand costs and giving cost predictability.

Option A is not correct because authorized views control data access/visibility, not query cost or runaway query prevention. Option D is not correct because materialized views can accelerate specific recurring queries but do not by themselves control costs or prevent runaway queries across a mixed workload.

Exam trap

PDE often tests the difference between cost-control mechanisms (quotas, partitioning, reservations) and access/performance mechanisms (authorized views, materialized views), causing candidates to pick visibility or caching features as cost controls.

74
MCQmedium

A company wants to use BigQuery materialized views to accelerate queries on a table that is updated every hour. Which statement about materialized views is true?

A.Materialized views cannot be clustered.
B.Materialized views must be manually refreshed by the user.
C.Materialized views can only be created on ingestion-time partitioned tables.
D.Materialized views are automatically updated when the base table changes.
AnswerD

BigQuery materialized views refresh automatically, but not instantly on every base-table change. They use a configurable refresh interval, with a default of 30 minutes, so hourly updates are covered. This satisfies the stem's hourly-update constraint without manual intervention, unlike logical views, which recompute on every query.

Why this answer

BigQuery materialized views are automatically refreshed when the base table changes, within a bounded staleness window, so queries against them return fresh results without manual intervention. This automatic refresh is a core property that makes them suitable for accelerating queries on hourly-updated tables.

Exam trap

PDE often tests materialized view refresh behavior and restrictions, so the trap is assuming manual refresh is required or that materialized views cannot be clustered or built on standard partitioned tables.

How to eliminate wrong answers

Option A is wrong because materialized views can be clustered, which further improves query performance and reduces bytes scanned. Option B is wrong because materialized views refresh automatically; manual refresh is not required (though you can trigger one). Option C is wrong because materialized views can be created on partitioned tables generally, not only ingestion-time partitioned tables, and can also be built on non-partitioned tables.

75
MCQeasy

You need to choose a messaging service for a real-time streaming application that requires low cost and can tolerate occasional message loss. Which service is MOST suitable?

A.Cloud Scheduler
B.Pub/Sub Lite
C.Pub/Sub
D.Cloud Tasks
AnswerB

Pub/Sub Lite provides zonal, low-cost messaging with no replication across zones, so it tolerates occasional message loss while meeting the low-cost constraint. Standard Pub/Sub replicates messages, incurring higher cost for durability the scenario does not require.

Why this answer

Pub/Sub Lite is designed for high-volume, low-cost streaming where occasional message loss is acceptable. Unlike standard Pub/Sub, it uses zonal or regional storage with pre-provisioned capacity, which dramatically lowers cost but does not guarantee the same durability or at-least-once delivery semantics under all failure conditions. This makes it the right fit when cost is prioritized over guaranteed delivery.

Exam trap

PDE often tests the distinction between Pub/Sub and Pub/Sub Lite by emphasizing cost tolerance for message loss — candidates incorrectly default to standard Pub/Sub for 'streaming' without weighing the durability/cost trade-off.

How to eliminate wrong answers

Option A is wrong because Cloud Scheduler is a cron-style job trigger, not a messaging/streaming service and cannot handle real-time message throughput. Option C is wrong because standard Pub/Sub provides strong durability and at-least-once delivery with replication, which is more expensive than needed when message loss is tolerable. Option D is wrong because Cloud Tasks is a task queue for asynchronous HTTP callbacks and single-consumer work items, not a high-throughput streaming messaging system.

Page 1 of 2 · 100 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Designing Data Processing Systems questions.