Courseiva

CCNA Designing Data Processing Systems Questions

25 of 100 questions · Page 2/2 · Designing Data Processing Systems · Answers revealed

76
MCQeasy

You are migrating on-premises Hadoop jobs to Google Cloud. The existing jobs use Spark for ETL and Hive for querying. You want to minimize changes to the existing code and maintain the ability to use Hive queries with the same metastore across multiple clusters. Which service combination should you use?

A.Cloud Dataflow with Beam SQL
B.Cloud Dataproc with Dataproc on GKE
C.Cloud BigQuery with external tables on Cloud Storage
D.Cloud Dataproc with Cloud Storage and Dataproc Metastore
AnswerD

Dataproc Metastore provides a managed Hive metastore service compatible with the Hive Metastore API, so existing Hive queries run unchanged and share metadata across multiple Dataproc clusters. Cloud Storage replaces HDFS as the storage layer, satisfying the requirement to minimise code changes during migration.

Why this answer

Cloud Dataproc runs managed Spark and Hive clusters, so existing Spark ETL jobs and Hive queries migrate with minimal code changes. Pairing Dataproc with Cloud Storage for data and Dataproc Metastore for a shared Hive metastore across clusters preserves the same table definitions and schema across multiple clusters.

Exam trap

The trap is picking a modern serverless option (Dataflow or BigQuery) that requires code rewrites, when the requirement explicitly says 'minimize changes' and 'same metastore across multiple clusters' — only Dataproc plus Dataproc Metastore satisfies both.

How to eliminate wrong answers

Option A is wrong because Dataflow with Beam SQL requires rewriting Spark/Hive logic into Beam pipelines and does not provide a Hive metastore. Option B is wrong because Dataproc on GKE changes the runtime model and does not by itself provide a shared Hive metastore across clusters. Option C is wrong because BigQuery with external tables replaces Hive semantics and requires rewriting queries, and it does not preserve a Hive metastore.

77
MCQeasy

You need to design a data processing system that ingests streaming data from thousands of IoT devices. The data must be processed in real-time to calculate average temperature per device over 1-minute intervals, and the results should be stored in BigQuery for analysis. You want a serverless solution with minimal management. Which combination of Google Cloud services should you use?

A.Cloud Pub/Sub for ingestion, Cloud Functions for processing, and BigQuery for storage.
B.Cloud Pub/Sub for ingestion, Cloud Dataflow for processing, and BigQuery for storage.
C.Cloud Pub/Sub for ingestion, Cloud Dataproc for processing, and Cloud Bigtable for storage.
D.Cloud IoT Core for ingestion, Cloud Dataproc for processing, and BigQuery for storage.
AnswerB

Pub/Sub is a scalable, serverless messaging service for ingesting streaming data. Dataflow is a serverless, fully managed service for stream processing that can compute 1-minute averages using windowing. BigQuery is a serverless data warehouse for storing and analyzing results. This combination requires minimal management and is ideal for real-time IoT analytics.

Why this answer

The scenario requires serverless ingestion, real-time processing with windowing, and analytical storage. Pub/Sub handles ingestion at scale, Dataflow provides serverless stream processing with windowing to compute 1-minute averages, and BigQuery stores results for analysis. This combination is fully managed and requires minimal operational effort.

Exam trap

The trap here is selecting services that are not serverless (like Dataproc) or not suited for stream processing (like Cloud Functions), or using deprecated services like Cloud IoT Core.

78
Multi-Selectmedium

A company uses Cloud Pub/Sub for event ingestion. They want to ensure that if a subscriber fails to process a message after 5 attempts, the message is sent to a dead letter topic for analysis. Which TWO configurations are needed?

Select 2 answers
A.Set max delivery attempts to 5 on the subscription.
B.Set the subscription's ack deadline to 600 seconds.
C.Enable message ordering on the subscription.
D.Create a dead letter topic and attach it to the subscription.
E.Set the subscription type to push.
AnswersA, D

Setting max delivery attempts to 5 on the subscription triggers dead-lettering once a message exceeds that threshold. Cloud Pub/Sub requires both this retry limit and a dead letter topic on the same subscription; without the limit, failed messages retry indefinitely, so this satisfies the "after 5 attempts" constraint in the stem.

Why this answer

Option A is correct because Pub/Sub's dead letter policy requires you to specify a maxDeliveryAttempts value on the subscription; setting it to 5 ensures the message is forwarded to the dead letter topic after exactly 5 failed delivery attempts. Option D is correct because a dead letter policy only takes effect when a separate dead letter topic exists and is attached to the subscription via the deadLetterPolicy.deadLetterTopic field, which is where unprocessable messages are published for analysis. Option B is not needed because the ack deadline (e.g., 600 seconds) controls how long a subscriber has to acknowledge a message before redelivery, not the number of attempts before dead-lettering.

Option C is unrelated because message ordering only preserves publish order for ordering keys and does not trigger dead-letter routing. Option E is incorrect because push versus pull is a delivery mode choice and has no bearing on the dead letter policy, which works with either subscription type.

79
MCQmedium

An organization runs periodic Apache Spark jobs on Dataproc to process data from Cloud Storage. They want to reduce costs by using preemptible instances for worker nodes. What is a key consideration when using preemptible instances in Dataproc?

A.Preemptible instances cannot be used with the standard cluster mode
B.Jobs must be designed to handle node preemption, and overall job runtime may increase
C.Preemptible instances are only available in certain regions
D.Jobs will automatically restart from the last checkpoint without any performance impact
AnswerB

Preemptible nodes can be reclaimed at any time, so Spark jobs must tolerate losing workers mid-execution, and recomputation plus retries typically lengthen total runtime. This matches the stem's cost-reduction goal while acknowledging the resilience and duration trade-offs preemptible instances impose.

Why this answer

Preemptible VMs in Dataproc can be reclaimed by Compute Engine at any time (with a 30-second shutdown notice), so Spark jobs must be designed to tolerate worker loss — for example, by using checkpointing, retries, or resilient data pipelines. Because preempted nodes are removed and replacements take time to spin up, overall job runtime typically increases compared to a cluster of standard VMs. This trade-off is the key operational consideration when choosing preemptible workers for cost savings.

Exam trap

PDE often tests the misconception that preemptible instances are transparent to the job — candidates must recognize that preemption causes task re-execution and longer runtimes, and that checkpointing is a developer responsibility, not an automatic Dataproc feature.

How to eliminate wrong answers

Option A is wrong because preemptible instances can absolutely be used with Dataproc standard clusters — in fact, Dataproc explicitly supports mixing preemptible workers with standard workers in a standard cluster (they cannot be the sole worker type in some configurations, but they are supported). Option C is wrong because preemptible VMs are available in most Compute Engine regions and zones, not restricted to a small subset — availability is broad, though capacity can vary. Option D is wrong because jobs do not automatically restart from the last checkpoint without performance impact; checkpointing must be explicitly implemented by the developer, and even then, restarting and re-processing increases runtime and resource consumption.

80
MCQmedium

You need to analyse streaming data from thousands of IoT devices, each sending temperature readings every second. You want to calculate the average temperature per device over the last 5 minutes, updating every minute. Which windowing strategy should you use in Dataflow?

A.Sliding windows of length 5 minutes with a period of 1 minute
B.Global windows with a trigger firing every minute
C.Fixed windows of 5 minutes
D.Session windows with a gap duration of 1 minute
AnswerA

Sliding windows of five minutes with a one-minute period recompute overlapping five-minute averages every minute, so each device's rolling average refreshes continuously. This matches the required five-minute lookback and one-minute update cadence, which fixed or session windows cannot provide.

Why this answer

Sliding windows of length 5 minutes with a period of 1 minute produce a new window every minute, each covering the previous 5 minutes of data — exactly matching the requirement to compute a 5-minute average that updates every minute. This is the canonical use case for sliding windows in Dataflow/Beam.

Exam trap

PDE often tests the confusion between sliding and fixed windows — candidates must recognize that 'updating every minute over the last 5 minutes' requires overlapping windows (sliding), not non-overlapping fixed windows or global windows with triggers.

How to eliminate wrong answers

Option B is wrong because global windows encompass the entire unbounded dataset and never close on their own; a trigger firing every minute would emit cumulative results over all data seen so far, not a 5-minute rolling average. Option C is wrong because fixed (tumbling) windows of 5 minutes are non-overlapping and emit only once every 5 minutes, so the average would not update every minute. Option D is wrong because session windows are defined by gaps in activity — with IoT devices sending every second, sessions would rarely close, and the gap duration of 1 minute does not produce a rolling 5-minute average.

81
MCQhard

A financial services firm runs a batch risk calculation on Dataproc. The job reads from Cloud Storage, processes data in memory, and writes results to BigQuery. The job must complete within a 2-hour window each night, and the cluster must be shut down automatically after completion to minimize cost. You want to orchestrate this with minimal operational overhead. What should you do?

A.Create a Dataproc workflow template that runs the job and deletes the cluster upon completion, and trigger it with a Cloud Scheduler job that calls the Dataproc API.
B.Submit the job to an ephemeral Dataproc cluster created with the gcloud dataproc clusters create command, and use a Cloud Function triggered by Cloud Scheduler to delete the cluster after a fixed 2-hour delay.
C.Use a persistent Dataproc cluster and submit the job via a cron job on the master node, relying on the cluster's idle timeout to shut it down.
D.Use Cloud Composer to create a Dataproc cluster, submit the job, and delete the cluster in a DAG, scheduling it daily.
AnswerA

Dataproc workflow templates can define a job and a cluster that is automatically deleted after the workflow finishes, which meets the auto-shutdown requirement. Cloud Scheduler can invoke the Dataproc API on a cron schedule, providing orchestration with minimal operational overhead. This combination directly addresses both the scheduling and cost cleanup needs.

Why this answer

Dataproc workflow templates natively support running a job on a cluster that is deleted automatically when the workflow completes, satisfying the auto-shutdown and cost requirements. Triggering the template with Cloud Scheduler provides a simple, serverless schedule. Persistent clusters, Composer, or fixed-delay deletion add cost or operational complexity without the same guarantee of completion-based cleanup.

Exam trap

The trap here is assuming that any scheduler plus a cluster deletion script is equivalent, when only a workflow template provides native completion-based cluster deletion.

82
MCQmedium

A company uses Dataproc Serverless for Spark batch jobs. They notice that some jobs are failing due to out-of-memory (OOM) errors. Which configuration parameter should they adjust to allocate more memory per executor?

A.Use a custom image with more memory
B.Set spark.driver.memory to a higher value
C.Increase the number of workers by setting --num-workers
D.Set spark.executor.memory to a higher value, e.g., 8g
AnswerD

spark.executor.memory controls the heap size allocated to each Spark executor JVM. Raising it to 8g gives executors more working memory, directly addressing the OOM failures caused by insufficient per-executor heap during Dataproc Serverless batch execution.

Why this answer

OOM errors in Spark executors are resolved by increasing executor memory, which is controlled by spark.executor.memory. Setting it to a higher value such as 8g allocates more heap to each executor, allowing it to process larger partitions without running out of memory. This is the direct and correct configuration for the reported symptom.

Exam trap

PDE often tests the confusion between driver memory and executor memory — candidates pick spark.driver.memory because it sounds like 'the memory setting' without distinguishing which process is failing.

How to eliminate wrong answers

Option A is wrong because a custom image with more memory does not change the JVM heap allocation for executors — memory must be configured via Spark properties, not the container image. Option B is wrong because spark.driver.memory affects the driver process, not executors; OOM in executors will not be fixed by increasing driver memory. Option C is wrong because increasing the number of workers adds more executors but does not increase memory per executor — if each executor still OOMs, more of them will also fail.

83
MCQeasy

A startup wants to analyze user clickstream data stored in Cloud Storage in Parquet format. They need to run ad-hoc SQL queries without managing any servers and want to pay only for the queries they run. Which Google Cloud service should they use?

A.BigQuery
B.Dataproc
C.Cloud Bigtable
D.Cloud SQL
AnswerA

BigQuery is a serverless, highly scalable data warehouse that supports querying external data in Cloud Storage via external tables. It charges based on the amount of data scanned per query, aligning with pay-per-query. It requires no infrastructure management, making it ideal for ad-hoc SQL analysis.

Why this answer

BigQuery provides serverless SQL analytics with pay-per-query pricing and can query external Parquet data in Cloud Storage. It requires no infrastructure management, perfectly matching the startup's needs. Other services either require provisioning, are not designed for SQL analytics, or lack external data querying.

Exam trap

The trap here is assuming that any data processing service can query Parquet files, but only BigQuery offers serverless SQL with external table support.

84
MCQeasy

A developer wants to create a BigQuery table that automatically expires data older than 30 days to reduce storage costs. Which table design feature should be used?

A.Authorized view
B.Clustered table
C.Materialized view
D.Partitioned table with partition expiration
AnswerD

Partition expiration automatically deletes partitions once their partition date passes the specified retention period, so data older than 30 days is removed without manual jobs. This satisfies the automatic expiry requirement while keeping recent data queryable, unlike table-level expiration which drops the entire table.

Why this answer

BigQuery partitioned tables support partition expiration, which automatically deletes partitions older than a specified number of days. Setting a partition expiration of 30 days on a time-partitioned table will drop data older than 30 days, reducing storage costs. This is the native, cost-effective way to enforce data retention.

Exam trap

PDE often tests the difference between partitioning, clustering, and expiration features, and candidates may confuse clustering with retention or pick materialized views for cost reduction.

How to eliminate wrong answers

Option A is wrong because an authorized view controls access to data, not retention or expiration. Option B is wrong because clustering only sorts data within partitions to improve query performance; it does not expire data. Option C is wrong because a materialized view precomputes query results for performance, not for data lifecycle management.

85
MCQmedium

A company uses BigQuery for analytics. They have a table that is queried frequently by date range. To reduce costs, they want to ensure queries only scan the relevant partitions. They also want to improve performance for queries filtering on a specific customer_id. Which table design should they use?

A.Partition by ingestion time and cluster by customer_id
B.Use a materialized view that filters by date and customer_id
C.Cluster by date column and partition by customer_id
D.Partition by date column and cluster by customer_id
AnswerD

Partitioning on the date column restricts each query to the relevant date range, cutting bytes scanned and cost. Clustering by customer_id then sorts data within each partition, so filters on that column skip irrelevant blocks, satisfying both the cost and performance constraints.

Why this answer

To reduce costs by ensuring queries only scan relevant partitions, you should partition by the date column (so date-range filters prune partitions). To improve performance for queries filtering on customer_id, you should cluster by customer_id (so BigQuery can skip blocks within partitions). Therefore, partition by date column and cluster by customer_id is the correct design.

Exam trap

PDE often tests the distinction between partitioning and clustering — candidates may reverse them or choose ingestion-time partitioning when a business date column is the filter, leading to unnecessary data scans.

How to eliminate wrong answers

Option A is wrong because partitioning by ingestion time does not align with date-range queries on a business date column; ingestion time may not match the query filter, leading to full scans. Option B is wrong because a materialized view can improve performance but does not change the underlying table design for cost reduction; it also adds storage cost and does not address clustering for customer_id. Option C is wrong because clustering by date and partitioning by customer_id reverses the optimal design — partitioning by customer_id would create too many partitions and date-range queries would not benefit from partition pruning.

86
MCQmedium

A retail company ingests point-of-sale events from thousands of stores into Cloud Pub/Sub. They need to process these events in a streaming Dataflow pipeline that enriches each event with store metadata from a slowly changing BigQuery table. The enrichment table is updated only once per day. The pipeline must minimize latency and avoid querying BigQuery for every event. Which approach should they use?

A.Use a CoGroupByKey transform to join the streaming events with a bounded read of the metadata table.
B.Load the BigQuery metadata table into a side input as a key-value map, refreshing it periodically via a scheduled pipeline.
C.Configure the pipeline to call the BigQuery Storage Read API for each event to fetch the latest metadata.
D.Use BigQueryIO.Read with a query that joins the events to the metadata table, and apply the join in the pipeline.
AnswerB

Side inputs in Dataflow allow you to supply additional data to each element in a pipeline without querying an external system per element. By loading the slowly changing metadata into a side input and refreshing it periodically, the pipeline can enrich events with minimal latency and without hitting BigQuery for each event. This matches the requirement to minimize latency and avoid per-event queries.

Why this answer

The correct approach is to load the slowly changing metadata into a side input that is refreshed periodically. This allows the streaming pipeline to enrich each event without querying BigQuery per event, keeping latency low. Side inputs are a standard Dataflow pattern for enriching streaming data with reference data that changes infrequently, and they avoid the high cost and latency of per-element external calls.

Exam trap

The trap here is assuming that a streaming pipeline must query BigQuery for each event or use a join transform, when side inputs are designed for exactly this kind of enrichment.

87
Multi-Selectmedium

A data team is building a near-real-time dashboard that displays aggregated metrics from Kafka topics. They want to use Pub/Sub as a managed messaging service and Dataflow for stream processing. They need to ingest data from Kafka into Pub/Sub with minimal custom code. Which THREE Google Cloud services should they use together? (Choose three.)

Select 3 answers
A.Dataflow
B.Pub/Sub
C.Kafka Connect (with Pub/Sub connector)
D.Cloud NAT
E.Cloud Functions
AnswersA, B, C

Dataflow provides the managed Apache Beam runtime that reads from Kafka and writes into Pub/Sub, satisfying the minimal-custom-code constraint. Its built-in Kafka-to-Pub/Sub template performs the ingestion without bespoke connectors, then continues stream processing for the dashboard's aggregated metrics.

Why this answer

Option A, Dataflow, is correct because it is Google Cloud's managed Apache Beam service for stream processing, and it can run a streaming pipeline that reads from Pub/Sub, performs windowed aggregations, and writes results for the near-real-time dashboard. Option B, Pub/Sub, is correct because it serves as the managed messaging service that decouples the Kafka ingestion layer from the Dataflow processing layer, buffering messages and enabling reliable, scalable delivery. Option C, Kafka Connect (with Pub/Sub connector), is correct because Kafka Connect provides a configuration-driven, low-code way to move data from Kafka topics into Pub/Sub using a Pub/Sub sink connector, satisfying the requirement for minimal custom code.

Option D, Cloud NAT, is not relevant because it provides outbound internet address translation for private VMs and does not ingest Kafka data into Pub/Sub. Option E, Cloud Functions, is not appropriate here because it is an event-driven serverless compute service for lightweight functions, not the managed stream-processing engine needed for aggregated Kafka metrics.

88
MCQeasy

A startup wants to build a data lake on Google Cloud to store raw JSON, CSV, and Parquet files from various sources. They need a storage solution that is highly durable, globally accessible, and integrates natively with BigQuery and Dataproc. They want to minimize management overhead. Which Google Cloud service should they use?

A.Filestore
B.Bigtable
C.Cloud Storage
D.Cloud SQL
AnswerC

Cloud Storage is a highly durable object store that integrates natively with BigQuery external tables and Dataproc. It requires no provisioning and scales automatically, making it ideal for a data lake with minimal management. It supports all mentioned file formats and is globally accessible.

Why this answer

Cloud Storage is the foundational object store for data lakes on Google Cloud. It offers high durability, global accessibility, and native integration with analytics services like BigQuery and Dataproc. It requires no capacity planning or server management, aligning with the goal of minimal overhead.

Relational, NoSQL, and NFS services are not suited for raw file storage at scale.

Exam trap

The trap here is assuming that any storage service can serve as a data lake; only object storage like Cloud Storage provides the necessary scale, durability, and native analytics integration.

89
MCQeasy

Which Google Cloud service provides a fully managed, serverless Spark environment without requiring cluster provisioning?

A.Dataproc on GKE
B.Dataflow
C.Dataproc Serverless
D.Cloud Data Fusion
AnswerC

Dataproc Serverless runs Spark workloads without any cluster provisioning, directly satisfying the stem's serverless requirement. Unlike standard Dataproc, which needs manual cluster creation and sizing, it provisions ephemeral compute automatically per job, so no infrastructure management is needed. This makes it the fully managed Spark environment the question describes.

Why this answer

Dataproc Serverless is a fully managed, serverless Spark environment on Google Cloud that eliminates the need to provision or manage clusters. It automatically scales resources and charges only for the duration of the workload, making it ideal for running Spark jobs without infrastructure overhead.

Exam trap

PDE often tests the distinction between serverless and managed services, and candidates may confuse Dataflow (Beam) with Dataproc Serverless (Spark) or think Dataproc on GKE is serverless when it still requires cluster management.

How to eliminate wrong answers

Option A is wrong because Dataproc on GKE requires managing a Kubernetes cluster, which involves cluster provisioning and configuration, not serverless. Option B is wrong because Dataflow is a fully managed service for Apache Beam, not Spark; it is designed for stream and batch processing but does not run Spark jobs. Option D is wrong because Cloud Data Fusion is a fully managed data integration service for building ETL pipelines, but it is not a serverless Spark environment; it may use Dataproc clusters under the hood but requires provisioning.

90
MCQeasy

You need to process a large Spark ML training job on a Dataproc cluster. The job is fault-tolerant and can handle occasional node failures. To reduce costs, which type of worker nodes should you use?

A.Preemptible worker nodes
B.Standard worker nodes
C.High-memory worker nodes
D.Sole-tenant nodes
AnswerA

Preemptible workers cost far less than standard VMs but can be reclaimed at any time. Because the job tolerates occasional node loss, Dataproc simply reschedules the lost tasks, so the fault-tolerance constraint is met while compute spend drops substantially.

Why this answer

Option A is correct because preemptible worker nodes in Dataproc are significantly cheaper than standard VMs but can be reclaimed by Compute Engine at any time. Since the Spark ML job is fault-tolerant and can handle occasional node failures, preemptible workers are the ideal cost-saving choice. Dataproc automatically replaces preempted workers, and Spark's lineage-based recomputation allows the job to continue without manual intervention.

Exam trap

PDE often tests the assumption that preemptible nodes are unsuitable for any production job — candidates miss that fault-tolerant, stateless workloads are exactly the intended use case.

How to eliminate wrong answers

Option B is wrong because standard worker nodes are billed at full on-demand rates and provide no cost advantage for a fault-tolerant workload. Option C is wrong because high-memory nodes cost more than standard nodes and are only justified for memory-intensive workloads, not for general cost reduction. Option D is wrong because sole-tenant nodes are dedicated physical hosts used for compliance/licensing isolation and are far more expensive than preemptible or standard nodes.

91
MCQmedium

You are designing a Dataflow pipeline that joins two unbounded PCollections from different sources. Which transform should you use?

A.ParDo
B.Flatten
C.CoGroupByKey
D.GroupByKey
AnswerC

CoGroupByKey performs a relational join of two or more PCollections sharing a common key type, emitting grouped values per key. It is the designated transform for joining unbounded PCollections, unlike side inputs or per-element lookups that cannot correlate streams.

Why this answer

CoGroupByKey performs a key-based join of multiple PCollections. It can handle unbounded streams with appropriate windowing.

92
MCQeasy

Which BigQuery feature allows you to share query results with specific users without giving them direct access to the underlying tables?

A.IAM roles
B.Authorized views
C.Dataset access controls
D.Materialized views
AnswerB

Authorized views run with the view owner's permissions, so grantees query the view and see results without any direct table access. This satisfies the requirement to share query results with specific users while withholding underlying table permissions.

Why this answer

Authorized views allow sharing results without granting access to the base tables.

93
MCQmedium

A logistics company collects GPS pings from delivery vehicles into Pub/Sub and needs to compute the distance traveled per vehicle per hour. The data volume is high and bursty, and the company wants a managed service that automatically scales the number of workers based on load while allowing custom windowing and stateful processing. Which service should they use?

A.Cloud Functions triggered by Pub/Sub messages
B.Dataflow with a streaming pipeline using windowing and stateful DoFn
C.Dataproc with a Spark Streaming job on a fixed-size cluster
D.BigQuery with a scheduled query that reads from a Pub/Sub subscription
AnswerB

Dataflow is a managed, autoscaling service for stream and batch processing. It supports windowing, triggers, and stateful DoFn, which are exactly the primitives needed to compute per-vehicle hourly distance while handling bursty load. The service scales workers automatically, matching the requirement for managed elasticity and custom stateful processing.

Why this answer

Dataflow provides managed autoscaling and native support for windowing and stateful processing, which are required to compute per-vehicle hourly distance from a bursty stream. The other options either lack stateful windowing, do not scale automatically, or are not designed for continuous stream processing.

Exam trap

The trap here is assuming that any service that can read Pub/Sub can also perform windowed, stateful aggregation, when only a stream processing engine like Dataflow provides those primitives natively.

94
MCQhard

A financial services firm must process payment events in strict order per account and cannot tolerate duplicates. The events arrive in Pub/Sub and must be written to BigQuery. The engineering team is designing the pipeline and wants to guarantee that each account's events are applied in the order they were published. Which approach should they take?

A.Write all events to a Cloud Storage bucket and run a nightly batch job that sorts by timestamp before loading into BigQuery.
B.Use a single Pub/Sub subscriber with one thread to process all events sequentially across all accounts.
C.Use a Pub/Sub pull subscription with multiple subscribers and rely on BigQuery to sort events by timestamp during queries.
D.Enable message ordering on the Pub/Sub subscription and use the ordering key set to the account ID, then process with a Dataflow pipeline that preserves order per key.
AnswerD

Pub/Sub message ordering with an ordering key ensures that messages with the same key are delivered in publish order. Setting the key to the account ID gives per-account ordering, and a Dataflow pipeline that respects the key can apply events in sequence. This directly satisfies the strict per-account ordering requirement.

Why this answer

Pub/Sub ordering keys combined with a key-aware Dataflow pipeline enforce per-account order without sacrificing parallelism across accounts. Sorting after the fact or using a single global consumer either fails to guarantee application order or imposes unacceptable throughput limits.

Exam trap

The trap here is thinking that timestamps alone can reconstruct publish order, when ordering keys are what Pub/Sub actually uses to sequence messages per key.

95
MCQeasy

A media company needs to process a large number of small JSON files stored in Cloud Storage. They want to use a serverless, SQL-based approach to transform and aggregate the data without managing infrastructure. Which Google Cloud service should they use?

A.Cloud Dataproc
B.Cloud Dataflow
C.Cloud SQL
D.BigQuery
AnswerD

BigQuery is a serverless, highly scalable data warehouse that supports SQL queries. It can directly query external data in Cloud Storage using external tables or load JSON files. This allows the company to transform and aggregate data using SQL without managing infrastructure. It is the ideal service for this requirement.

Why this answer

BigQuery is a serverless, SQL-based analytics service that can query data directly from Cloud Storage, including JSON files, using external tables or loading jobs. It eliminates infrastructure management and allows the media company to transform and aggregate data using familiar SQL. This makes it the most suitable choice among the options.

Exam trap

The trap here is confusing serverless with code-free SQL; Dataflow is serverless but requires coding, while BigQuery is both serverless and SQL-based.

96
MCQmedium

Your company ingests millions of events per second into a Pub/Sub topic. The downstream consumer must process events with minimal latency and high throughput. However, the consumer occasionally falls behind during traffic spikes, and you need to ensure no data loss while minimizing costs. Which subscription type and configuration should you choose?

A.Push subscription with a load balancer
B.Pull subscription with flow control settings
C.Push subscription with endpoint on Cloud Run
D.Pull subscription with exactly-once delivery disabled
AnswerB

A pull subscription lets the consumer control retrieval rate, and flow control settings cap outstanding messages and bytes so the subscriber is not overwhelmed during spikes. Pub/Sub retains unacknowledged messages, so throttled consumption prevents data loss while avoiding the cost of over-provisioned resources.

Why this answer

A pull subscription with flow control settings allows the consumer to control the rate of message delivery, preventing overload during traffic spikes. Pull subscriptions are ideal for high-throughput, low-latency processing because the consumer can batch and process messages at its own pace. Flow control settings (e.g., max outstanding messages, max bytes) help avoid overwhelming the consumer, ensuring no data loss while minimizing costs.

Exam trap

PDE often tests the misconception that push subscriptions are always better for low latency, but pull with flow control provides better backpressure and cost control for high-throughput, spiky workloads.

How to eliminate wrong answers

Option A is wrong because a push subscription with a load balancer adds complexity and does not inherently provide flow control; push subscriptions push messages to an endpoint, which can overwhelm the consumer during spikes. Option C is wrong because a push subscription to Cloud Run may have cold starts and concurrency limits, and it lacks fine-grained flow control, risking data loss or high latency. Option D is wrong because disabling exactly-once delivery does not address the need for flow control; it may reduce costs but does not prevent the consumer from falling behind, and exactly-once is not the primary concern here.

97
Multi-Selecteasy

Your team is using Cloud Dataprep to clean and transform a dataset. Which TWO features of Cloud Dataprep help you understand data quality issues before running the pipeline? (Choose 2.)

Select 2 answers
A.Scheduling data quality jobs
B.Column histograms
C.Joining datasets
D.Recipe steps
E.Data quality profiling
AnswersB, E

Column histograms render the value distribution for each column, exposing skew, outliers, nulls and unexpected cardinality before the pipeline runs. This satisfies the requirement to understand data quality issues during profiling, letting you spot malformed or dominant values that would otherwise corrupt downstream transformation output.

Why this answer

Option B, column histograms, is correct because Cloud Dataprep generates interactive histograms for each column that visually reveal value distributions, outliers, nulls, and anomalies, letting you spot data quality issues before executing a pipeline. Option E, data quality profiling, is correct because Cloud Dataprep's profiling feature automatically scans the dataset and reports metrics such as missing values, distinct counts, mismatched types, and invalid values, which is exactly the pre-pipeline assessment of data quality described. Option A, scheduling data quality jobs, is not correct because scheduling controls when transformation jobs run, not how you inspect data quality beforehand.

Option C, joining datasets, is not correct because joins are transformation operations that combine data rather than diagnose quality issues. Option D, recipe steps, is not correct because recipe steps define the transformations to apply, not the profiling or histogram analysis used to understand data quality first.

Exam trap

PDE often tests the confusion between transformation features (joins, recipe steps) and diagnostic features (histograms, profiling), so candidates must distinguish understanding data from transforming it.

98
MCQmedium

Your team is designing a data processing system that ingests JSON messages from millions of IoT devices. The ingestion rate is highly variable, with spikes up to 500,000 messages per second. You need a fully managed, serverless messaging service that can buffer messages and decouple producers from consumers. Which Google Cloud service should you choose?

A.Cloud Bigtable
B.Cloud Dataflow
C.Cloud Dataproc
D.Cloud Pub/Sub
AnswerD

Cloud Pub/Sub is a fully managed, serverless messaging service designed for high-throughput, variable workloads. It automatically scales to handle millions of messages per second, buffers messages durably, and decouples producers from consumers. It integrates natively with Dataflow and other GCP services, making it ideal for this IoT ingestion scenario.

Why this answer

Cloud Pub/Sub is the correct choice because it is a serverless, fully managed messaging service that automatically scales to handle variable ingestion rates, buffers messages durably, and decouples producers from consumers. It is purpose-built for high-throughput event ingestion and integrates seamlessly with downstream processing services like Dataflow.

Exam trap

The trap here is confusing a processing service like Dataflow with a messaging service like Pub/Sub, overlooking that decoupling and buffering are messaging concerns.

99
MCQmedium

A data engineer needs to run an existing Spark job on Google Cloud with minimal code changes. The job requires Hive metastore access. Which Dataproc feature should they use to provide a managed Hive metastore?

A.Cloud SQL for MySQL
B.Dataproc Metastore
C.BigQuery as a Hive metastore
D.Dataproc on GKE
AnswerB

Dataproc Metastore is a fully managed, Hive-compatible metastore service that existing Spark jobs connect to via the standard Hive metastore interface, so no code changes are needed. It satisfies the managed Hive metastore requirement directly, unlike cluster-local or self-managed alternatives.

Why this answer

Dataproc Metastore is a fully managed, highly available Hive metastore service (based on Hive Metastore 2.3/3.1) that can be attached to Dataproc clusters, providing a persistent, serverless metastore without running a separate Hive metastore on a cluster. It allows existing Spark jobs that require Hive metastore access to run with minimal code changes, since the metastore endpoint is configured via cluster properties.

Exam trap

PDE often tests the confusion between a backing database (Cloud SQL) and a managed metastore service (Dataproc Metastore); candidates pick Cloud SQL thinking MySQL is the Hive metastore, missing that the managed service is the correct answer.

How to eliminate wrong answers

Option A (Cloud SQL for MySQL) is wrong because while Hive metastore can technically use MySQL as a backing RDBMS, Cloud SQL alone is not a managed Hive metastore service — you would have to install and manage Hive metastore yourself. Option C (BigQuery as a Hive metastore) is wrong because BigQuery is an analytics data warehouse, not a Hive metastore implementation; it does not expose the Thrift metastore API that Spark/Hive clients require. Option D (Dataproc on GKE) is wrong because it is a deployment model for running Dataproc workloads on Kubernetes, not a managed Hive metastore feature.

100
MCQhard

A logistics company ingests GPS telemetry from delivery vehicles into Pub/Sub. They need to process the stream in Dataflow to calculate real-time estimated arrival times (ETAs). The pipeline must handle late-arriving data up to 2 hours and must emit results every 5 minutes. The team wants to use Apache Beam's windowing and triggering. Which windowing strategy and trigger should they use to meet these requirements?

A.Fixed windows of 5 minutes with an AfterWatermark trigger that includes a late firing with a 2-hour allowed lateness.
B.Sliding windows of 10 minutes every 5 minutes with an early trigger that fires when the watermark passes the end of the window.
C.Fixed windows of 5 minutes with a default trigger, allowing late data to be dropped after the watermark passes.
D.Session windows with a gap duration of 5 minutes and a repeating trigger every 5 minutes.
AnswerA

Fixed windows of 5 minutes align with the desired emission interval. The AfterWatermark trigger with a late firing allows the pipeline to emit results when the watermark passes (on-time) and again when late data arrives, up to the allowed lateness of 2 hours. This precisely handles late-arriving data and emits updates every 5 minutes, meeting both requirements.

Why this answer

The requirement is to emit results every 5 minutes and handle late data up to 2 hours. Fixed windows of 5 minutes provide the regular emission cadence. The AfterWatermark trigger with a late firing ensures that on-time results are emitted when the watermark passes, and late data triggers additional firings.

Setting allowed lateness to 2 hours ensures that late data is not dropped prematurely. This combination satisfies both the timing and lateness requirements.

Exam trap

The trap here is choosing sliding or session windows for a regular emission interval, or forgetting to configure a late trigger and allowed lateness, which would drop late data.

← PreviousPage 2 of 2 · 100 questions total

Ready to test yourself?

Try a timed practice session using only Designing Data Processing Systems questions.