Courseiva

Google Professional Data Engineer (PDE) — Questions 226–300

747 questions total · 10pages · All types, answers revealed

Page 3

Page 4 of 10

Page 5
226
MCQhard

You manage a Cloud Dataflow streaming pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline uses the BigQueryIO write transform with STREAMING_INSERTS. You need to ensure exactly-once processing semantics for the BigQuery writes. What should you do?

A.Add a BigQuery insertId to each record and rely on BigQuery's built-in deduplication for streaming inserts.
B.Use BigQueryIO.Write with withMethod(STORAGE_WRITE_API) and set withNumStorageWriteApiStreams to a value greater than zero.
C.Switch to BigQueryIO.Write with withMethod(FILE_LOADS) and set withTriggeringFrequency to a low value.
D.Enable exactly-once by setting the pipeline option --experiments=use_runner_v2 and using BigQueryIO.Write with withMethod(STREAMING_INSERTS).
AnswerB

The Storage Write API supports exactly-once semantics when used with Dataflow. Configuring withMethod(STORAGE_WRITE_API) and specifying one or more streams enables the necessary deduplication and transactional guarantees. This is the recommended approach for exactly-once streaming writes to BigQuery, replacing the older STREAMING_INSERTS method.

Why this answer

To achieve exactly-once processing for BigQuery writes in a Dataflow streaming pipeline, use the BigQuery Storage Write API with withMethod(STORAGE_WRITE_API) and configure at least one stream. This method provides transactional writes and deduplication, ensuring each record is written exactly once. Other methods like STREAMING_INSERTS with insertId offer only best-effort deduplication and cannot guarantee exactly-once semantics.

Exam trap

The trap here is assuming that the insertId field in streaming inserts guarantees exactly-once semantics, when in reality it only provides best-effort deduplication within a limited time window.

227
MCQmedium

A company wants to store backups of on-premises databases in Google Cloud for long-term retention. They need WORM (Write Once, Read Many) compliance and object-level retention policies. What should they use?

A.Firestore backups
B.Cloud Storage with Object Lock
C.BigQuery table snapshots
D.Cloud Storage with retention policies
AnswerD

Retention policies are bucket-level, not per-object. Object Lock is needed for WORM.

Why this answer

Cloud Storage supports WORM (Write Once, Read Many) compliance through bucket retention policies and object retention (object-level retention lock), along with legal holds. A bucket retention policy sets a minimum retention period for all objects in the bucket, while object retention allows you to set or extend retention on individual objects, and legal holds prevent deletion or modification until removed. Together these provide the object-level retention controls required for long-term backup retention and regulatory compliance, making Cloud Storage with retention policies the correct choice.

Exam trap

Google often tests the distinction between bucket-level retention policies and object-level retention controls. Candidates may incorrectly assume a separate 'Object Lock' feature exists, but in Google Cloud, WORM compliance is implemented using Cloud Storage retention policies (bucket-level) combined with object retention and legal holds (object-level).

How to eliminate wrong answers

Option A is wrong because Firestore backups are designed for Firestore databases and do not support WORM compliance or object-level retention policies; they are intended for point-in-time recovery of Firestore data, not for long-term archival with immutable storage. Option C is wrong because BigQuery table snapshots are used for preserving table data at a specific point in time for querying or recovery, but they do not provide WORM compliance or object-level retention policies; they are not a storage service for backup files. Option D is wrong because Cloud Storage with retention policies applies a uniform retention period to all objects in a bucket, but it does not support object-level retention policies; Object Lock is required for granular, per-object retention settings and legal holds.

228
MCQeasy

An online retailer uses BigQuery for analytics. They have a time-series table with 5 billion rows and new data arrives every day. They want to optimize query performance and reduce costs by ensuring that queries scan only the partitions they need. Which table design should they use?

A.Use a table partitioned on the timestamp column.
B.Use a table clustered on the timestamp column.
C.Use a table with no partitioning but use LIMIT in queries.
D.Use a table partitioned by ingestion time with a partition expiration.
AnswerA

Partitioning the table on the timestamp column lets BigQuery prune partitions, so queries filtering by date scan only relevant partitions rather than all 5 billion rows. This cuts bytes processed, improving performance and lowering on-demand query costs.

Why this answer

Partitioning on the timestamp column allows BigQuery to perform partition pruning, so queries with filters on that column only scan the relevant partitions. This directly reduces the amount of data read, lowering both query cost (pay-per-byte) and improving performance. For a 5-billion-row table with daily data arrival, time-unit partitioning is the standard design to meet the stated goals.

Exam trap

Google Cloud often tests the distinction between partitioning (which prunes data at the storage level) and clustering (which only sorts data within a partition or table), leading candidates to mistakenly believe clustering alone can reduce bytes scanned for time-range queries.

How to eliminate wrong answers

Option B is wrong because clustering only sorts data within partitions or within the table, but does not enable partition pruning; without partitioning, queries still scan the entire table unless a filter matches the clustering key, and clustering alone does not reduce the bytes billed to only the needed time range. Option C is wrong because using LIMIT does not reduce the amount of data scanned; BigQuery still reads all bytes from the entire table before applying the LIMIT, so costs remain high and performance is not improved. Option D is wrong because partitioning by ingestion time (using _PARTITIONTIME or _PARTITIONDATE) only works for append-only streaming or load jobs and does not allow querying on an arbitrary timestamp column; also, partition expiration would delete old data automatically, but the requirement is to scan only needed partitions, not to expire them.

229
MCQmedium

A team wants to transfer data from an on-premises Hadoop cluster to Cloud Storage for processing. The cluster is located in a remote area with limited bandwidth. They need to transfer 500 TB of data. Which service should they use?

A.Transfer Appliance
B.BigQuery Data Transfer Service
C.Storage Transfer Service
D.Dataproc with gsutil
AnswerA

Transfer Appliance ships a physical storage device to the remote site, letting the team copy 500 TB locally and return it to Google for upload. This bypasses the limited bandwidth that makes network-based transfer of that volume impractical.

Why this answer

Transfer Appliance is designed for petabyte-scale offline transfers when bandwidth is limited.

230
MCQhard

An organization is implementing a data lake on Google Cloud using Cloud Storage. They need to process both batch and streaming data with a unified pipeline. The team has experience with Apache Beam. Which architecture should they use to minimize operational overhead?

A.Kappa architecture with Cloud Dataflow using the same pipeline for batch and streaming
B.Use Cloud Dataproc for batch and Cloud Dataflow for streaming
C.Lambda architecture with Cloud Dataflow for batch and Cloud Pub/Sub for streaming
D.Use Cloud Data Fusion for both batch and streaming
AnswerA

Cloud Dataflow runs the same Apache Beam pipeline for both bounded and unbounded sources, so one codebase serves batch and streaming. This satisfies the unified-pipeline requirement while remaining fully managed, minimising operational overhead for the Beam-experienced team.

Why this answer

Kappa architecture uses a single streaming pipeline for both batch and streaming, simplifying operations. Dataflow implements Beam and supports both modes.

231
MCQmedium

A data engineer needs to create a BigQuery table that is partitioned by ingestion time and clustered by customer_id and transaction_date. They also want to limit access so that only users from a specific domain can query the table. Which approach should they use?

A.Create the table with partitioning only, then use a materialized view to restrict access
B.Create the table without clustering, use row-level security to filter by domain, and grant access to the table
C.Create the table with partitioning and clustering, then create an authorized view on the table and grant the view access to the domain users
D.Create the table with partitioning and clustering, then grant bigquery.dataViewer to the domain via IAM at the dataset level
AnswerC

Partitioning by ingestion time and clustering by customer_id and transaction_date optimises the table's physical layout. An authorised view then runs with the owner's permissions, so granting domain users access to the view alone satisfies the domain-restriction requirement without exposing the base table.

Why this answer

Authorized views allow sharing query results with specific users/groups without giving direct table access. Clustering and partitioning are defined at table creation. IAM roles at dataset level are too broad.

Row-level security filters rows but doesn't restrict domain.

232
MCQhard

A financial services company needs to explain predictions from a complex ensemble model for regulatory compliance. Which Vertex AI service should they use?

A.Vertex AI Explainable AI
B.Vertex AI Vizier
C.Vertex AI Feature Store
D.Vertex AI Prediction
AnswerA

Vertex AI Explainable AI provides feature attributions, such as Shapley values and sampled Shapley methods, for complex models. These explanations document which features drove each prediction, satisfying the regulatory requirement to justify ensemble model outputs to auditors.

Why this answer

Vertex AI Explainable AI is the correct service because it provides feature attributions and other explainability techniques (e.g., Shapley value approximations, integrated gradients) that help interpret predictions from complex ensemble models. This is essential for regulatory compliance, where the company must demonstrate how input features influence each prediction, ensuring transparency and auditability.

Exam trap

Google Cloud often tests the distinction between services that optimize or deploy models versus those that interpret them, so the trap here is assuming that Vertex AI Prediction includes built-in explainability, when in fact it only serves predictions and requires a separate Explainable AI request for attributions.

How to eliminate wrong answers

Option B (Vertex AI Vizier) is wrong because it is a hyperparameter tuning and optimization service, not designed for explaining model predictions. Option C (Vertex AI Feature Store) is wrong because it serves as a centralized repository for feature management and serving, not for generating post-hoc explanations of model outputs. Option D (Vertex AI Prediction) is wrong because it handles model deployment and online/batch inference requests, but does not natively provide interpretability or attribution explanations for individual predictions.

233
MCQmedium

A company needs to process high-throughput streaming data with low latency. They are considering Cloud Pub/Sub for ingestion and Cloud Dataflow for processing. However, they are concerned about cost. Which alternative to Cloud Pub/Sub would reduce costs while still meeting the throughput requirements?

A.Cloud Pub/Sub with pull subscriptions
B.Cloud Tasks
C.Cloud Pub/Sub Lite
D.Cloud Pub/Sub with push subscriptions
AnswerC

Pub/Sub Lite provides the same high-throughput, low-latency ingestion as Pub/Sub but at a substantially lower cost, because it trades away automatic multi-zone replication and requires you to provision capacity in a specific zone. That trade-off satisfies the throughput requirement while directly addressing the stated cost concern.

Why this answer

Cloud Pub/Sub Lite is a zonal, cost-optimized messaging service designed for high-throughput streaming with predictable, low latency. Unlike standard Pub/Sub, which replicates data across zones and charges for data transfer and storage, Pub/Sub Lite lets you choose a capacity (in MiB/s) and charges a flat hourly rate, significantly reducing costs for steady, high-volume workloads. It integrates with Cloud Dataflow via the Pub/Sub Lite I/O connector, so the processing pipeline remains viable.

Thus, it meets the throughput and latency requirements while addressing cost concerns.

Exam trap

PDE often tests the misconception that pull or push subscriptions are cost-saving alternatives, when they are merely delivery methods for standard Pub/Sub, and confuses Cloud Tasks with a streaming ingestion service.

How to eliminate wrong answers

Option A is wrong because pull subscriptions are a consumption method for standard Pub/Sub, not a separate service; they do not inherently reduce costs compared to push subscriptions and still incur standard Pub/Sub pricing for storage and data transfer. Option B is wrong because Cloud Tasks is a task queue for asynchronous, single-consumer task execution, not a high-throughput streaming ingestion service; it lacks the scalability and ordering guarantees needed for streaming data and would not meet throughput requirements. Option D is wrong because push subscriptions are another consumption method for standard Pub/Sub, not a cost-reducing alternative; they can actually increase costs due to HTTP endpoint overhead and retries, and do not change the underlying Pub/Sub pricing model.

234
MCQmedium

A company wants to ingest data from an on-premises Oracle database into BigQuery in near real-time with minimal latency. The database has a high volume of inserts and updates. Which service should they use?

A.Datastream
B.BigQuery Data Transfer Service
C.Pub/Sub
D.Storage Transfer Service
AnswerA

Datastream provides serverless change data capture (CDC) by reading Oracle redo logs, streaming inserts and updates continuously into BigQuery with sub-minute latency. This directly satisfies the near real-time, high-volume insert-and-update constraint, unlike batch extraction tools that cannot capture ongoing changes.

Why this answer

Datastream is correct because it is a serverless change data capture (CDC) service that replicates data from Oracle to BigQuery in near real-time with minimal latency. It reads the Oracle redo logs (via LogMiner or binary log reader) to capture inserts, updates, and deletes and streams them to BigQuery.

Exam trap

PDE often tests the difference between batch transfer services and CDC services; candidates incorrectly choose BigQuery Data Transfer Service or Pub/Sub for real-time Oracle replication, not realizing Datastream is the managed CDC solution.

How to eliminate wrong answers

Option B is wrong because BigQuery Data Transfer Service is designed for scheduled, batch data transfers from various sources (e.g., SaaS apps, Cloud Storage) and does not support real-time CDC from Oracle. Option C is wrong because Pub/Sub is a messaging service, not a CDC tool; it cannot read Oracle redo logs or replicate database changes without custom code. Option D is wrong because Storage Transfer Service is for moving files between object stores (e.g., S3 to GCS), not for database replication.

235
MCQhard

Refer to the exhibit. A data scientist notices that the evaluation component rarely passes the threshold, causing the pipeline to fail often. What should they do to improve efficiency?

A.Reduce the training dataset size
B.Add a conditional component that only runs evaluation if training metrics are above a certain level
C.Remove the evaluation component
D.Increase the threshold value
AnswerB

A conditional component gates evaluation on training metrics exceeding a threshold, so evaluation runs only when the trained model is worth assessing. This prevents repeated pipeline failures from evaluation on poor models, improving overall pipeline efficiency.

Why this answer

Adding a conditional component that only runs evaluation when training metrics exceed a certain threshold prevents unnecessary evaluation runs on poorly performing models. This reduces pipeline failures by ensuring that evaluation, which may be resource-intensive or prone to failure with low-quality inputs, is only triggered when the model has demonstrated sufficient training performance. This approach optimizes resource usage and pipeline reliability without sacrificing the evaluation step entirely.

Exam trap

Google Cloud often tests the misconception that simply adjusting thresholds or removing components is the solution, when the correct approach is to add conditional logic to gate resource-intensive steps based on upstream quality metrics.

How to eliminate wrong answers

Option A is wrong because reducing the training dataset size would likely degrade model quality and does not address the root cause of evaluation failures; it may even increase variance and instability. Option C is wrong because removing the evaluation component entirely would eliminate the ability to validate model performance, which is critical for ensuring model quality and compliance in production pipelines. Option D is wrong because increasing the threshold value would make it even harder for the evaluation component to pass, exacerbating the failure rate rather than improving efficiency.

236
MCQmedium

A company uses BigQuery ML to create a classification model. The model is used for batch prediction on a weekly basis. After six months, the data distribution shifts, and model accuracy drops. Which approach should the company take to maintain model performance?

A.Use Cloud Dataflow to preprocess the data and then update the model with new features.
B.Perform hyperparameter tuning on the original training data.
C.Apply model quantization to reduce model size and improve inference speed.
D.Schedule automatic retraining of the model using the most recent three months of data.
AnswerD

Scheduled retraining on the most recent three months of data refreshes the model against current distributions, countering the drift that degraded accuracy. This satisfies the requirement to maintain performance after six months of distribution shift, without manual intervention each cycle.

Why this answer

The model's accuracy drop is due to data distribution shift (concept drift). Scheduling automatic retraining using the most recent three months of data ensures the model adapts to the new patterns without manual intervention. BigQuery ML supports scheduled queries and automatic model retraining via the `CREATE OR REPLACE MODEL` statement, making this approach both practical and aligned with MLOps best practices for batch prediction pipelines.

Exam trap

Google Cloud often tests the misconception that hyperparameter tuning or feature engineering alone can fix data drift, when in fact only retraining on fresh data addresses the shift.

How to eliminate wrong answers

Option A is wrong because Cloud Dataflow is a data processing tool, not a solution for retraining; preprocessing and adding new features does not address the distribution shift unless the model is retrained on the new data. Option B is wrong because hyperparameter tuning on the original training data optimizes the model for the old distribution, not the shifted one, and will not recover accuracy. Option C is wrong because model quantization reduces model size and speeds up inference but does not improve accuracy or address data drift; it may even slightly degrade performance.

237
MCQeasy

A data engineer schedules a Cloud Composer 2 environment to run a DAG that triggers a Dataflow batch job every night. The DAG sometimes fails because the Dataflow job takes longer than the default task timeout. The engineer wants the DAG to wait for the Dataflow job to finish rather than timing out. Which change should the engineer make?

A.Add a retry policy to the DAG so failed tasks are retried until the Dataflow job finishes.
B.Increase the DAG's schedule interval so runs start less frequently.
C.Set the Dataflow operator's execution_timeout and deferrable parameters so the task waits for job completion.
D.Move the Dataflow launch into a Cloud Function and call it from the DAG.
AnswerC

The Dataflow operators in Cloud Composer can wait for the job to complete, and setting an appropriate execution_timeout prevents the task from failing early. Using the deferrable mode releases the worker slot while waiting, which is the recommended pattern for long-running jobs in Composer 2. This directly addresses the timeout problem.

Why this answer

Cloud Composer 2 Dataflow operators can block until the job reaches a terminal state, and the deferrable variant frees the worker while waiting. Setting execution_timeout to a value longer than the expected job duration prevents premature task failure. Schedule interval, Cloud Function wrapping, and retry policies do not make a task wait for the underlying Dataflow job.

Exam trap

The trap here is confusing task retries or schedule changes with making the operator wait for an asynchronous job to finish.

238
MCQmedium

A financial services firm stores customer transaction data in BigQuery. Compliance requires that a nightly Cloud Composer DAG verify that the previous day's partition is complete before downstream reporting DAGs run, and that the reporting DAG never start if the verification fails. The engineer wants the dependency expressed inside orchestration rather than by polling from the reporting DAG. What should the engineer do?

A.Create a Cloud Scheduler job that calls the reporting DAG's trigger endpoint only after a Cloud Function confirms verification succeeded.
B.Combine the verification and reporting logic into a single DAG so that task order enforces the dependency.
C.Have the reporting DAG poll the verification table in a loop with a sensor until a success row appears, then proceed.
D.Use a Dataset (Dataset-scheduled DAG) in Cloud Composer: have the verification DAG produce the dataset and schedule the reporting DAG to trigger on it.
AnswerD

Airflow datasets provide a push-based dependency: the verification DAG marks the dataset as updated on success, and the reporting DAG is scheduled to run only when that dataset is updated. This expresses the dependency declaratively in orchestration, satisfies the no-polling requirement, and naturally blocks reporting when verification fails.

Why this answer

Airflow datasets in Cloud Composer create a producer-consumer dependency where the verification DAG's successful completion updates a dataset and the reporting DAG is scheduled on that update. This keeps the dependency inside orchestration, avoids polling, and inherently prevents the reporting DAG from running when verification fails.

Exam trap

The trap here is defaulting to a sensor or an external scheduler for cross-DAG dependencies, when Airflow's dataset scheduling is the push-based mechanism designed for exactly this.

239
MCQeasy

A data engineer needs to orchestrate a series of tasks that include calling external APIs, running BigQuery queries, and sending notifications. The workflow involves conditional branching and parallel steps. Which Google Cloud service should be used?

A.Workflows
B.Cloud Scheduler
C.Cloud Composer
D.Dataflow
AnswerA

Workflows is a fully managed orchestration service whose YAML-based syntax natively supports conditional branching, parallel branches and connector calls to external APIs, BigQuery and Pub/Sub. This directly satisfies the stem's requirement for branching and parallel steps without managing compute infrastructure.

Why this answer

Workflows is the correct choice because it is a fully managed orchestration service designed specifically for coordinating multi-step, event-driven workflows that involve conditional branching, parallel execution, and integration with external APIs, BigQuery, and notifications via HTTP calls or service integrations. It provides built-in error handling, retries, and a declarative YAML-based syntax that directly supports the described requirements without needing to manage infrastructure or schedule tasks.

Exam trap

The trap here is that candidates often confuse Cloud Scheduler as an orchestrator because it can trigger workflows, but it lacks the conditional branching and parallel execution capabilities required for this scenario, while Cloud Composer is mistakenly chosen due to its familiarity with Airflow, despite being heavier than necessary for a simple orchestration task.

How to eliminate wrong answers

Option B is wrong because Cloud Scheduler is a cron-based job scheduler that triggers tasks on a fixed schedule, not an orchestrator for complex workflows with conditional branching and parallel steps. Option C is wrong because Cloud Composer is a managed Apache Airflow service that can orchestrate workflows, but it is overkill for this use case, requires managing DAGs and infrastructure overhead, and is not the simplest or most cost-effective choice for a lightweight, event-driven workflow. Option D is wrong because Dataflow is a stream and batch data processing service based on Apache Beam, designed for transforming and analyzing data pipelines, not for orchestrating tasks like API calls, BigQuery queries, and notifications with conditional logic.

240
Multi-Selectmedium

Which TWO actions should you take to ensure model reliability in a production Vertex AI Endpoint?

Select 2 answers
A.Use only batch predictions to avoid real-time issues
B.Monitor prediction accuracy in production with logging and alerts
C.Disable request/response logging to reduce latency
D.Use a single model endpoint for all traffic
E.Gradually shift traffic to new model versions (canary deployment)
AnswersB, E

Logging requests and responses to BigQuery or Cloud Logging, then alerting on drift or accuracy degradation, detects reliability failures in live traffic. This satisfies the production reliability requirement by surfacing silent model decay before it affects business outcomes, enabling timely retraining or rollback.

Why this answer

Option B is correct because monitoring prediction accuracy in production via request/response logging and Cloud Monitoring alerts is essential to detect model drift, data skew, and degradation, allowing timely retraining or rollback to maintain reliability. Option E is correct because gradually shifting traffic to new model versions using canary deployment on a Vertex AI Endpoint lets you validate the new model against live traffic and roll back quickly if metrics degrade, minimizing risk. Option A is not appropriate because batch predictions cannot serve real-time production traffic and do not by themselves ensure reliability.

Option C is wrong because disabling request/response logging removes the observability needed to detect accuracy issues and violates reliability monitoring. Option D is wrong because routing all traffic to a single model endpoint without versioning or traffic splitting prevents safe canary testing and rollback.

Exam trap

Google Cloud often tests the misconception that disabling logging improves reliability by reducing latency, when in fact it removes the observability needed to detect and diagnose failures, which is a core tenet of MLOps reliability.

241
MCQhard

A financial services company runs a Dataflow streaming pipeline that reads from Pub/Sub and writes enriched records to BigQuery. Compliance requires that the raw Pub/Sub messages be retained for seven years so that any record can be reprocessed if the enrichment logic is later found to be incorrect. The pipeline currently has no archival step. What should the data engineer do to satisfy the retention requirement with the least operational overhead?

A.Configure a Pub/Sub dead-letter topic and set the maximum delivery attempts high enough that failed messages persist for the retention period.
B.Enable BigQuery table snapshots on the destination table and set a seven-year expiration on each snapshot.
C.Add a Cloud Storage sink to the Dataflow pipeline that writes each raw message to a Cloud Storage bucket with a lifecycle policy moving objects to Coldline or Archive storage after 30 days.
D.Increase the Pub/Sub topic message retention duration to the maximum allowed and rely on the subscription backlog for seven years.
AnswerC

Writing the unmodified messages to Cloud Storage preserves the raw payload independently of the BigQuery output, and a lifecycle policy automatically transitions old objects to colder, cheaper classes so seven years of retention stays affordable. The Dataflow sink adds little operational burden because the pipeline already runs continuously and only needs one additional write branch.

Why this answer

Retaining raw messages for years requires a durable, low-cost object store, and Cloud Storage with lifecycle-based class transitions is the standard fit on Google Cloud. By branching the existing Dataflow pipeline to write unmodified messages to a bucket, the company keeps a complete, reprocessable archive without building a separate ingestion path or managing long-lived Pub/Sub backlogs.

Exam trap

The trap here is assuming Pub/Sub retention can be stretched to meet a multi-year compliance window, when its retention is bounded and it is not an archival system.

242
MCQeasy

A data science team needs to ensure that a deployed Vertex AI model can handle varying traffic patterns with minimal latency and cost. What should they do?

A.Use Vertex AI Prediction with autoscaling
B.Use batch prediction instead of online
C.Pre-warm all instances
D.Deploy to a single large machine type
AnswerA

Autoscaling adjusts replica count to match incoming request volume, absorbing traffic spikes while scaling down during quiet periods. This satisfies both constraints in the stem: minimal latency under load and minimal cost when demand drops.

Why this answer

Vertex AI Prediction with autoscaling dynamically adjusts the number of serving instances based on incoming traffic, ensuring minimal latency during spikes and cost efficiency during lulls. This is the recommended approach for handling variable traffic patterns in production, as it leverages Google Cloud's managed infrastructure to scale from zero to thousands of nodes automatically.

Exam trap

Google Cloud often tests the misconception that batch prediction can substitute for online serving in variable traffic scenarios, but the key distinction is that batch prediction lacks real-time latency guarantees and cannot scale dynamically per request.

How to eliminate wrong answers

Option B is wrong because batch prediction is designed for asynchronous, large-scale offline inference on static datasets, not for real-time traffic with varying patterns; it cannot handle low-latency online requests. Option C is wrong because pre-warming all instances defeats the purpose of autoscaling, leading to constant high cost regardless of actual traffic, and is not a dynamic solution. Option D is wrong because deploying to a single large machine type creates a single point of failure and cannot scale horizontally to handle traffic spikes, resulting in either over-provisioning cost or latency under load.

243
MCQhard

A financial services company must ensure that predictions from a deployed model do not become biased against protected groups. They have a monitoring system in place. Which metric should they track?

A.Prediction latency
B.Prediction distribution across demographic segments
C.Per-query input feature distribution
D.Model accuracy over time
AnswerB

Tracking prediction distribution across demographic segments reveals whether outcomes diverge between protected groups, exposing disparate impact. Comparing approval or score rates per segment detects bias that aggregate accuracy metrics would hide, satisfying the fairness monitoring requirement.

Why this answer

Tracking prediction distribution across demographic segments (option B) directly monitors for bias by comparing the model's output rates for different protected groups. If the distribution diverges significantly, it indicates potential disparate impact, which is the core concern for fairness in deployed models. This aligns with monitoring for algorithmic fairness, not just operational performance.

Exam trap

The trap here is that candidates confuse operational metrics (latency, accuracy) with fairness metrics, assuming high accuracy guarantees fairness, but the Google Professional Data Engineer exam tests that bias can exist even with high accuracy if the model performs differently across demographic segments.

How to eliminate wrong answers

Option A is wrong because prediction latency measures the time taken to serve a prediction, which is a performance metric unrelated to bias or fairness against protected groups. Option C is wrong because per-query input feature distribution tracks individual input values, not the aggregated output predictions across demographic segments needed to detect bias. Option D is wrong because model accuracy over time measures overall predictive performance, which can remain high even when the model is biased against a specific group (e.g., accuracy may be high for a majority class while failing for a minority class).

244
MCQhard

A financial services firm runs a Dataflow batch pipeline that joins a 2 TB transaction dataset with a 40 GB customer reference dataset. The reference data changes only once per day and is currently read from a BigQuery table with a side-input transform on every element. Job cost is dominated by repeated BigQuery reads, and the pipeline occasionally hits quota errors. The team wants to minimize cost and quota pressure while keeping the daily refresh. What should the data engineer change?

A.Load the reference dataset into a BigQuery table partitioned by ingestion date and read only the latest partition as the side input.
B.Enable BigQuery BI Engine on the reference table so side-input queries are served from memory instead of disk.
C.Replace the side input with a CoGroupByKey transform that groups transactions and customer records by customer ID before the join.
D.Stage the reference dataset as an Avro file in Cloud Storage, load it into the pipeline with a file-based source, and pass it as a side input refreshed by a scheduled daily export from BigQuery.
AnswerD

Exporting the slowly changing 40 GB reference data once per day to Cloud Storage removes per-element BigQuery reads entirely, and Avro gives compact, schema-carrying storage that Dataflow reads efficiently. The single daily export keeps the reference fresh while eliminating the quota pressure and repeated scan cost that the BigQuery side input caused on every pipeline run.

Why this answer

Moving the slowly changing reference dataset out of BigQuery into a compact Cloud Storage format and refreshing it once daily eliminates read amplification. A single scheduled export replaces thousands of per-element reads, cutting both cost and quota pressure while preserving the daily update cadence. CoGroupByKey, partition pruning on a side input, and BI Engine all leave the repeated BigQuery access pattern intact.

Exam trap

The trap here is believing that a side input is always cheap, when a large side input is re-materialized per worker and re-read from its source.

245
MCQhard

A company uses Vertex AI Feature Store for serving features to both training and prediction. The team notices that predictions made shortly after training use different feature values, causing a training-serving skew. What is the most effective way to prevent this skew?

A.Configure the Feature Store to use point-in-time lookup using the training timestamp
B.Retrain the model more frequently to adapt to the new feature distributions
C.Use batch prediction instead of online prediction to ensure consistent features
D.Ensure that the training and prediction environments use identical compute resources
AnswerA

Point-in-time lookup retrieves feature values as they existed at the training timestamp, rather than the latest values. This aligns serving-time features with those seen during training, directly eliminating the training-serving skew caused by time-dependent feature drift.

Why this answer

Point-in-time lookup ensures that feature values used during training are exactly the same as those used during prediction by retrieving the feature value as it existed at the training timestamp. This directly addresses training-serving skew caused by time-dependent feature changes, which is a common issue in Vertex AI Feature Store when features are updated after training.

Exam trap

A common pitfall is assuming that retraining more frequently or switching prediction methods can resolve training-serving skew. The root cause is temporal inconsistency in feature values, which requires point-in-time lookups to ensure the same feature values are used during training and prediction.

How to eliminate wrong answers

Option B is wrong because retraining more frequently does not prevent the skew; it only reduces the window of time during which stale features are used, but the fundamental mismatch between training-time and prediction-time feature values remains. Option C is wrong because batch prediction does not inherently use consistent features; it still retrieves the latest feature values unless point-in-time lookup is explicitly configured, and it does not solve the skew for online serving scenarios. Option D is wrong because identical compute resources have no impact on feature value consistency; the skew arises from feature value changes over time, not from hardware differences.

246
MCQhard

You are designing a Dataflow pipeline that reads from Pub/Sub and writes to BigQuery. The pipeline uses the Storage Write API with exactly-once semantics. You need to ensure that the pipeline can handle a sudden spike in traffic without losing data or causing duplicates. Which configuration should you adjust to improve throughput while maintaining exactly-once?

A.Increase the number of Storage Write API streams.
B.Increase the maximum number of workers.
C.Enable autoscaling for the Dataflow job.
D.Switch to STREAMING_INSERTS method.
AnswerA

The Storage Write API allows multiple streams to be used concurrently. Increasing the number of streams (via withNumStorageWriteApiStreams) can improve throughput by parallelizing writes. This is a key tuning parameter for high-volume pipelines while preserving exactly-once semantics, as each stream maintains its own offset and deduplication.

Why this answer

The Storage Write API uses a configurable number of streams to write data to BigQuery. Each stream can handle a certain throughput, and increasing the number of streams allows more parallel writes, improving overall throughput. This adjustment maintains exactly-once semantics because each stream is managed with its own offset and deduplication.

Autoscaling and worker count help with processing but do not directly increase the write throughput limit imposed by the number of streams.

Exam trap

The trap here is assuming that scaling the Dataflow workers alone will increase BigQuery write throughput, but the Storage Write API has its own stream-based throughput limit that must be tuned separately.

247
MCQhard

A manufacturing company wants to detect anomalies in sensor data from thousands of IoT devices in real time. The data is streaming into Pub/Sub. The best solution should use a machine learning model served from AI Platform that scores sensor readings aggregated over 5-minute windows. Which pipeline design meets these requirements?

A.Use Cloud Dataproc with Spark Streaming to aggregate data, and use a Spark ML model embedded in the pipeline
B.Use BigQuery streaming inserts and run scheduled queries that call the ML model
C.Use Cloud Dataflow with sliding windows to aggregate sensor readings every 5 minutes, then call a trained model hosted on AI Platform Prediction for each window
D.Use Cloud Functions triggered by Pub/Sub to process each sensor reading individually
AnswerC

Cloud Dataflow sliding windows aggregate the Pub/Sub stream into five-minute intervals, and each window is scored by calling the AI Platform Prediction endpoint. This satisfies both constraints: real-time streaming aggregation and serving the trained model rather than embedding it in the pipeline.

Why this answer

Cloud Dataflow's sliding windows natively handle the 5-minute aggregation requirement for streaming data, and its ability to call external services via a DoFn allows integration with AI Platform Prediction for real-time model scoring. This design aligns with the need for low-latency, scalable processing of Pub/Sub streams without managing infrastructure.

Exam trap

Google Cloud often tests the distinction between stream processing (Dataflow) and batch-oriented services (BigQuery scheduled queries), and the trap here is assuming that BigQuery's streaming inserts combined with scheduled queries can achieve real-time aggregation, when in fact scheduled queries introduce minutes of delay and are not window-aware for sliding time intervals.

How to eliminate wrong answers

Option A is wrong because Cloud Dataproc with Spark Streaming requires managing a cluster and embedding a Spark ML model in the pipeline, which adds operational overhead and does not leverage AI Platform's managed prediction service as specified. Option B is wrong because BigQuery streaming inserts and scheduled queries introduce latency (scheduled queries run at intervals, not in real time) and are not designed for per-window scoring of streaming data. Option D is wrong because Cloud Functions triggered by Pub/Sub process each sensor reading individually, which cannot aggregate data over 5-minute windows as required.

248
Multi-Selecteasy

You are using Cloud Workflows to orchestrate a series of API calls. You need to handle errors and retries. Which THREE features of Cloud Workflows can you use? (Choose THREE.)

Select 3 answers
A.Use try/except blocks to catch and handle errors.
B.Integrate with Cloud Load Balancing for high availability.
C.Use conditional branches (if-else) based on step results.
D.Define a retry policy on a step.
E.Enable automatic logging for each step.
AnswersA, C, D

Cloud Workflows supports try/except syntax within a step, letting you catch a raised error and route execution to a handler instead of failing the whole execution. This directly satisfies the need to handle errors during orchestration of the API call sequence.

Why this answer

Option A is correct because Cloud Workflows supports try/except blocks, allowing you to catch errors raised by a step and execute fallback logic or handle the failure gracefully. Option C is correct because Cloud Workflows supports conditional branching with if/else constructs, letting you evaluate step results and choose different execution paths based on those outcomes. Option D is correct because Cloud Workflows lets you attach a retry policy to a step, specifying parameters such as max_retries, initial_delay, max_delay, and multiplier to automatically retry failed calls.

Option B is not correct because Cloud Load Balancing is a traffic-distribution service for workloads, not a Cloud Workflows error-handling or retry feature. Option E is not correct because logging is handled by Cloud Logging integration and is not a configurable per-step error-handling or retry feature of Cloud Workflows.

Exam trap

The trap is selecting logging or load balancing as error-handling features; candidates must distinguish between observability (logging) and actual error/retry control mechanisms (try/except, retry policies).

249
MCQmedium

You are building a Dataflow pipeline that reads Avro-formatted files from Cloud Storage. The files use a schema that is updated frequently. You want to minimize pipeline restarts due to schema changes. Which approach should you take?

A.Convert all Avro files to CSV before ingestion and use a fixed CSV schema.
B.Use a BigQuery load job with schema autodetect instead of Dataflow.
C.Read the Avro files using a schema inferred at runtime from the Avro file metadata.
D.Use a fixed Avro schema in your pipeline code and manually update it each time the source schema changes.
AnswerC

Avro files embed their schema in the file header, so a Dataflow pipeline can dynamically read that schema and adapt to changes without code changes or restarts. This is the standard approach for handling evolving Avro schemas in a streaming or batch pipeline.

Why this answer

Avro files include their schema in the file header, so a Dataflow pipeline can read that schema at runtime and adapt to changes automatically. This avoids the need to hardcode or manually update schemas, reducing pipeline restarts when the source schema evolves.

Exam trap

The trap here is assuming that Avro schemas must be defined statically in the pipeline code, when in fact Avro's self-describing format allows dynamic schema resolution.

250
MCQhard

A Dataflow streaming pipeline that uses global windows and triggers every 5 seconds is experiencing increasing lag and high system latency. The pipeline reads from Pub/Sub, transforms data with a ParDo, and writes to BigQuery. Which action is most likely to reduce lag?

A.Use a session window to group related events.
B.Replace the global window with a sliding window of 1 minute.
C.Change the trigger to processing time instead of event time.
D.Increase the number of workers manually.
AnswerB

A sliding window reduces the number of elements per trigger and improves latency by distributing state across workers.

Why this answer

B is correct because sliding windows of 1 minute allow the pipeline to process data in overlapping fixed-size windows, which can reduce the buildup of data in memory compared to global windows. Global windows with frequent triggers (every 5 seconds) can cause unbounded state growth and high latency as the pipeline must maintain state for all elements until the trigger fires, whereas sliding windows naturally bound the data per window and enable more efficient watermark and trigger management in Dataflow.

Exam trap

Google Cloud often tests the misconception that increasing workers or changing trigger timing alone can fix lag caused by inappropriate windowing strategy, when the real issue is that global windows with frequent triggers create unbounded state that overwhelms the pipeline's memory and shuffle capacity.

How to eliminate wrong answers

Option A is wrong because session windows group events based on inactivity gaps, which does not address the core issue of unbounded state from global windows and can actually increase state size if sessions are long. Option C is wrong because changing the trigger to processing time instead of event time does not reduce lag; it may cause data to be processed based on when it arrives rather than when it occurred, potentially increasing latency due to watermark misalignment and still requiring global window state. Option D is wrong because manually increasing the number of workers can help with throughput but does not fix the fundamental design flaw of using global windows with frequent triggers, which leads to excessive state accumulation and shuffling; autoscaling in Dataflow already handles worker count based on backlog.

251
MCQhard

A company runs large batch prediction jobs on Vertex AI every day. They want to minimize costs while ensuring the jobs complete within a 4-hour window. The model requires significant memory. What is the most cost-effective approach?

A.Use Cloud TPUs to accelerate predictions
B.Use a smaller machine type (e.g., n1-standard-4) to reduce cost
C.Use preemptible VMs with a machine type that meets memory requirements
D.Use standard VMs and reduce parallelization
AnswerC

Preemptible VMs cost substantially less than standard VMs and suit batch workloads that tolerate interruption. Pairing them with a machine type meeting the model's memory requirement keeps the daily job within the four-hour window at minimal cost.

Why this answer

Preemptible VMs (now called Spot VMs) are significantly cheaper than standard VMs (up to 60-80% discount) and are ideal for fault-tolerant batch prediction jobs that can handle interruptions. Since the job has a 4-hour window and the model requires significant memory, using preemptible VMs with a machine type that meets the memory requirements minimizes cost while allowing the job to complete if restarted within the time limit.

Exam trap

Google Cloud often tests the misconception that preemptible VMs are unreliable for any production workload, but the trap here is that batch prediction jobs are inherently fault-tolerant and can leverage preemptible VMs to drastically reduce costs without violating the completion window.

How to eliminate wrong answers

Option A is wrong because Cloud TPUs are specialized hardware for training and inference of large models, but they are more expensive and not necessary for batch prediction; they also do not directly address the memory requirement or cost minimization for a 4-hour window. Option B is wrong because using a smaller machine type (e.g., n1-standard-4) would likely cause out-of-memory errors or severe performance degradation since the model requires significant memory, making the job fail or exceed the 4-hour window. Option D is wrong because reducing parallelization would increase job duration, potentially exceeding the 4-hour window, and standard VMs are more expensive than preemptible VMs, so this approach does not minimize costs.

252
MCQmedium

A financial company needs to process batch trades data daily and ensure that if a transformation step fails, the entire daily run is retried from the beginning. Which design pattern is appropriate?

A.Use idempotent writes with checkpointing
B.Use an orchestrator like Cloud Composer with retry logic
C.Retry the failed step only
D.Use a transactional staging area
AnswerB

Cloud Composer orchestrates the batch pipeline as a directed acyclic graph, so a failed transformation triggers a configurable retry of the whole daily run from the start. This satisfies the stem's all-or-nothing rerun constraint, which per-task retries or event-driven serverless functions cannot guarantee.

Why this answer

The requirement states that if any transformation step fails, the entire daily run must be retried from the beginning. An orchestrator like Cloud Composer (Apache Airflow) provides native DAG-level retry logic that can be configured to restart the entire workflow on failure, ensuring atomicity of the batch run. This pattern is essential for maintaining data consistency when partial processing cannot be tolerated.

Exam trap

Google Cloud often tests the misconception that checkpointing or idempotent writes are sufficient for full-run retries, but the trap is that checkpointing enables partial resumption, not the complete restart from scratch that the question explicitly demands.

How to eliminate wrong answers

Option A is wrong because idempotent writes with checkpointing allow resumption from the last successful checkpoint, which contradicts the requirement to retry the entire run from the beginning; checkpointing is designed for partial retries, not full restarts. Option C is wrong because retrying only the failed step would leave the daily run in an inconsistent state, as earlier steps may have already committed partial results that cannot be rolled back without a full restart. Option D is wrong because a transactional staging area ensures atomic writes but does not provide orchestration or retry logic to restart the entire pipeline from the start upon failure.

253
MCQhard

Your organization runs a Cloud Composer 2 environment that executes dozens of DAGs. Several DAGs share a connection to an external REST API that enforces a rate limit of 100 requests per minute. During peak hours, DAGs fail with HTTP 429 errors. You want to prevent these failures without changing the external API's limits. What should you do?

A.Increase the number of Celery workers in the Composer environment.
B.Set the DAG's max_active_runs to 1 for each affected DAG.
C.Add exponential backoff retries to the API-calling tasks.
D.Create an Airflow pool with a limited number of slots and assign the affected tasks to that pool.
AnswerD

Airflow pools limit the number of concurrent task instances that can use a set of slots. By creating a pool sized to stay under the API's rate limit and assigning the API-calling tasks to it, you throttle concurrency across all DAGs. This directly prevents exceeding 100 requests per minute without modifying the external service.

Why this answer

Airflow pools provide a shared concurrency limit that spans DAGs. Sizing a pool to keep total in-flight API calls under the provider's rate limit prevents 429 responses. Increasing workers, limiting per-DAG active runs, or relying on retries does not coordinate request volume across the many DAGs that share the external API.

Exam trap

The trap here is treating retries or per-DAG concurrency limits as a substitute for a cross-DAG concurrency control like an Airflow pool.

254
MCQhard

A company needs to store logs in Cloud Storage for compliance, with a requirement that logs cannot be deleted or overwritten for a period of 7 years. Which Cloud Storage feature should they enable?

A.Bucket Lock with a retention policy of 7 years.
B.Requester Pays bucket setting.
C.Versioning enabled with Object Hold.
D.Object lifecycle management with a delete rule after 7 years.
AnswerA

Bucket Lock applies a retention policy that prevents objects from being deleted or overwritten until the specified period elapses, directly enforcing the seven-year compliance requirement. It is the Cloud Storage mechanism designed for immutable, tamper-resistant retention.

Why this answer

Bucket Lock with a retention policy of 7 years is the correct feature because it enforces a WORM (Write Once, Read Many) model on the bucket. Once a retention policy is locked, objects cannot be deleted or overwritten until the retention period expires, meeting the compliance requirement for immutable log storage.

Exam trap

Google often tests the distinction between features that prevent deletion (Bucket Lock) versus features that only track versions or automate cleanup (Versioning, Lifecycle), leading candidates to choose Versioning or Lifecycle rules thinking they enforce immutability.

How to eliminate wrong answers

Option B is wrong because Requester Pays shifts storage costs to the requester but does not prevent deletion or overwriting of objects. Option C is wrong because Versioning enabled with Object Hold can prevent deletion of specific object versions but does not prevent overwriting of the current version, and holds are not a bucket-wide immutable policy. Option D is wrong because Object lifecycle management with a delete rule only automates deletion after a set time but does not prevent manual deletion or overwriting before that time, so it cannot enforce a non-deletion guarantee.

255
MCQhard

You have a Dataflow batch pipeline that processes data from Cloud Storage and writes to BigQuery. The pipeline uses a custom DoFn that sometimes throws exceptions due to malformed input records. You want to ensure that the pipeline continues processing valid records while logging the malformed ones for later analysis, without failing the entire job. Which Dataflow feature should you use?

A.Configure the pipeline to use a dead-letter queue in Pub/Sub.
B.Use a side output to capture and log the malformed records.
C.Set the pipeline's failure mode to 'continue' in the pipeline options.
D.Use a ParDo with a try-catch block and log the exception, then drop the record.
AnswerB

Dataflow supports side outputs, which allow you to route elements that fail processing to a separate PCollection. By catching exceptions in your DoFn and emitting the malformed records to a side output, you can continue processing valid records without failing the pipeline. The side output can then be written to a dead-letter sink, such as Cloud Storage or BigQuery, for later analysis. This is the recommended pattern for handling bad records in Dataflow.

Why this answer

The correct approach is to use a side output to capture malformed records. Side outputs allow you to emit elements that fail processing to a separate PCollection, which can then be written to a dead-letter sink for analysis. This enables the pipeline to continue processing valid records without failing.

Other options either do not exist, drop data, or are not applicable to batch pipelines. Side outputs are a core Dataflow feature for error handling.

Exam trap

The trap here is assuming that simply logging and dropping bad records is sufficient, when the requirement is to preserve them for later analysis, which necessitates a side output to a separate sink.

256
MCQhard

A data pipeline uses Cloud Data Fusion to perform ETL jobs. The pipeline reads from BigQuery, transforms data using Wrangler, and writes to Cloud Storage. The team notices that the pipeline runs slower than expected. They suspect the Data Fusion instance is under-provisioned. Which action should be taken to improve performance?

A.Add more Dataproc Metastore instances
B.Change the Data Fusion instance type from Basic to Enterprise
C.Enable Data Fusion accelerator for BigQuery
D.Rewrite the pipeline using Cloud Dataprep instead
AnswerB

Data Fusion instance type determines the available compute and memory for pipeline execution. Basic instances are limited to a small, fixed profile, so upgrading to Enterprise provides greater resources and the ability to scale executors, directly addressing the suspected under-provisioning causing slow ETL runs.

Why this answer

Cloud Data Fusion instance types (Basic, Enterprise, Developer) determine the compute resources available to run pipelines. Upgrading from Basic to Enterprise increases the number of Dataproc worker nodes and resources available to the CDAP runtime, improving pipeline throughput and reducing runtime. This directly addresses the under-provisioned instance suspicion.

Exam trap

PDE often tests the confusion that adding metadata or auxiliary services (like Dataproc Metastore) improves ETL performance, when the real lever is the Data Fusion instance type and Dataproc cluster sizing.

How to eliminate wrong answers

Option A is wrong because Dataproc Metastore is a metadata service for Hive/Spark catalogs and does not provide compute capacity for Data Fusion pipelines. Option C is wrong because there is no 'Data Fusion accelerator for BigQuery' feature; Data Fusion uses plugins/connectors, not accelerators. Option D is wrong because rewriting in Cloud Dataprep changes the tooling but does not address the under-provisioned Data Fusion instance and Dataprep is more of a data preparation UI than a full ETL runtime.

257
MCQhard

A data scientist developed a model using custom training on Vertex AI. They want to automate the entire training-to-deployment process. Which service should they use?

A.Cloud Composer
B.Vertex AI Pipelines
C.Cloud Build
D.Cloud Functions
AnswerB

Vertex AI Pipelines orchestrates the full ML workflow as a directed acyclic graph, chaining custom training jobs into evaluation and endpoint deployment steps. This automates the entire training-to-deployment process the stem requires, unlike standalone training or model registry components.

Why this answer

Vertex AI Pipelines is the correct choice because it provides a fully managed, serverless orchestration service specifically designed to automate ML workflows, including custom training, hyperparameter tuning, evaluation, and deployment. It integrates natively with Vertex AI services and supports Kubeflow Pipelines SDK or TFX for defining reproducible, end-to-end pipelines, making it the ideal solution for automating the entire training-to-deployment process.

Exam trap

The trap here is that candidates often confuse general-purpose orchestration (Cloud Composer) with ML-specific pipeline orchestration (Vertex AI Pipelines), overlooking that Vertex AI Pipelines provides built-in ML artifact tracking and native integration with Vertex AI training and prediction services.

How to eliminate wrong answers

Option A is wrong because Cloud Composer is a workflow orchestration service based on Apache Airflow, which is more general-purpose and requires custom operators or hooks to interact with Vertex AI, adding unnecessary complexity and not providing native ML pipeline capabilities. Option C is wrong because Cloud Build is a CI/CD service focused on building, testing, and deploying software artifacts (e.g., containers), not on orchestrating ML training workflows or managing model deployment steps like evaluation and versioning. Option D is wrong because Cloud Functions is a serverless compute service for event-driven, short-lived functions, which lacks the state management, sequencing, and artifact tracking needed for multi-step ML pipelines.

258
MCQeasy

Which BigQuery feature allows you to write data with exactly-once semantics, high throughput, and the ability to buffer data before making it available for queries?

A.BigQuery load jobs
B.BigQuery Data Transfer Service
C.Legacy streaming inserts
D.Storage Write API with buffered mode
AnswerD

Storage Write API's buffered mode holds records in a stream until you commit them, delivering exactly-once semantics through stream-level offsets and deduplication. This satisfies the stem's requirement to buffer data before query visibility, while sustaining high throughput — unlike the legacy streaming API, which inserts rows immediately.

Why this answer

The Storage Write API with buffered mode (option D) is correct because it provides exactly-once semantics for data ingestion, high throughput via gRPC streaming, and the ability to buffer data in memory before making it available for queries. This mode allows you to commit rows in a stream, ensuring no duplicates, while the buffering stage gives you control over when data becomes visible in BigQuery.

Exam trap

A common trap is assuming that legacy streaming inserts (option C) provide exactly-once semantics, but they actually offer at-least-once delivery. The Storage Write API with buffered mode is the only correct choice for exactly-once, high-throughput, buffered writes.

How to eliminate wrong answers

Option A is wrong because BigQuery load jobs offer at-least-once semantics (duplicates possible on retry) and do not support buffering before query availability; they write data directly to tables. Option B is wrong because BigQuery Data Transfer Service is a scheduled, managed service for importing data from external sources (e.g., Google Ads, Amazon S3) and does not provide high-throughput streaming or buffered write semantics. Option C is wrong because legacy streaming inserts (tabledata.insertAll) provide at-least-once semantics (duplicates can occur) and lack the buffered mode that defers query visibility; they also have lower throughput and no exactly-once guarantee.

259
MCQeasy

You need to store petabytes of data in a data warehouse that supports ANSI SQL, automatic scaling, and real-time analytics. The data is primarily used for ad-hoc queries and business intelligence. Which Google Cloud service should you use?

A.Cloud SQL
B.Cloud Bigtable
C.Cloud Spanner
D.BigQuery
AnswerD

BigQuery is a fully managed, petabyte-scale data warehouse that supports ANSI SQL and provides automatic scaling and real-time analytics. It is designed for ad-hoc queries and business intelligence workloads, with separation of storage and compute, and integrates with BI tools. It is the ideal choice for this scenario.

Why this answer

BigQuery is the correct choice because it is a serverless, petabyte-scale data warehouse that supports ANSI SQL, automatically scales, and is optimized for ad-hoc queries and business intelligence. Its separation of storage and compute allows independent scaling and cost-effective analytics on large datasets.

Exam trap

The trap here is confusing a transactional database like Cloud Spanner with an analytical data warehouse, overlooking that BigQuery is purpose-built for large-scale SQL analytics.

260
MCQeasy

Which Google Cloud service is a fully managed relational database for MySQL, PostgreSQL, and SQL Server, offering automatic replication and backups?

A.Cloud Spanner
B.AlloyDB
C.Bigtable
D.Cloud SQL
AnswerD

Cloud SQL is Google Cloud's fully managed relational service supporting MySQL, PostgreSQL and SQL Server, with automated replication and backups handled by the platform. This directly satisfies the stem's requirement for a managed engine covering all three database engines without manual administration.

Why this answer

Cloud SQL is the correct answer because it is Google Cloud's fully managed relational database service that supports MySQL, PostgreSQL, and SQL Server. It provides automatic replication across zones and automated backups, making it the ideal choice for traditional relational database workloads without the need for manual administration.

Exam trap

The trap here is that candidates often confuse Cloud SQL with Cloud Spanner because both are relational databases, but Cloud Spanner is designed for global scale and does not support MySQL, PostgreSQL, or SQL Server compatibility.

How to eliminate wrong answers

Option A is wrong because Cloud Spanner is a globally distributed, horizontally scalable relational database service that supports strong consistency and SQL, but it is not a fully managed service for MySQL, PostgreSQL, or SQL Server; it uses its own proprietary SQL dialect and is designed for sharded, multi-region deployments. Option B is wrong because AlloyDB is a fully managed PostgreSQL-compatible database service optimized for high performance and transactional workloads, but it does not support MySQL or SQL Server. Option C is wrong because Bigtable is a fully managed, scalable NoSQL wide-column database service, not a relational database, and it does not support MySQL, PostgreSQL, or SQL Server.

261
MCQmedium

A media analytics company ingests 8 TB of new JSON event logs into BigQuery every day and keeps all history for 5 years. Analysts almost never filter on the raw event payload column and only occasionally select it, but they frequently filter on event_date and user_id. Storage cost is the top concern, and query performance on the frequently filtered columns must stay fast. What should the data engineer do?

A.Store the raw event payload in a BigQuery column of type JSON so it is parsed at query time and billed as a separate data type.
B.Partition the table by event_date and cluster it by user_id, then rely on the default columnar storage to avoid scanning the payload.
C.Enable the BigQuery long-term storage pricing tier so data older than 90 days is billed at a lower rate automatically.
D.Store the raw event payload in a Cloud Storage bucket and keep only the structured, frequently queried columns in the BigQuery table, referencing the object path.
AnswerD

Moving the rarely filtered payload out of BigQuery and into Cloud Storage Standard or Nearline removes the largest column from the columnar table, cutting both active storage cost and bytes scanned for the common queries. The structured columns remain in BigQuery for fast filtering on event_date and user_id, and the object path can be used with external tables or remote functions when the payload is genuinely needed.

Why this answer

The workload filters on structured columns but almost never touches the large payload, so the payload is dead weight in the columnar store. Keeping only the structured columns in BigQuery preserves fast partitioned and clustered filtering, while relocating the payload to Cloud Storage removes the dominant storage and scan cost. Partitioning, clustering, the JSON type, and long-term storage all leave that large column inside the table, so none of them eliminate the core expense.

Exam trap

The trap here is assuming that partitioning, clustering, or long-term storage pricing reduces the cost of a large column that queries still scan, when only removing that column from the table actually does.

262
MCQmedium

A company wants to build a data lake on Cloud Storage for raw, curated, and processed data zones. They need to enforce data governance including column-level security and row-level filtering for BigQuery queries. Which solution should they use?

A.BigLake tables over Cloud Storage
B.BigQuery external tables reading from GCS
C.Dataproc with Spark SQL
D.Cloud Storage with IAM and VPC Service Controls
AnswerA

BigLake tables let BigQuery query Cloud Storage data while applying fine-grained governance: column-level security via policy tags and row-level filtering via row access policies. This satisfies the requirement to enforce governance across raw, curated, and processed zones without copying data into native BigQuery storage.

Why this answer

BigLake tables provide a unified governance layer over Cloud Storage data, enabling fine-grained access control such as column-level security and row-level filtering directly on BigQuery queries. This is achieved by integrating BigQuery's access control policies with the external data stored in GCS, without needing to move data into BigQuery native storage. The other options either lack these granular security features or require complex workarounds.

Exam trap

Google often tests the misconception that BigQuery external tables (Option B) can support the same fine-grained security as BigLake, but they cannot because external tables lack the integrated policy engine for column and row-level controls.

How to eliminate wrong answers

Option B is wrong because BigQuery external tables reading from GCS only support table-level IAM permissions and cannot enforce column-level security or row-level filtering; they treat the external data as a flat file without fine-grained access controls. Option C is wrong because Dataproc with Spark SQL does not natively provide column-level or row-level security on Cloud Storage data; it requires manual implementation via Spark's security APIs and does not integrate with BigQuery's governance model. Option D is wrong because Cloud Storage with IAM and VPC Service Controls only provides bucket- and object-level access control and network perimeter security, but cannot enforce column-level or row-level filtering on queries executed in BigQuery.

263
MCQeasy

A team has trained a model using AutoML Tables. They want to deploy it for batch predictions on a schedule. What is the simplest approach?

A.Write a Cloud Function triggered by Cloud Scheduler
B.Export model to Cloud Storage and use Dataflow
C.Deploy to App Engine
D.Use Vertex AI Batch Prediction with a scheduled pipeline
AnswerD

Vertex AI Batch Prediction accepts an AutoML Tables model and a BigQuery or Cloud Storage input, and wrapping it in a scheduled pipeline automates recurring runs. This satisfies the scheduled batch constraint with minimal custom code.

Why this answer

Vertex AI Batch Prediction is the simplest approach because it is a managed service that directly supports batch predictions on AutoML Tables models without requiring additional infrastructure. By wrapping it in a scheduled Vertex AI pipeline, you can automate the entire workflow—triggering predictions on a schedule, handling input/output to Cloud Storage, and managing compute resources—all within the Vertex AI ecosystem, minimizing operational overhead.

Exam trap

Google Cloud often tests the misconception that you must export an AutoML model to use it outside Vertex AI, but the simplest path is to use Vertex AI's native batch prediction service, which avoids the overhead of custom infrastructure like Dataflow or Cloud Functions.

How to eliminate wrong answers

Option A is wrong because Cloud Functions are designed for lightweight, event-driven tasks and lack native support for AutoML Tables model serving; you would need to manually load the model and handle scaling, which adds complexity and is not the simplest approach. Option B is wrong because exporting the model to Cloud Storage and using Dataflow introduces unnecessary steps—Dataflow requires writing a custom pipeline to load the exported model and perform predictions, whereas Vertex AI Batch Prediction handles this natively. Option C is wrong because App Engine is a platform for hosting web applications, not designed for batch prediction workloads; it would require building a custom prediction service and managing scaling, which is more complex than using Vertex AI's built-in batch prediction.

264
MCQmedium

A logistics company uploads nightly shipment manifests as newline-delimited JSON files into a Cloud Storage bucket. Analysts need to query these files with standard SQL immediately after upload, but the schema evolves frequently and the company wants to avoid managing a load job. Which approach should they use?

A.Create a BigQuery external table over the Cloud Storage bucket using the BigLake connection with JSON format and schema autodetect.
B.Mount the bucket in Dataproc and run a Hive external table over the JSON files.
C.Use the Storage Transfer Service to copy the JSON into a BigQuery dataset, then query the dataset.
D.Load each file into a native BigQuery table with a scheduled query that runs bq load every night.
AnswerA

BigQuery external tables over Cloud Storage let you query files in place with standard SQL, and BigLake connections add governance plus support for schema autodetect on newline-delimited JSON. Because the data stays in Cloud Storage, no load job or pipeline maintenance is required, and new files matching the URI prefix are picked up automatically, which fits a frequently evolving manifest schema.

Why this answer

External tables let BigQuery read files directly from Cloud Storage, so analysts can run standard SQL over the manifests without a load job, and BigLake-backed external tables support schema autodetection for newline-delimited JSON. Because the schema changes often, autodetect combined with in-place querying avoids brittle load pipelines while still exposing the data through the BigQuery SQL surface.

Exam trap

The trap here is assuming that querying files in Cloud Storage requires moving them into BigQuery storage first, when external tables can query them in place.

265
MCQeasy

A data engineer must ensure that a Cloud Composer DAG which loads a BigQuery table runs every day at 02:00 UTC and that a dependent downstream report DAG runs only after the load succeeds. The report DAG lives in the same Composer environment but is a separate DAG file. Which feature should be used to coordinate the two DAGs?

A.A shared XCom written by the load DAG and read by the report DAG using xcom_pull across DAGs.
B.A single trigger_rule of all_success applied to the report DAG's root task.
C.A TimeSensor in the report DAG set to the expected load completion time.
D.An ExternalTaskSensor in the report DAG referencing the load DAG and task, with a matching execution date.
AnswerD

ExternalTaskSensor is designed to wait on a task in a different DAG, which matches the cross-DAG dependency here. With the correct execution_delta or execution_date_fn so the two runs align, the report DAG waits until the load task succeeds before proceeding, giving the required success-based coordination.

Why this answer

ExternalTaskSensor is the purpose-built mechanism for waiting on a task in another DAG. Configuring it with the correct execution date alignment makes the report DAG block until the load task succeeds, which is exactly the cross-DAG dependency described. Time-based and intra-DAG mechanisms cannot observe another DAG's success state.

Exam trap

The trap here is using a time-based wait or an intra-DAG trigger rule when the dependency crosses DAG boundaries.

266
MCQeasy

You are loading a CSV file into BigQuery. The file has a column `price` that contains values like `$1,234.56`. You want to load it into a NUMERIC column. What should you do?

A.Use the `--null_marker` option to treat the dollar sign as null.
B.Specify the schema with `price` as a NUMERIC and set the `--allow_quoted_newlines` option.
C.Preprocess the CSV file to remove the dollar signs and commas before loading, or use a transformation in a data pipeline.
D.Load the data as a STRING and then use a CAST to NUMERIC in a query to create a new table.
AnswerC

BigQuery's load job expects numeric columns to contain plain numbers without currency symbols or thousands separators. The safest approach is to clean the data before loading, either by editing the CSV or using a tool like Cloud Dataflow or Dataprep to transform the values. Alternatively, you can load as STRING and then transform in BigQuery, but preprocessing is often simpler. This ensures the load succeeds without errors.

Why this answer

BigQuery requires numeric columns to contain plain numeric values without formatting characters. To load a CSV with values like `$1,234.56` into a NUMERIC column, you must remove the dollar signs and commas, either by preprocessing the file or using a data transformation step. Options that suggest loading as-is or using unrelated options will fail.

Exam trap

The trap here is assuming BigQuery can automatically parse currency-formatted strings into numbers, when it requires clean numeric input.

267
MCQmedium

A machine learning engineer needs to deploy a custom TensorFlow model for online predictions with low latency. The model is already trained and saved in SavedModel format. Which Vertex AI service should they use?

A.Vertex AI Workbench
B.Vertex AI Prediction
C.Vertex AI Feature Store
D.Vertex AI AutoML
AnswerB

Vertex AI Prediction deploys the SavedModel to an endpoint serving online predictions with low latency, satisfying the stem's requirement. Unlike batch prediction, it provisions a persistent endpoint for real-time inference, and unlike custom training, it consumes the already-trained artefact directly without retraining.

Why this answer

Vertex AI Prediction allows you to deploy custom models (including TensorFlow SavedModel) to an endpoint for online predictions. It supports autoscaling and low-latency serving.

268
MCQmedium

Your team uses dbt to transform data in BigQuery. You need to schedule dbt runs to refresh materialized tables and views every hour. The transformations include both full refreshes and incremental models. What is the most efficient way to orchestrate these dbt runs on Google Cloud?

A.Use Cloud Composer (Airflow) to schedule and run dbt commands.
B.Use Cloud Build with a trigger to run dbt every hour.
C.Use Cloud Scheduler to trigger a Cloud Function that runs dbt.
D.Set up a cron job on a Compute Engine instance to run dbt.
AnswerA

Cloud Composer runs Airflow DAGs that invoke dbt commands, giving dependency management, retries and hourly scheduling across both full-refresh and incremental models. Airflow's operators orchestrate the dbt runs and surface failures, which raw cron or Cloud Scheduler cannot manage as cleanly.

Why this answer

Cloud Composer (Airflow) is the recommended orchestration tool for complex workflows like dbt runs, supporting dependencies, retries, and scheduling. Cloud Scheduler alone cannot run dbt directly; it can trigger a Cloud Function to run dbt, but that is less maintainable. Cloud Build is CI/CD, not scheduling.

Using a cron job on Compute Engine is possible but not managed.

269
MCQeasy

A mobile app needs a NoSQL database that supports offline synchronization when the device goes offline and later reconnects. Which Google Cloud database should be used?

A.Firestore
B.Cloud Bigtable
C.Cloud Spanner
D.Cloud SQL
AnswerA

Firestore is a NoSQL document database with built-in offline persistence: the client SDK caches data locally and synchronises changes automatically when connectivity returns. This directly satisfies the stem's offline synchronisation constraint, unlike Cloud SQL or Bigtable, which lack native mobile offline support.

Why this answer

Firestore is a NoSQL document database that provides built-in offline synchronization for mobile and web apps. It caches data locally and automatically syncs changes when the device reconnects, making it ideal for offline-capable applications.

Exam trap

The trap is confusing Firestore with other NoSQL databases like Bigtable, or assuming that Spanner's global distribution supports offline sync, when only Firestore provides native offline synchronization for mobile apps.

How to eliminate wrong answers

Option B is wrong because Cloud Bigtable is a high-throughput, low-latency NoSQL database for analytics and time-series data, but it does not support offline synchronization for mobile apps. Option C is wrong because Cloud Spanner is a globally distributed relational database with strong consistency, not designed for offline mobile sync. Option D is wrong because Cloud SQL is a managed relational database service, which lacks native offline sync capabilities for mobile devices.

270
Multi-Selecthard

Your company is building a data processing system that ingests sensor data from millions of devices, processes it in near real-time to detect anomalies, and stores raw and processed data for long-term analytics. The system must meet a 99.9% uptime SLA and minimize data loss. Which THREE design choices are best? (Choose three.)

Select 3 answers
A.Use Cloud Pub/Sub as the ingestion layer with a dead-letter topic to capture unprocessed messages.
B.Store raw data in Cloud Bigtable and processed data in Cloud Storage.
C.Use Dataflow with at-least-once processing guarantees and perform deduplication downstream.
D.Use Cloud Storage for raw data archival and BigQuery for processed analytics data.
E.Use a global Cloud Load Balancer in front of the Dataflow workers.
AnswersA, C, D

Cloud Pub/Sub decouples ingestion from processing, absorbing millions of device events while buffering against consumer failures, which supports the 99.9% uptime SLA. Its dead-letter topic captures messages that repeatedly fail processing, preventing silent data loss and satisfying the minimise-data-loss constraint.

Why this answer

Option A is correct because Cloud Pub/Sub provides a highly available, globally distributed ingestion layer that decouples millions of devices from downstream processing, and a dead-letter topic captures messages that repeatedly fail processing so they are not silently lost, supporting the 99.9% uptime and minimal data-loss goals. Option C is correct because Dataflow with at-least-once processing guarantees ensures no message is dropped during failures, and performing deduplication downstream handles the duplicate records that at-least-once semantics can produce, which is the standard trade-off for minimizing data loss in near real-time pipelines. Option D is correct because Cloud Storage is the durable, low-cost archival store for raw sensor data, while BigQuery is the serverless analytics warehouse suited for long-term processed-data analytics at scale.

Option B is not best because Bigtable is optimized for high-throughput random reads/writes of time-series or keyed data rather than long-term analytical querying, and Cloud Storage alone for processed data lacks BigQuery's SQL analytics capability. Option E is not appropriate because a global Cloud Load Balancer distributes HTTP(S)/TCP traffic to backends and is not used to front Dataflow workers, which are managed by the Dataflow service and scale internally.

Exam trap

Google Cloud often tests the misconception that a load balancer is needed to scale Dataflow workers, when in fact Dataflow auto-scales its own workers and uses Pub/Sub's pull subscriptions to distribute messages evenly across workers without a separate load balancer.

271
MCQhard

Your team uses Looker to develop a model on top of BigQuery. The data is partitioned by ingestion time, and analysts frequently query the last 7 days. However, Looker queries are scanning the entire table, causing high costs. Which change should you implement?

A.Enable BI Engine on the BigQuery table to accelerate queries.
B.Apply a LookML access_filter to dynamically filter on the partition column.
C.Create a materialized view that aggregates data daily and point Looker to that view.
D.Use clustering on the order_date column to improve query performance.
AnswerB

Access filters in LookML can be used to restrict queries to a specific partition range, reducing full table scans.

Why this answer

The single best approach is to apply a partition filter requirement in LookML and enable partition pruning in BigQuery. Other options are not directly about Looker or are suboptimal.

Exam trap

The question initially stated 'Pick two' but then contradicted itself, potentially confusing candidates. The correct answer is a single best approach.

272
MCQmedium

A company uses BigQuery for analytics and needs to ensure that certain columns containing PII are encrypted at query time so that only authorized users can decrypt. What should they use?

A.BigQuery AEAD encryption functions
B.VPC Service Controls
C.Customer-managed encryption keys (CMEK)
D.Fine-grained IAM roles
AnswerA

BigQuery AEAD encryption functions let you encrypt specific PII columns at query time using keys managed through Cloud KMS, so only principals holding the key can decrypt. This satisfies the requirement for column-level, query-time encryption rather than storage-level or dataset-level controls.

Why this answer

BigQuery AEAD encryption functions allow you to encrypt sensitive columns (e.g., PII) at query time using a user-managed key, so that only authorized users who possess the key can decrypt the data. This is the correct approach because it provides column-level, application-layer encryption that is transparent to the query engine and ensures that unauthorized users see only ciphertext.

Exam trap

In Google Cloud exams, a common trap is to assume that Customer-managed encryption keys (CMEK) provide column-level, application-layer encryption or query-time decryption control. CMEK only protects data at rest at the storage level, not at query time. The correct approach for column-level query-time encryption is to use BigQuery AEAD encryption functions.

How to eliminate wrong answers

Option B is wrong because VPC Service Controls provide network-level security boundaries to prevent data exfiltration, not column-level encryption at query time. Option C is wrong because Customer-managed encryption keys (CMEK) encrypt data at rest (storage layer), not at query time, and do not control per-user decryption access. Option D is wrong because Fine-grained IAM roles control access to tables or rows via row-level security, but they do not encrypt the data itself; authorized users still see plaintext PII.

273
MCQeasy

A logistics company stores delivery manifests as newline-delimited JSON files in a Cloud Storage bucket. Analysts want to run SQL queries over these files immediately without importing them into a BigQuery dataset or paying for duplicate storage. Which BigQuery capability should they use?

A.Use the BigQuery Storage Write API to stream the files into a dataset.
B.Mount the bucket as a BigQuery dataset using the Cloud Storage connector.
C.Create an external table over the Cloud Storage bucket using the JSON type.
D.Load the JSON files with a scheduled query that runs every hour.
AnswerC

BigQuery external tables can point directly at newline-delimited JSON objects in Cloud Storage, letting analysts run SQL without loading the data into managed storage. This satisfies the requirement to query immediately and avoid duplicate storage costs, since the data stays in the bucket and is read at query time.

Why this answer

External tables let BigQuery read newline-delimited JSON directly from Cloud Storage at query time, so analysts get immediate SQL access without loading data into managed storage. The files remain in the bucket, avoiding duplicate storage charges while still supporting standard SQL over the manifests.

Exam trap

The trap here is reaching for a loading or streaming mechanism when the requirement explicitly forbids duplicating storage and demands immediate access.

274
MCQmedium

A company needs to load data from a MySQL database into BigQuery daily. The data volume is 10 GB per day and the schema changes occasionally. They want to minimize costs and operational overhead. What is the MOST appropriate approach?

A.Use Datastream to stream changes from MySQL to BigQuery
B.Use BigQuery Data Transfer Service for MySQL
C.Export MySQL data to CSV, upload to GCS, and use BigQuery load jobs
D.Use Cloud SQL federated query from BigQuery
AnswerC

This approach requires manual handling of schema changes and daily exports, which is more operational overhead.

Why this answer

For a daily batch load of 10 GB with occasional schema changes, the most cost-effective and low-overhead approach is to export the MySQL data to CSV or Parquet, upload it to Cloud Storage, and use BigQuery load jobs. BigQuery load jobs are free (you only pay for storage and queries), and they can handle schema changes via schema autodetect or by updating the table schema. Datastream is designed for continuous change data capture with low latency, which is unnecessary and likely more expensive for a once-daily batch load.

BigQuery Data Transfer Service does not support MySQL as a source, and Cloud SQL federated queries are for querying external data, not for loading it into BigQuery.

Exam trap

A common misconception is that Datastream is always the best choice for moving data from MySQL to BigQuery. However, Datastream is intended for continuous replication and is not cost-optimal for periodic batch loads. For a daily batch load, the classic extract-load pattern (export to GCS, then BigQuery load job) is more appropriate and cost-effective.

How to eliminate wrong answers

Option B is wrong because BigQuery Data Transfer Service for MySQL is not a supported service; BigQuery Data Transfer Service supports sources like Google Ads, Amazon S3, and Teradata, but not direct MySQL connections. Option C is wrong because exporting MySQL to CSV, uploading to GCS, and using load jobs incurs higher operational overhead (manual scripting, schema management) and does not handle schema changes gracefully, requiring manual intervention for each change. Option D is wrong because Cloud SQL federated queries from BigQuery are designed for ad-hoc querying of live Cloud SQL data, not for daily bulk ingestion, and they do not persist data in BigQuery, leading to repeated query costs and no historical retention.

275
MCQhard

A logistics company streams vehicle telemetry into a Pub/Sub topic. Each message contains a vehicle ID and a timestamp, and the Dataflow pipeline computes per-vehicle distance using a stateful DoFn with a ValueState timer set to fire after 5 minutes of event-time inactivity. During a regional network outage, some vehicles stop sending data for 20 minutes and then resume with correctly ordered timestamps. After the outage, operators notice that some late-arriving records are being dropped before the stateful computation. Which pipeline setting should be adjusted to retain those records for processing?

A.Switch the pipeline's windowing from event-time to processing-time windows.
B.Set an allowed lateness on the windowing strategy that exceeds the 20-minute outage gap.
C.Increase the Pub/Sub subscription acknowledgement deadline so messages remain outstanding longer.
D.Change the stateful DoFn timer from event-time to processing-time so it fires only when data resumes.
AnswerB

Allowed lateness extends the window's lifetime beyond the watermark so records arriving after the watermark passes the window end are still processed instead of being dropped. A value greater than 20 minutes covers the outage gap. The stateful DoFn timers continue to fire, but late elements are routed into the still-open window and contribute to distance calculations.

Why this answer

Late data handling is governed by allowed lateness on the window, not by Pub/Sub delivery settings or the timer domain. When the watermark advances past a window end after a 20-minute gap, records with earlier timestamps are dropped unless the window is kept alive. Setting allowed lateness beyond the outage gap lets the stateful computation include the resumed telemetry in the correct event-time window.

Exam trap

The trap here is confusing message-level delivery guarantees in Pub/Sub with event-time window semantics in Beam, where late data is discarded based on the watermark rather than on acknowledgement timing.

276
MCQeasy

A company has deployed a classification model on Vertex AI. They want to detect data drift in real-time for the model's input features. Which service should they use?

A.Cloud Monitoring
B.Cloud Data Loss Prevention
C.Cloud Logging
D.Vertex AI Model Monitoring
AnswerD

Vertex AI Model Monitoring continuously evaluates incoming prediction requests against the training baseline, computing drift metrics on feature distributions. This satisfies the real-time detection requirement, unlike batch-only tools. It natively integrates with Vertex AI endpoints, so no separate pipeline is needed to surface skew or drift alerts.

Why this answer

Vertex AI Model Monitoring is the correct service because it is specifically designed to detect data drift and feature skew for models deployed on Vertex AI. It continuously monitors input features against a baseline distribution and alerts when drift exceeds a configured threshold, enabling real-time detection without requiring custom code.

Exam trap

The trap here is that candidates confuse general monitoring (Cloud Monitoring) with ML-specific drift detection, assuming any monitoring tool can detect data drift, when in fact Vertex AI Model Monitoring is the only service that performs statistical distribution comparison for model inputs.

How to eliminate wrong answers

Option A is wrong because Cloud Monitoring is a general-purpose observability service for metrics, uptime checks, and dashboards; it lacks built-in statistical drift detection for ML model features. Option B is wrong because Cloud Data Loss Prevention (DLP) is used for inspecting, classifying, and masking sensitive data, not for monitoring feature distributions or drift. Option C is wrong because Cloud Logging captures and stores log entries from services but does not perform statistical analysis or drift detection on model inputs.

277
MCQeasy

A startup is building a mobile app that needs to sync user data across devices in real time. They expect millions of concurrent users and need a NoSQL database with offline support and automatic multi-region replication. Which Google Cloud service meets these requirements?

A.Cloud Bigtable
B.Cloud Spanner
C.Firestore
D.Cloud SQL
AnswerC

Firestore provides a serverless NoSQL document store with built-in offline persistence via client SDKs and automatic multi-region replication, directly satisfying the real-time sync, offline support and global scale constraints. Its real-time listeners push updates across devices, meeting the millions-of-concurrent-users requirement.

Why this answer

Firestore is a NoSQL, serverless document database that provides real-time synchronization, offline support via local persistence, and automatic multi-region replication. It is designed for mobile and web apps with millions of concurrent users, making it the ideal choice for this use case.

Exam trap

The trap here is that candidates often confuse Cloud Spanner's global SQL capabilities with NoSQL requirements, or assume Cloud Bigtable's NoSQL label fits all NoSQL workloads, ignoring the specific need for real-time sync and offline support.

How to eliminate wrong answers

Option A is wrong because Cloud Bigtable is a wide-column NoSQL database optimized for high-throughput analytical workloads (e.g., time-series, IoT), not for real-time sync or offline mobile app support, and it lacks built-in multi-region replication. Option B is wrong because Cloud Spanner is a globally distributed, strongly consistent relational SQL database, not a NoSQL database, and while it supports multi-region replication, it does not provide offline support for mobile clients. Option D is wrong because Cloud SQL is a managed relational SQL database (MySQL, PostgreSQL, SQL Server) that is not NoSQL, does not support offline mobile sync, and requires manual configuration for multi-region replication.

278
MCQeasy

A data analyst needs to transform nested and repeated fields in BigQuery. They have a table with a column of type ARRAY<STRUCT<...>>. Which SQL function should they use to flatten the array into individual rows for analysis?

A.STRUCT
B.CAST
C.UNNEST
D.REPLACE
AnswerC

UNNEST expands an ARRAY into a set of rows, and combined with a CROSS JOIN or comma join it flattens ARRAY<STRUCT<...>> columns so each struct becomes an individual row. This directly satisfies the flattening requirement for analysis.

Why this answer

UNNEST is used to flatten arrays into rows. STRUCT is used to group fields. CAST is for type conversion.

REPLACE is for string replacement.

279
MCQmedium

You are designing a BigQuery data warehouse for a retail company. Queries frequently filter on order_date and customer_id. To optimize query performance and cost, which table design should you use?

A.Cluster by order_date and partition by customer_id
B.Partition by ingestion_time and cluster by order_date
C.Use a clustered table without partitioning
D.Partition by order_date and cluster by customer_id
AnswerD

Partitioning by order_date prunes scanned data when queries filter on that column, while clustering by customer_id co-locates rows sharing that value within each partition. This combination satisfies both frequent filter predicates, reducing bytes scanned and query cost compared with clustering or partitioning alone.

Why this answer

Partitioning by order_date and clustering by customer_id aligns the physical storage with the query access pattern: partition pruning eliminates irrelevant date ranges, and clustering sorts data within each partition by customer_id so filter and aggregation operations scan fewer blocks. This combination minimizes bytes scanned, which directly reduces both query latency and cost in BigQuery's on-demand pricing model.

Exam trap

PDE often tests the partition-then-cluster ordering — candidates reverse the two or choose ingestion_time partitioning, not realizing that the partition column must match the dominant filter predicate to enable pruning.

How to eliminate wrong answers

Option A is wrong because it reverses the roles — clustering by order_date and partitioning by customer_id would create a partition per customer, which is a poor cardinality choice and prevents effective date-range pruning. Option B is wrong because partitioning by ingestion_time does not match the order_date filter, so queries filtering on order_date cannot prune partitions and must scan all ingestion-time partitions. Option C is wrong because a clustered table without partitioning forgoes partition pruning entirely, so date-range filters scan the full table even though clustering helps with customer_id.

280
MCQeasy

A company wants to monitor the performance of a deployed model in production. Which metric indicates that the model's predictions are degrading?

A.Increase in prediction error rate
B.Increase in prediction latency
C.Decrease in throughput
D.Increase in number of requests
AnswerA

A rising prediction error rate, measured against ground-truth labels or a proxy, directly signals that model accuracy is deteriorating in production. Other signals such as latency or throughput reflect infrastructure health, not predictive quality, so they cannot indicate degradation of the model itself.

Why this answer

An increase in prediction error rate directly indicates that the model's outputs are deviating from the expected or ground-truth values, signaling degradation in predictive performance. This metric captures the core concept of model drift, where the statistical properties of the input data or the relationship between features and labels change over time, leading to less accurate predictions. In production ML monitoring, tracking error rate (e.g., classification accuracy, RMSE) is the primary method to detect when a model needs retraining or updating.

Exam trap

Google Cloud often tests the distinction between operational metrics (latency, throughput) and model performance metrics (error rate), trapping candidates who confuse system health with prediction quality.

How to eliminate wrong answers

Option B is wrong because prediction latency measures the time taken for the model to return a prediction, which reflects infrastructure or model complexity issues, not the accuracy or degradation of the predictions themselves. Option C is wrong because throughput (requests per second) is a measure of system capacity and scalability, not a direct indicator of prediction quality or model drift. Option D is wrong because an increase in the number of requests indicates higher demand or usage, which does not imply that the model's predictions are becoming less accurate or degrading.

281
Multi-Selectmedium

A team is designing a Spanner database for a global inventory system. They need to optimize query performance for frequently joined tables. Which THREE design decisions help achieve this? (Choose 3.)

Select 3 answers
A.Use Cloud SQL instead if joins are needed.
B.Design primary keys to distribute write load evenly across splits.
C.Use interleaved tables to co-locate related rows.
D.Store all data in a single table with JSON columns to avoid joins.
E.Create secondary indexes on columns used in WHERE clauses.
AnswersB, C, E

Evenly distributed primary keys prevent hotspots by spreading writes across splits, avoiding a single split becoming a bottleneck. This keeps load balanced across nodes, so frequently joined tables are not throttled by one overloaded split during concurrent inventory updates.

Why this answer

Option B is correct because Spanner scales by splitting data into ranges, and a monotonically increasing or skewed primary key concentrates writes on a single split; designing keys (for example, using hash prefixes or reversed timestamps) to distribute writes evenly avoids hotspots and keeps join-driving lookups performant. Option C is correct because interleaved tables physically co-locate child rows with their parent row in the same split, so joins between parent and child on the interleaved key prefix are executed locally without network shuffling, dramatically improving join performance. Option E is correct because secondary indexes let Spanner satisfy WHERE-clause predicates by index scan rather than full table scan, reducing the rows read before the join and thus lowering latency and cost.

Option A is not appropriate because Cloud SQL is a regional, non-horizontally-scalable relational service and does not provide Spanner's global distribution, external consistency, or interleaving; joins are fully supported in Spanner. Option D is not appropriate because collapsing everything into a single table with JSON columns discards Spanner's relational join and interleaving capabilities, prevents effective secondary indexing on nested fields, and typically worsens performance and schema maintainability.

Exam trap

Google often tests the misconception that avoiding joins entirely (Option D) is a better optimization than properly using Spanner's native features like interleaving and secondary indexes, which are designed to handle joins efficiently at scale.

282
MCQhard

You need to process a large volume of event data from Cloud Storage, apply complex transformations using Apache Spark, and then load the results into BigQuery. The data arrives in batches every hour. You want to minimize costs by using preemptible VMs. Which service should you use?

A.Cloud Composer
B.BigQuery
C.Dataproc
D.Dataflow
AnswerC

Dataproc runs Apache Spark natively and supports preemptible VMs for worker nodes, cutting compute costs substantially. It handles hourly batch ingestion from Cloud Storage, applies the complex Spark transformations, and writes results into BigQuery, matching the cost-minimisation constraint.

Why this answer

Dataproc is a managed Apache Spark and Hadoop service that supports preemptible VMs, making it the ideal choice for batch processing with complex transformations using Spark while minimizing costs. It integrates natively with Cloud Storage for input and BigQuery for output, and preemptible VMs can reduce compute costs by up to 80%. The hourly batch pattern aligns with Dataproc's job-based execution model.

Exam trap

PDE often tests the confusion between Dataflow and Dataproc, where candidates pick Dataflow for Spark workloads, but Dataflow uses Apache Beam, not Spark, and does not support preemptible VMs in the same cost-optimized way.

How to eliminate wrong answers

Option A is wrong because Cloud Composer is a workflow orchestration service (managed Airflow) used to schedule and monitor pipelines, not to run Spark transformations itself. Option B is wrong because BigQuery is a serverless data warehouse for analytics, not a Spark processing engine; it cannot run custom Spark code. Option D is wrong because Dataflow is based on Apache Beam, not Apache Spark, and while it supports batch processing, it does not natively run Spark jobs and does not use preemptible VMs in the same way (it uses worker VMs but the programming model is Beam, not Spark).

283
MCQmedium

A company uses dbt on BigQuery to transform data. They want to run dbt models on a schedule and manage environments (dev, prod). Which GCP service should they use to run dbt jobs?

A.Dataflow
B.Cloud Composer
C.Cloud Scheduler
D.Cloud Build
AnswerB

Cloud Composer provides managed Apache Airflow scheduling, letting dbt jobs run on cron-like schedules with separate dev and prod environments configured as distinct DAG variables or connections. This satisfies both the scheduling and environment-management requirements.

Why this answer

Cloud Composer is an Apache Airflow managed service that can schedule dbt runs.

284
MCQmedium

A company runs a real-time anomaly detection system on Google Cloud. Streaming data from IoT devices is ingested via Pub/Sub, processed by Dataflow (Apache Beam), and results are written to Bigtable for low-latency serving. Recently, the system has been experiencing increased latency and occasional data loss. The Dataflow pipeline shows high system lag and backlog in Pub/Sub. The Bigtable cluster has 3 nodes and is reporting high CPU utilization (over 90%). The team suspects the issue is with the pipeline configuration. They have already verified that there are no errors in the pipeline code and no network issues. Which action should they take to resolve the issue?

A.Increase the number of Bigtable nodes to handle the write throughput.
B.Change the Dataflow worker machine type to n2-standard-8.
C.Decrease the batch size in the Dataflow pipeline to reduce latency.
D.Increase the number of Dataflow workers to process messages faster.
AnswerA

Bigtable CPU above 90% with only three nodes indicates the cluster cannot absorb the pipeline's write throughput, which backs up Dataflow and Pub/Sub. Adding nodes distributes write load across more tablet servers, directly relieving the saturation causing the latency and data loss.

Why this answer

The high CPU utilization on Bigtable (over 90%) indicates that the cluster is saturated and cannot keep up with the write throughput from Dataflow. This causes backpressure in the pipeline, leading to increased system lag and backlog in Pub/Sub, and eventually data loss when Pub/Sub messages expire. Increasing the number of Bigtable nodes directly addresses the bottleneck by distributing the write load and reducing CPU pressure, which allows the pipeline to drain the backlog and reduce latency.

Exam trap

Google Cloud often tests the misconception that scaling Dataflow workers or changing machine types always resolves pipeline latency, but the trap here is that the bottleneck is at the sink (Bigtable), so you must scale the sink first to relieve backpressure.

How to eliminate wrong answers

Option B is wrong because changing the Dataflow worker machine type to n2-standard-8 would increase compute capacity for processing, but the bottleneck is at the Bigtable sink, not the Dataflow workers; the pipeline is already experiencing backpressure from Bigtable, so more worker CPU would not resolve the write throughput limitation. Option C is wrong because decreasing the batch size in Dataflow would increase the number of smaller writes to Bigtable, which actually increases overhead and CPU usage on Bigtable, worsening the latency and backlog issue. Option D is wrong because increasing the number of Dataflow workers would increase the parallelism of writes to Bigtable, further amplifying the write pressure on the already saturated Bigtable cluster, making the high CPU utilization and backlog worse.

285
MCQhard

A company is building a data lake on Cloud Storage with data from multiple sources. They need to apply schema-on-read and support ad-hoc SQL queries. Which architecture is most suitable?

A.Ingest to Cloud Spanner, query directly.
B.Ingest to Cloud SQL, then export to Cloud Storage for queries.
C.Ingest to Cloud Storage, create BigQuery external tables.
D.Ingest to Cloud Storage, load into Dataproc for queries.
AnswerC

BigQuery external tables query Cloud Storage data directly, preserving the raw files without loading or transformation. This satisfies schema-on-read, since the schema is applied at query time rather than ingest, and supports ad-hoc SQL through BigQuery's engine. The multi-source constraint is met because each external table maps to its own source path.

Why this answer

BigQuery external tables allow schema-on-read by defining the schema at query time over data stored in Cloud Storage, enabling ad-hoc SQL queries without loading data into a separate system. This architecture directly supports the requirement for schema-on-read and SQL-based analysis, as BigQuery provides a serverless, scalable SQL engine.

Exam trap

Google Cloud often tests the distinction between schema-on-read (BigQuery external tables) and schema-on-write (traditional databases like Cloud Spanner or Cloud SQL), where candidates mistakenly choose a transactional database for analytical workloads.

How to eliminate wrong answers

Option A is wrong because Cloud Spanner is a globally distributed, strongly consistent relational database designed for transactional workloads, not for schema-on-read or ad-hoc SQL queries over raw data in a data lake. Option B is wrong because Cloud SQL is a managed relational database for OLTP workloads, and exporting to Cloud Storage for queries adds unnecessary latency and complexity, failing to leverage schema-on-read directly. Option D is wrong because Dataproc is a managed Spark/Hadoop service that requires data loading and cluster management, which is not as efficient or serverless as BigQuery external tables for ad-hoc SQL queries on a data lake.

286
MCQeasy

An engineer needs to create a reusable Dataflow pipeline that can be executed with different parameters without modifying code. Which Dataflow feature should they use?

A.Dataflow Shuffle
B.Dataflow Flex Templates
C.Dataflow SQL
D.Dataflow Classic Templates
AnswerB

Flex Templates package the pipeline as a Docker image with a metadata specification, so the same template runs repeatedly with different runtime parameters and no code changes. This satisfies the stem's reusability and parameterisation constraint.

Why this answer

Dataflow Flex Templates allow packaging a pipeline as a Docker image with a metadata file, enabling reuse with different parameters at runtime without code changes. They support dynamic parameters and are the recommended approach for reusable pipelines.

Exam trap

PDE often tests the difference between Classic and Flex Templates, where candidates incorrectly choose Classic Templates for dynamic parameterization.

How to eliminate wrong answers

Option A is wrong because Dataflow Shuffle is a service for shuffling data, not for pipeline templating. Option C is wrong because Dataflow SQL is a way to run SQL queries on Dataflow, not for creating reusable parameterized pipelines. Option D is wrong because Classic Templates require parameters to be defined at compile time and are less flexible than Flex Templates.

287
MCQeasy

Your company uses Cloud Composer to run a daily ETL workflow. The workflow consists of several tasks that must run in a specific order. You want to receive an alert if any task fails. Which Cloud Monitoring feature should you use?

A.Create an alerting policy based on the Cloud Composer metric for task failures.
B.Set up a Cloud Function that checks the Airflow UI periodically.
C.Configure Airflow email alerts in the DAG's default_args.
D.Use Cloud Logging to create a log-based metric for task failures.
AnswerA

Cloud Composer exports metrics to Cloud Monitoring, including metrics for task failures. You can create an alerting policy that triggers when the number of failed tasks exceeds a threshold. This is the most direct way to get notified of task failures in your ETL workflow. It leverages the native integration between Cloud Composer and Cloud Monitoring, allowing you to set up notifications via email, SMS, or other channels.

Why this answer

The recommended way to alert on task failures in Cloud Composer is to use Cloud Monitoring alerting policies based on the built-in Cloud Composer metrics for task failures. These metrics are automatically exported and can be used to trigger alerts when failures occur. This approach is native, scalable, and integrates with various notification channels.

Other methods, such as log-based metrics or polling, are either redundant or less efficient.

Exam trap

The trap here is overlooking the built-in Cloud Composer metrics and instead opting for custom log-based metrics or polling, which are unnecessary and more complex.

288
Multi-Selecthard

You are designing a BigQuery table to store clickstream events. Queries will frequently filter by a user_id column and a TIMESTAMP column named event_time, and the table will grow to several petabytes. You want to minimize bytes scanned and cost for these queries. Which TWO actions should you take? (Choose two.)

Select 2 answers
A.Create a materialized view that pre-aggregates events by user_id and event_time, and query the view instead of the base table.
B.Partition the table by the DATE of event_time using time-unit partitioning, and require queries to include a filter on event_time.
C.Set the table's expiration time to 30 days so that older partitions are automatically deleted, reducing the data volume.
D.Cluster the table on the user_id column so that queries filtering on user_id scan only the relevant blocks within each partition.
E.Enable the BigQuery BI Engine reservation and route all queries through it to cache results in memory.
AnswersB, D

Partitioning by event_time allows BigQuery to prune partitions when queries filter on that column, dramatically reducing bytes scanned. Requiring a partition filter prevents accidental full-table scans. This is the primary cost-control mechanism for large time-series tables and directly addresses the frequent event_time filters.

Why this answer

For a petabyte-scale clickstream table filtered by event_time and user_id, partitioning on event_time enables partition pruning, and clustering on user_id enables block-level pruning within partitions. Together they minimize bytes scanned. Materialized views, table expiration, and BI Engine do not provide the same structural cost reduction for arbitrary filters.

Exam trap

The trap here is treating BI Engine or materialized views as general-purpose cost reducers, when they only accelerate specific cached or aggregated query patterns and do not replace partitioning and clustering.

289
Multi-Selectmedium

Your team is running a Dataflow streaming pipeline that reads from Pub/Sub, transforms data, and writes to BigQuery. You notice that the pipeline's backlog is growing and the processing latency has increased from seconds to minutes. You need to diagnose and resolve the issue. Which TWO actions should you take? (Choose two.)

Select 2 answers
A.Stop the pipeline, increase the number of workers in the streaming engine configuration, and restart it.
B.Increase the batch size in the WriteToBigQuery transform to reduce I/O operations.
C.Configure a dead-letter queue in Cloud Storage for failed messages to reduce reprocessing load.
D.Increase the maximum number of workers in the pipeline's autoscaling configuration to allow more compute resources.
E.Examine the Dataflow monitoring dashboard for metrics like system lag, data freshness, and worker throughput.
AnswersD, E

Raising the maximum worker cap lifts the ceiling that autoscaling cannot exceed, so Dataflow can add compute to drain the growing Pub/Sub backlog. Since latency rose with backlog, the pipeline is under-provisioned; the existing maximum was throttling horizontal scaling, and a higher limit restores throughput.

Why this answer

Option E is correct because the Dataflow monitoring dashboard exposes streaming-specific metrics such as system lag, data freshness, and worker throughput, which are exactly the signals needed to determine whether the growing backlog is caused by insufficient compute, a hot key, or a slow sink. Option D is correct because a streaming pipeline with a growing backlog and rising latency is typically under-provisioned, and raising the maximum number of workers in the autoscaling configuration lets Dataflow horizontally scale the worker pool to absorb the increased load. Option A is not appropriate because stopping and restarting the pipeline interrupts streaming state and is unnecessary when autoscaling can be adjusted on a running job.

Option B is not appropriate because increasing the WriteToBigQuery batch size does not address a Pub/Sub backlog and can actually increase latency and memory pressure. Option C is not appropriate because a dead-letter queue for failed messages does not relieve the reprocessing load of successfully processed but backlogged messages.

Exam trap

Google Cloud often tests the misconception that you must stop a streaming pipeline to change worker count or that increasing batch size always improves throughput, when in fact Dataflow supports live autoscaling and larger batches can worsen latency.

290
MCQhard

You are a data engineer at a global e-commerce company. Your team manages a real-time recommendation system that ingests user clickstream events from a Pub/Sub topic (topic-clickstream). The pipeline uses Dataflow to read events, join with user profile data from Cloud Bigtable, compute recommendations using a machine learning model hosted on Cloud Run, and write results to a BigQuery table for analytics. The pipeline has been running smoothly for months, but recently the Dataflow job started failing with the error: "Workflow failed. Causes: S01:ReadPubSub/Read+Transform/ParDo(ExtractUserID)+ ... (5a3b2c1d) The job failed because a worker encountered an out-of-memory error." The Dataflow job uses the Streaming Engine feature with a worker type of n2-standard-8 (8 vCPU, 32 GB memory) and autoscaling from 2 to 20 workers. The clickstream event rate has increased from 500 events/second to 5000 events/second over the past week. The user profile data in Bigtable has also grown, with average row size increasing from 1 KB to 10 KB due to additional fields. You need to resolve the out-of-memory errors without completely redesigning the pipeline. What should you do?

A.Increase the maximum number of workers in autoscaling from 20 to 50.
B.Change the worker machine type to n2-highmem-8 (8 vCPU, 64 GB memory) in the Dataflow job configuration.
C.Reduce the batch size in the Dataflow pipeline by setting the `max_batch_size` parameter to a lower value.
D.Increase the number of Bigtable nodes to improve read throughput.
AnswerB

Doubling worker memory to 64 GB directly addresses the out-of-memory error, since the larger 10 KB Bigtable rows inflate per-element state during the join. Streaming Engine already offloads shuffle, so the constraint is worker heap, not pipeline design.

Why this answer

The out-of-memory error is caused by the increased per-worker memory load from larger Bigtable rows (1 KB to 10 KB) and higher event throughput (500 to 5000 events/sec). Switching to n2-highmem-8 doubles the memory from 32 GB to 64 GB, giving each worker more headroom to cache user profiles and process larger batches without OOM. This directly addresses the root cause without redesigning the pipeline.

Exam trap

Google Cloud often tests the misconception that scaling out (more workers) solves memory issues, when in fact the per-worker memory limit is the bottleneck and must be increased via a higher-memory machine type.

How to eliminate wrong answers

Option A is wrong because increasing the maximum number of workers spreads the load across more machines but does not increase the memory per worker; each worker still has only 32 GB, so the same OOM condition persists on individual workers. Option C is wrong because reducing batch size lowers memory per batch but increases the number of batches and overhead, which can worsen performance and still not prevent OOM if the per-row memory footprint (10 KB) is the dominant factor. Option D is wrong because Bigtable node count affects read throughput and latency, not the memory consumed by the Dataflow worker when caching or processing rows; the OOM is on the Dataflow side, not Bigtable.

291
MCQmedium

A retail analytics team has a 400 TB BigQuery table partitioned by DATE on a transaction_date column. Analysts almost always filter by a single store_id and a date range, and each store has roughly 12 years of history. Queries currently scan the entire partition for the requested dates. The team wants to reduce bytes billed without changing the table name or rewriting the ingestion pipeline. What should they do?

A.Create a materialized view that pre-aggregates transactions by store_id and date, and point analysts at the view.
B.Enable partition expiration set to 30 days so older partitions are dropped automatically.
C.Add a clustering specification on store_id to the existing partitioned table.
D.Convert the table to use the US multi-region instead of a single region to increase scan parallelism.
AnswerC

BigQuery can co-locate and sort data within each date partition by the clustering columns, so a filter on store_id lets the engine read only the blocks containing that store rather than every block in the partition. This directly cuts bytes billed for the store-plus-date-range pattern and requires no change to the table name or pipeline.

Why this answer

Clustering a partitioned table on the frequently filtered column lets BigQuery prune blocks within each partition, so a query that filters on both date and store reads far less data than a full partition scan. The optimization applies transparently to the same table name, so ingestion and existing SQL keep working, and bytes billed drop in proportion to how well the clustering column discriminates rows inside each partition.

Exam trap

The trap here is assuming that partitioning alone prunes all irrelevant data, when partition pruning only eliminates whole partitions and a within-partition filter still scans the entire partition unless clustering is defined.

292
Multi-Selecthard

A company is migrating ML workflows to Vertex AI Pipelines. They want to ensure best practices for pipeline reproducibility and debugging. Which THREE actions should they take? (Choose three.)

Select 3 answers
A.Set a random seed for all training components
B.Store all artifacts in Cloud Storage with versioned prefixes
C.Pin all dependencies in training images
D.Use dynamic pipeline parameters for each run
E.Use conditional execution based on previous component outputs
AnswersA, B, C

Setting a random seed makes stochastic operations such as weight initialisation and data shuffling deterministic, so identical pipeline runs produce identical outputs. This directly satisfies the reproducibility requirement by eliminating run-to-run variation in training components.

Why this answer

Option A is correct because setting a random seed for all training components makes stochastic operations (weight initialization, data shuffling, dropout) deterministic, so identical inputs and parameters yield identical outputs, which is essential for reproducibility. Option B is correct because storing artifacts in Cloud Storage with versioned prefixes preserves immutable, uniquely addressable lineage for each pipeline run, allowing you to trace and compare outputs across executions for debugging. Option C is correct because pinning all dependencies (exact library versions in the training images) prevents silent behavioral changes from dependency drift, ensuring the same code produces the same results over time.

Option D is not correct because dynamic pipeline parameters vary inputs per run and, by themselves, do not improve reproducibility or debugging; they are a flexibility feature, not a best practice for deterministic reruns. Option E is not correct because conditional execution based on previous component outputs changes the pipeline's control flow at runtime, which can make runs harder to reproduce and debug rather than more deterministic.

Exam trap

Google Cloud often tests the distinction between features that improve workflow flexibility (like dynamic parameters or conditional execution) and those that enforce reproducibility and debuggability, leading candidates to confuse operational convenience with best practices for deterministic pipelines.

293
Multi-Selectmedium

A data engineer is designing a batch processing system using Cloud Dataproc. Which TWO practices improve performance and reduce costs? (Choose TWO.)

Select 2 answers
A.Always use persistent disks for all nodes.
B.Set autoscaling policies based on YARN memory.
C.Store intermediate data in HDFS.
D.Use preemptible VMs for worker nodes.
E.Use the largest machine types for master nodes.
AnswersB, D

Optimizes resource utilization.

Why this answer

Autoscaling policies based on YARN memory allow the cluster to dynamically add or remove worker nodes in response to actual resource demand from running jobs. This prevents over-provisioning (reducing costs) and ensures sufficient resources for job completion (improving performance), as Cloud Dataproc directly monitors YARN memory metrics to trigger scaling actions.

Exam trap

The trap here is that candidates often confuse HDFS with Cloud Storage, assuming intermediate data must be stored locally for performance, but Cloud Storage is actually faster and cheaper for transient data in Dataproc due to its native integration and lack of replication overhead.

294
MCQeasy

Refer to the exhibit. A Cloud Build step fails when pushing a Docker image to Artifact Registry. What is the missing IAM role for the Cloud Build service account?

A.roles/artifactregistry.writer
B.roles/containerregistry.admin
C.roles/storage.objectCreator
D.roles/cloudbuild.builds.editor
AnswerA

roles/artifactregistry.writer grants the push permission Cloud Build requires to upload image layers to the repository. Without it, the step fails with a permission denied error, so this role directly resolves the failing push described in the exhibit.

Why this answer

The Cloud Build service account needs the `roles/artifactregistry.writer` role to push Docker images to Artifact Registry. This role grants the necessary permissions to upload artifacts, including images, to the registry. Without it, the build step fails with an authorization error.

Exam trap

Google Cloud often tests the distinction between Artifact Registry and Container Registry roles, and the trap here is that candidates confuse `roles/containerregistry.admin` (for Container Registry) with the correct Artifact Registry role, or assume that Cloud Build's own editor role includes artifact push permissions.

How to eliminate wrong answers

Option B is wrong because `roles/containerregistry.admin` is for Container Registry (gcr.io), not Artifact Registry, and the question specifies Artifact Registry. Option C is wrong because `roles/storage.objectCreator` applies to Cloud Storage buckets, not Artifact Registry repositories. Option D is wrong because `roles/cloudbuild.builds.editor` allows managing Cloud Build builds but does not grant permissions to push artifacts to Artifact Registry.

295
Multi-Selecthard

A data engineer is designing a BigQuery table that will store billions of rows of application logs. Queries will almost always filter on a log_date column and then join on a user_id column, and the team wants to minimize both storage cost and bytes scanned. The table is append-only and grows by about 50 GB per day. Which two design choices should the engineer make? (Choose two.)

Select 2 answers
A.Cluster the table by user_id.
B.Set the table's expiration to 30 days to reduce storage cost.
C.Enable streaming inserts for the log ingestion pipeline.
D.Partition the table by log_date.
E.Create a materialized view that aggregates logs by user_id.
AnswersA, D

Clustering by user_id sorts data within each partition so that filters and joins on user_id read fewer blocks. Combined with partitioning by log_date, it addresses both the date filter and the user_id join in the stated workload. Clustering also improves the efficiency of the join by co-locating related rows within partitions.

Why this answer

Partitioning by log_date lets date-filtered queries prune partitions, and clustering by user_id reduces the blocks read for user_id filters and joins. Together they shrink bytes scanned and lower query cost for the dominant access pattern. Streaming inserts, table expiration, and materialized views do not address the general filtering and join pattern and can add cost or risk data loss.

Exam trap

The trap here is believing that clustering alone or a materialized view can substitute for partitioning when queries filter on a date column, when partition pruning is what removes whole segments of data from the scan.

296
MCQeasy

A data scientist trains a TensorFlow model using Vertex AI Training and wants to deploy it for online prediction. Which Vertex AI resource should the data scientist use to create an endpoint for serving predictions?

A.Vertex AI Batch Prediction Job
B.Vertex AI Endpoint
C.Vertex AI Feature Store
D.Vertex AI Model Registry
AnswerB

A Vertex AI Endpoint provides the managed serving resource that hosts the trained model and exposes an online prediction API. Deploying the model to an endpoint satisfies the online prediction requirement, whereas training jobs, datasets or model registry entries do not serve live requests.

Why this answer

Vertex AI Endpoint is the correct resource for deploying a trained model to serve online predictions. It provides a managed endpoint that exposes a REST API for real-time inference requests, which is exactly what the data scientist needs for online prediction.

Exam trap

Google often tests the distinction between batch prediction and online prediction, leading candidates to mistakenly choose Batch Prediction Job when the question explicitly asks for 'online prediction' or 'real-time serving'.

How to eliminate wrong answers

Option A is wrong because Vertex AI Batch Prediction Job is designed for asynchronous, batch processing of large datasets, not for real-time online prediction serving. Option C is wrong because Vertex AI Feature Store is a managed repository for storing, serving, and sharing feature data, not a service for deploying models or creating prediction endpoints. Option D is wrong because Vertex AI Model Registry is a central repository for managing model versions and metadata, but it does not directly serve predictions; models must be deployed to an endpoint for online serving.

297
MCQmedium

A data pipeline using Cloud Pub/Sub and Cloud Dataflow is experiencing duplicate messages. The source system publishes messages at least once. What Dataflow technique ensures exactly-once processing?

A.Use idempotent sinks
B.Use GlobalWindows
C.Set watermark threshold
D.Enable streaming engine
AnswerA

Idempotent sinks deduplicate writes at the destination, so repeated delivery of the same Pub/Sub message produces one persisted record. This satisfies the exactly-once requirement despite at-least-once publication, because Dataflow's streaming engine pairs sink idempotency with per-message deduplication rather than relying on the source.

Why this answer

Idempotent sinks ensure that even if Cloud Pub/Sub delivers the same message multiple times (due to its at-least-once delivery semantics), the Dataflow pipeline can deduplicate or safely reapply the same data without causing duplicates in the output. This is achieved by designing the sink (e.g., BigQuery with insertId, Cloud Storage with unique filenames) to recognize and ignore repeated writes, effectively providing exactly-once processing semantics downstream.

Exam trap

The trap here is that candidates confuse 'exactly-once processing' with 'exactly-once delivery' from the source, but Pub/Sub only guarantees at-least-once delivery, so the responsibility for deduplication falls on the Dataflow pipeline and its sink design, not on windowing or engine settings.

How to eliminate wrong answers

Option B is wrong because GlobalWindows groups all elements into a single window for batch-like processing, but it does not address message duplication; it only changes how data is windowed, not how duplicates are handled. Option C is wrong because setting a watermark threshold controls how long the pipeline waits for late data, which affects completeness and latency but does not prevent duplicate messages from being processed. Option D is wrong because enabling Streaming Engine improves scalability and reduces checkpoint latency in Dataflow, but it does not provide deduplication or exactly-once guarantees; duplicates can still occur from Pub/Sub's at-least-once delivery.

298
MCQeasy

A small startup wants to run a nightly batch job that transforms a 2 GB CSV file in Cloud Storage and writes the result back as Parquet. The team has no cluster administration experience, wants per-job pricing rather than an always-on cluster, and needs the transformation to run in under thirty minutes. Which Google Cloud approach should they choose?

A.Use Cloud Data Fusion to build a visual pipeline and run it on a dedicated ephemeral Dataproc cluster each night.
B.Create a persistent Dataproc cluster with a fixed number of worker nodes and submit the job to it every night.
C.Use a serverless Spark job on Dataproc Serverless, submitting the transformation and letting the platform provision and tear down resources per run.
D.Run the transformation on a Compute Engine VM with a cron job that invokes a local Spark installation each night.
AnswerC

Dataproc Serverless runs Spark workloads without provisioning or managing a cluster, and billing is based on the resources the workload consumes during execution. For a nightly 2 GB transformation, the job completes quickly and the team pays only for that run. This matches the requirements for no administration and per-job cost.

Why this answer

Dataproc Serverless executes Spark workloads without any cluster provisioning, and you pay for the resources consumed while the workload runs. A nightly 2 GB CSV-to-Parquet transformation finishes well within the thirty-minute window, so the team gets per-job pricing and avoids cluster administration entirely.

Exam trap

The trap here is equating Dataproc with clusters you must manage, when Dataproc Serverless runs Spark workloads on demand with no cluster to provision or tear down.

299
MCQhard

A company stores highly sensitive financial data in BigQuery. They need to encrypt certain columns (e.g., credit card numbers) with customer-managed encryption keys (CMEK) at the column level. Which BigQuery feature should they use?

A.Customer-managed encryption keys (CMEK) on the dataset
B.VPC Service Controls
C.AEAD encryption functions with Cloud KMS
D.BigQuery Data Catalog with policy tags
AnswerD

Policy tags classify columns and enforce access control through fine-grained permissions; they do not apply customer-managed encryption keys to individual columns. It is tempting because policy tags deliver column-level security, but that is access governance, not column-level CMEK encryption as the stem requires.

Why this answer

BigQuery supports column-level encryption with customer-managed encryption keys (CMEK) through column-level encryption using policy tags in BigQuery Data Catalog. You create a policy tag, associate a Cloud KMS key with that policy tag, and apply the policy tag to a column. BigQuery then encrypts that column with the customer-managed key.

AEAD functions with Cloud KMS provide application-layer encryption of individual values, but they are not the native BigQuery CMEK feature and should not be described as CMEK.

Exam trap

Candidates often confuse dataset-level CMEK (encrypts all data at rest in the dataset) with column-level CMEK via policy tags (encrypts specific columns with customer-managed keys), leading them to mistakenly choose dataset CMEK when the requirement is for column-level granularity.

How to eliminate wrong answers

Option A is wrong because CMEK on a dataset encrypts all data at rest in that dataset, not at the column level; it cannot selectively encrypt specific columns like credit card numbers. Option B is wrong because VPC Service Controls provide network security boundaries to prevent data exfiltration, not column-level encryption of data within BigQuery. Option D is wrong because BigQuery Data Catalog with policy tags is used for fine-grained access control and data classification (e.g., masking or row-level security), not for encrypting column data with customer-managed keys.

300
MCQhard

You are building a real-time fraud detection system using Dataflow. Events from Pub/Sub need to be grouped by user_id within a 5-minute window to detect suspicious patterns. Some events may be delayed by up to 2 minutes. How should you configure the window and trigger to balance accuracy and latency?

A.Sliding window of 5 minutes with a 1-minute period and no allowed lateness
B.Session window with a gap duration of 5 minutes
C.Fixed window of 5 minutes with no allowed lateness and default trigger
D.Fixed window of 5 minutes with allowed lateness of 2 minutes and early trigger every 1 minute
AnswerD

Fixed windows align to epoch boundaries, so a 5-minute window groups events by user_id without overlap. Allowed lateness of 2 minutes retains late-arriving events, satisfying the stated delay constraint. Early triggers every minute emit speculative results, balancing latency against the accuracy gained from late data.

Why this answer

Option D is correct because it directly addresses both the accuracy requirement (handling up to 2 minutes of late data) and the latency requirement (emitting early results every 1 minute). A fixed 5-minute window groups events into non-overlapping intervals, which is appropriate for per-user fraud pattern detection. Setting allowed lateness to 2 minutes ensures that events delayed by up to 2 minutes are still included in the correct window, and an early trigger (e.g., repeating every 1 minute) provides low-latency partial results before the window closes.

This combination balances completeness and timeliness.

Exam trap

PDE often tests the misconception that a sliding window is needed for overlapping patterns, but the question specifies grouping by user_id within a 5-minute window, which implies non-overlapping fixed windows; also, candidates may overlook the need for allowed lateness to handle delayed events, opting for no lateness and thus sacrificing accuracy.

How to eliminate wrong answers

Option A is wrong because a sliding window with a 1-minute period creates overlapping windows, which would cause the same event to be counted in multiple windows, leading to duplicate fraud alerts and incorrect aggregation; also, no allowed lateness means delayed events are dropped, reducing accuracy. Option B is wrong because session windows group events based on gaps of inactivity, not fixed time intervals; a 5-minute gap duration would create windows that vary in length and do not align with the requirement to group events within a 5-minute window, making it unsuitable for detecting patterns in fixed time frames. Option C is wrong because a fixed window with no allowed lateness and the default trigger will drop any event arriving after the window closes, so the 2-minute delayed events would be lost, compromising accuracy; the default trigger also emits results only at the end of the window, increasing latency.

Page 3

Page 4 of 10

Page 5

All pages