Courseiva

Google Professional Data Engineer (PDE) — Questions 1–75

747 questions total · 10pages · All types, answers revealed

Page 1 of 10

Page 2
1
MCQmedium

Your team uses a CI/CD pipeline with Cloud Build to train and deploy ML models on Vertex AI. You want to ensure that only models that pass validation checks (e.g., accuracy threshold, fairness metrics) are promoted to production. What is the best way to implement this?

A.Use Cloud Scheduler to trigger retraining and only deploy if the new model outperforms the previous one on a holdout set.
B.Use Vertex AI Model Registry's automatic promotion feature that moves models to production based on evaluation results.
C.Configure Cloud Functions to re-evaluate the model daily and promote if it passes.
D.In the Cloud Build pipeline, after training, run validation scripts. If validation passes, deploy to a staging endpoint for manual approval, then promote to production.
AnswerD

Validation scripts inside Cloud Build gate promotion on accuracy and fairness metrics, satisfying the stem's requirement that only models passing checks reach production. Staging with manual approval adds human review before promotion, though it introduces delay. This keeps enforcement within the existing pipeline rather than relying on post-deployment monitoring.

Why this answer

It integrates validation directly into the CI/CD pipeline using Cloud Build, ensuring that only models passing specific checks (e.g., accuracy threshold, fairness metrics) are promoted. By running validation scripts after training and requiring manual approval before production promotion, this approach provides both automated gatekeeping and human oversight, aligning with MLOps best practices for safe model deployment.

Exam trap

Google Cloud often tests the misconception that Vertex AI Model Registry has built-in automatic promotion based on evaluation metrics, but in reality, it requires external orchestration (like Cloud Build) to implement such logic.

How to eliminate wrong answers

Option A is wrong because Cloud Scheduler triggers retraining on a schedule, not based on validation results, and it does not integrate with the CI/CD pipeline to enforce promotion gates. Option B is wrong because Vertex AI Model Registry does not have an automatic promotion feature based on evaluation results; it stores and manages models but requires external logic to decide promotion. Option C is wrong because Cloud Functions re-evaluating the model daily is reactive and does not tie into the build pipeline's validation step, potentially promoting a model that was not validated at the time of training.

2
MCQhard

A logistics company wants to optimize their delivery routes using historical GPS data. The data is stored in BigQuery and is updated daily. They need to run a complex machine learning model that requires iterative processing over the entire dataset using Apache Spark. The model training takes several hours and must be run weekly. They want to minimize cost and operational overhead. Which approach should they take?

A.Run the Spark job on a long-running Dataproc cluster that is always available.
B.Use Dataproc Serverless for Spark with the BigQuery connector to read data directly from BigQuery.
C.Use BigQuery ML to train the model directly in BigQuery without Spark.
D.Export the data to Cloud Storage and run a Dataproc cluster with autoscaling, then delete the cluster after training.
AnswerB

Dataproc Serverless for Spark can read data directly from BigQuery using the BigQuery connector, eliminating the need to export data. It is a serverless solution, so it automatically provisions and scales resources, and charges only for the duration of the job. This minimizes both cost and operational overhead, as there is no cluster to manage. It supports complex Spark MLlib models and iterative processing, making it ideal for this scenario.

Why this answer

Dataproc Serverless for Spark with the BigQuery connector allows the company to run Spark jobs directly on BigQuery data without exporting it. It is serverless, so it minimizes operational overhead and costs by charging only for job execution. It supports complex Spark MLlib models and iterative processing, making it the best fit for weekly training on historical GPS data.

This approach aligns with the goals of minimizing cost and operational overhead while leveraging Spark's capabilities.

Exam trap

The trap here is assuming that data must be exported from BigQuery to Cloud Storage for Spark processing, overlooking the direct BigQuery connector available in Dataproc Serverless.

3
Multi-Selecthard

You are setting up Dataplex data quality rules for a BigQuery table. You want to define rules that check for non-null values in key columns and also validate that a column's values fall within a certain range. Which TWO rule types must you use? (Choose 2)

Select 2 answers
A.Table rule (e.g., row count)
B.Row rule (e.g., not null)
C.Partition rule
D.Column rule (e.g., value range)
E.Custom SQL rule
AnswersB, D

A row rule evaluates each record independently, so a not-null condition on key columns is expressed as a row-level rule. Range validation on a column's values is also row-scoped, checking every value against the defined bounds. Both satisfy the non-null and range checks required.

Why this answer

Dataplex data quality rules include row rules (for null checks) and column rules (for range or value checks). Table rules apply to the entire table; custom SQL can be used but row and column rules are the standard.

4
MCQhard

A company processes IoT sensor data in near real-time. They ingest data via Cloud Pub/Sub, then a Dataflow streaming pipeline writes to Bigtable for low-latency queries. Recently, they observed increased Pub/Sub message backlog during traffic spikes. What is the most effective scaling strategy?

A.Increase Pub/Sub subscription throughput by increasing the number of partitions
B.Increase Dataflow worker count and adjust autoscaling configuration
C.Use a Cloud Scheduler to throttle Pub/Sub publishing
D.Add a Cloud Function to pre-process messages before they are consumed by Dataflow
AnswerB

Increasing Dataflow worker count with tuned autoscaling directly drains the Pub/Sub backlog, since backlog itself is the autoscaling signal for streaming pipelines. Horizontal worker scaling raises parallel processing throughput, matching the traffic spikes, whereas Bigtable or Pub/Sub-side changes would not accelerate consumption of the queued messages.

Why this answer

The increased Pub/Sub backlog during traffic spikes indicates that the Dataflow pipeline is unable to consume messages as fast as they are being published. Increasing the Dataflow worker count and adjusting autoscaling configuration allows the pipeline to scale horizontally, processing more messages per second and reducing the backlog. Pub/Sub itself is designed to handle high throughput, so the bottleneck is the consumer (Dataflow), not the ingestion layer.

Exam trap

The trap here is that candidates mistakenly think Pub/Sub's throughput is limited by partitions (like Kafka) or that throttling the publisher is a valid scaling strategy, when in fact the bottleneck is the streaming pipeline's processing capacity, which must be scaled horizontally.

How to eliminate wrong answers

Option A is wrong because Pub/Sub does not use partitions like Kafka; increasing partitions is not a valid concept for Pub/Sub subscriptions, and throughput is managed by the subscriber's ability to pull messages, not by partitioning. Option C is wrong because throttling Pub/Sub publishing with Cloud Scheduler would reduce the incoming data rate, but this is counterproductive for near real-time processing and does not address the root cause of insufficient consumer capacity. Option D is wrong because adding a Cloud Function to pre-process messages would introduce an additional processing step that could further increase latency and does not directly solve the Dataflow pipeline's inability to keep up with the message volume.

5
Multi-Selectmedium

A Dataflow streaming job is processing data from Pub/Sub and writing to BigQuery. The job is stuck with the message 'No progress has been made' for several minutes. Which TWO actions should the team take to troubleshoot and resolve the issue? (Choose TWO.)

Select 2 answers
A.Set the updateCompatibility flag to true and restart the pipeline.
B.Increase the persistent disk size for all workers to reduce I/O contention.
C.Examine the worker logs in Cloud Logging for any error messages or exceptions.
D.Force stop the pipeline and update it with a new version using the --update flag.
E.Enable Dataflow Streaming Engine to move state to the backend and reduce worker load.
AnswersC, E

Worker logs in Cloud Logging capture exceptions, permission errors, and BigQuery write failures that stall a streaming pipeline, directly explaining the 'No progress' state. Examining them satisfies the troubleshooting requirement by surfacing the underlying error before any configuration change.

Why this answer

Option C is correct because when a Dataflow streaming job reports 'No progress has been made,' the first troubleshooting step is to inspect the worker logs in Cloud Logging, which surface exceptions, stuck DoFns, BigQuery write errors, or Pub/Sub backlog issues that explain the stall. Option E is correct because enabling Dataflow Streaming Engine moves pipeline state (such as timers and stateful processing) off the worker VMs and into the Dataflow backend, reducing worker CPU/memory pressure and often resolving stalls caused by state-heavy or resource-constrained workers. Option A is incorrect because updateCompatibility is not a valid Dataflow pipeline option for resolving a stuck job.

Option B is incorrect because increasing persistent disk size addresses disk I/O capacity, not the typical causes of a streaming job making no progress. Option D is incorrect because force-stopping and redeploying with --update does not diagnose the root cause and can lose in-flight state without first examining logs.

Exam trap

Google Cloud often tests the misconception that increasing resources (like disk size) or restarting the pipeline is the default fix, when in reality the first step is always to inspect logs to understand the failure mode.

6
MCQmedium

A media company ingests thousands of small JSON files per hour into a Cloud Storage bucket and wants to analyze them with BigQuery. Analysts frequently filter by event date and by device type, and they want to minimize query cost. The team wants a managed approach that avoids writing custom transformation code. Which BigQuery feature should the engineer use?

A.Create an external table over the bucket and rely on automatic schema detection.
B.Use a BigQuery load job to create a native table partitioned by event date and clustered by device type.
C.Create a federated query that reads the bucket through the Cloud Storage connector.
D.Define a BigLake external table with a metadata cache and partition pruning enabled.
AnswerB

Loading the JSON into a native table consolidates the small files into BigQuery managed storage, and partitioning by event date lets queries prune to the relevant date range. Clustering by device type further reduces bytes scanned for filters on that column. A standard load job requires no custom transformation code, matching the managed requirement, and native storage gives the best query performance and cost profile.

Why this answer

Loading the JSON into a native, partitioned, and clustered table moves the data into BigQuery managed columnar storage, where filters on event date prune partitions and filters on device type benefit from clustering. This reduces bytes scanned and therefore cost, and it requires no custom transformation code. External and federated approaches keep querying the raw small files, which is less efficient and more expensive for repeated analysis.

Exam trap

The trap here is assuming that any external or federated access to Cloud Storage is automatically cheaper, when repeated analytical queries over raw small files usually cost more than loading into partitioned native storage.

7
MCQhard

A company uses Pub/Sub with push subscriptions to deliver events to a Cloud Run service. Recently, the service has been returning HTTP 429 (Too Many Requests), causing messages to be retried and eventually sent to the dead letter topic. What is the MOST likely cause?

A.The subscription ackDeadline is set too low, causing messages to be redelivered
B.The push endpoint is not acknowledging messages quickly enough, causing a backlog
C.The dead letter topic is misconfigured, causing messages to be sent to it prematurely
D.The Cloud Run service needs more instances to handle the incoming request rate
AnswerD

Cloud Run scales based on requests; if max instances reached, it returns 429. Increasing instances or adjusting concurrency resolves this.

Why this answer

Push subscriptions can be rate limited by the receiving service. Increasing ackDeadline gives more time but doesn't reduce rate. Using pull subscriptions shifts the rate control to the subscriber.

Adjusting max delivery attempts only affects how many retries before dead letter, not rate limiting.

8
MCQhard

A healthcare organization stores patient data in BigQuery. They need to encrypt a specific column (e.g., SSN) using a key they manage, and decrypt it only for authorized queries via a user-defined function. Which approach should they use?

A.Use BigQuery AEAD encryption functions with a Cloud KMS key
B.Use BigQuery column-level access controls
C.Use Cloud Key Management Service (Cloud KMS) with CMEK for the BigQuery dataset
D.Use Cloud Data Loss Prevention (DLP) to de-identify the column
AnswerA

BigQuery AEAD functions encrypt and decrypt individual column values using a Cloud KMS key the organisation controls, and can be invoked inside a user-defined function. This satisfies both the customer-managed key and authorised-query decryption constraints.

Why this answer

BigQuery AEAD encryption functions (e.g., `AEAD.ENCRYPT` and `AEAD.DECRYPT`) allow you to encrypt a specific column using a customer-managed key stored in Cloud KMS, and then decrypt it only within a user-defined function (UDF) that enforces access controls. This meets the requirement of per-column encryption with key management and authorized decryption via a UDF.

Exam trap

A common mistake is to choose dataset-level encryption (CMEK) because it involves Cloud KMS, but CMEK does not allow per-column encryption or UDF-controlled decryption. The correct approach uses BigQuery AEAD encryption functions with a Cloud KMS key for column-level encryption and authorized decryption via a UDF.

How to eliminate wrong answers

Option B is wrong because BigQuery column-level access controls only restrict who can see the column, but they do not encrypt the data at rest or in transit, so the data remains in plaintext and does not satisfy the encryption requirement. Option C is wrong because Cloud KMS with CMEK encrypts the entire BigQuery dataset at the storage level, not a specific column, and decryption is automatic for authorized users, not controlled via a UDF. Option D is wrong because Cloud DLP de-identifies data (e.g., masking or tokenization) but is not designed for reversible encryption with a customer-managed key and UDF-based decryption; it is typically used for static de-identification, not dynamic per-query decryption.

9
Multi-Selecteasy

Which TWO actions can reduce the cost of running a Dataproc cluster for a nightly batch job?

Select 2 answers
A.Increase the number of worker nodes for faster processing.
B.Use high-memory machine types for master node.
C.Use preemptible VMs for worker nodes.
D.Attach local SSDs to all nodes.
E.Delete the cluster after the job completes.
AnswersC, E

Preemptible VMs cost substantially less than standard worker instances, directly cutting the dominant compute expense of the nightly batch job. Because Dataproc tolerates worker preemption by re-running affected tasks, the job's fault-tolerant batch nature satisfies the availability constraint that would otherwise rule this option out.

Why this answer

Option C is correct because preemptible VMs in Dataproc cost significantly less than standard VMs (up to 80% cheaper), and for a nightly batch job that can tolerate interruptions, using preemptible worker nodes reduces compute costs substantially. Option E is correct because Dataproc charges for cluster resources while the cluster exists; deleting the cluster after the nightly job completes ensures you only pay for the duration of the job rather than leaving idle nodes running. Option A is incorrect because increasing worker nodes raises the number of billable VM instances, increasing cost rather than reducing it.

Option B is incorrect because high-memory machine types for the master node increase the per-hour price of that node without benefiting a typical batch workload. Option D is incorrect because attaching local SSDs adds cost (and local SSD pricing) to every node, increasing rather than reducing the cluster's expense.

Exam trap

Google Cloud often tests the misconception that scaling up resources (more nodes or faster hardware) always reduces cost by shortening runtime, but in reality, the increased per-hour cost usually outweighs the time savings for batch jobs.

10
MCQmedium

Your team runs a Cloud Composer (Airflow) environment that executes a nightly BigQuery ELT DAG. A downstream task must only run after an upstream task that loads a partitioned table completes, and you want the downstream task to wait for the load's completion signal without polling BigQuery repeatedly. Which Airflow mechanism should you implement to coordinate these tasks within the DAG?

A.Define a direct downstream dependency (loader >> downstream) so Airflow's scheduler marks the downstream task runnable only after the loader task succeeds.
B.Use a task dependency with an ExternalTaskSensor pointed at the loader task's DAG and task ID, configured with the correct execution_delta.
C.Set the loader task's trigger_rule to all_done and add a dependency edge so the downstream task runs immediately after the loader reaches any terminal state.
D.Use an XCom to push a completion flag from the loader task and have the downstream task pull it via xcom_pull, combined with a trigger rule of all_success on the dependency edge.
AnswerA

In Airflow, a dependency edge makes the scheduler hold the downstream task until the upstream task reaches a successful terminal state. This is the native, no-polling coordination mechanism inside a single DAG, exactly matching the requirement that the downstream task wait for the load's completion signal without repeatedly querying BigQuery.

Why this answer

A direct dependency edge is the canonical Airflow way to sequence tasks inside one DAG: the scheduler keeps the downstream task in a waiting state until the upstream load succeeds. Sensors and XComs serve other purposes and would add complexity or incorrect semantics here. The dependency edge provides the required completion gating with no polling of BigQuery.

Exam trap

The trap here is assuming a sensor or XCom is needed when a simple dependency edge already enforces ordering within the same DAG.

11
MCQmedium

A company has a production model deployed on Vertex AI that shows declining accuracy over time. The model uses features from a BigQuery feature store. The data science team suspects data drift. What is the most efficient way to monitor and detect drift for this model?

A.Enable Vertex AI Model Monitoring on the endpoint to automatically detect skew and drift
B.Periodically export training data and production data to CSV and compare distributions manually
C.Create a scheduled retraining pipeline that runs weekly
D.Set up Cloud Monitoring dashboards to track prediction request volumes and error rates
AnswerA

Vertex AI Model Monitoring attaches to the deployed endpoint and continuously compares incoming prediction feature distributions against the training baseline, automatically flagging training-serving skew and drift. This satisfies the drift-detection requirement without building custom BigQuery comparison pipelines.

Why this answer

Vertex AI Model Monitoring is purpose-built for detecting feature skew and drift in production models. It automatically compares the distribution of prediction request data against training data statistics, alerting when significant divergence occurs — this is the most efficient and integrated approach for a Vertex AI endpoint.

Exam trap

Google Cloud often tests the distinction between monitoring for data drift (which requires distribution comparison) and monitoring for operational metrics (like latency or error rates), leading candidates to confuse Cloud Monitoring dashboards with drift detection.

How to eliminate wrong answers

Option B is wrong because manually exporting data to CSV and comparing distributions is inefficient, error-prone, and does not scale; Vertex AI Model Monitoring automates this process with statistical tests like Jensen-Shannon divergence. Option C is wrong because scheduled retraining does not detect drift — it blindly retrains on a fixed schedule, which may be unnecessary or miss drift between cycles, and it does not provide monitoring or alerts. Option D is wrong because Cloud Monitoring dashboards for request volumes and error rates track operational health, not data drift; they cannot detect changes in feature distributions that cause accuracy decline.

12
MCQmedium

You are building a binary classification model using AutoML Tables on Vertex AI. The dataset has a severe class imbalance (1% positive class). Which strategy should you use to handle the imbalance?

A.Oversample the minority class using SMOTE before training.
B.Do nothing; AutoML Tables automatically handles class imbalance.
C.Downsample the majority class to match the minority class size.
D.Use a weight column to assign higher weights to the minority class.
AnswerD

AutoML Tables allows setting class weights to address imbalance; it will adjust the loss function accordingly.

Why this answer

The correct option is D: use a weight column in AutoML Tables to assign higher weights to the minority class. In Vertex AI AutoML Tables, you can add a weight column to the training data and specify it as the weight column, so minority-class rows contribute more to the loss. This directly counteracts the 1% positive-class imbalance without changing the data distribution.

Options A and C are not supported preprocessing steps in AutoML Tables, since it manages data splitting and training internally and does not expose SMOTE or manual downsampling. Option B is incorrect because AutoML Tables does not automatically correct severe class imbalance; you must supply weights or adjust the optimization objective. Option D is therefore the appropriate, supported mechanism for this scenario.

Exam trap

Confusing scikit-learn/XGBoost-style class_weight parameters with Vertex AI AutoML Tables, which instead uses a dataset weight column to handle class imbalance.

13
Multi-Selecthard

A data engineer needs to create a unified table that combines data from Cloud Storage (Parquet files) and BigQuery native tables, with fine-grained access control and governance. Which three Google Cloud features should they use together? (Choose THREE.)

Select 3 answers
A.BigQuery
B.Dataproc
C.BigLake
D.Cloud Storage
E.Cloud SQL
AnswersA, C, D

BigQuery is used for both native tables and the unified query engine.

Why this answer

BigQuery is correct because it serves as the unified query engine that can read data from both Cloud Storage (via external tables or BigLake) and native BigQuery tables, enabling a single SQL interface for analysis. It also integrates with fine-grained access control through row-level security and column-level access policies, and supports governance via Data Catalog and VPC Service Controls.

Exam trap

The trap here is that candidates may confuse Dataproc (a processing engine) with a storage or query service, or think Cloud SQL can handle Parquet files, when the correct combination requires BigQuery, BigLake, and Cloud Storage to achieve unified querying and governance.

14
MCQeasy

A company is designing a streaming data pipeline to process real-time clickstream events. They need to aggregate events by session window with a 5-minute gap and enable exactly-once processing semantics. Which Google Cloud service should they use?

A.Cloud Pub/Sub with Cloud Functions
B.Cloud Dataflow with Apache Beam
C.Cloud Dataproc with Spark Streaming
D.Cloud Bigtable with Dataflow templates
AnswerB

Cloud Dataflow with Apache Beam satisfies both constraints: Beam's session windows natively aggregate events separated by a 5-minute gap, and Dataflow's exactly-once processing guarantees deduplicate records end-to-end. Pub/Sub alone lacks windowing, while Dataproc requires manual checkpointing, so neither meets the stated semantics.

Why this answer

Cloud Dataflow with Apache Beam is the correct choice because it provides native support for session windows with a 5-minute gap duration and exactly-once processing semantics via its sink and source integrations. Dataflow's Beam SDK allows you to define session windows using `Window.into(Sessions.withGapDuration(Duration.standardMinutes(5)))`, and its checkpointing and idempotent writes ensure exactly-once delivery even in failure scenarios.

Exam trap

Google Cloud often tests the distinction between stateless serverless services (like Cloud Functions) and stateful stream processing engines (like Dataflow), leading candidates to incorrectly choose Cloud Pub/Sub with Cloud Functions because they overlook the need for session window state management and exactly-once semantics.

How to eliminate wrong answers

Option A is wrong because Cloud Pub/Sub with Cloud Functions does not support session windowing natively; Cloud Functions are stateless and cannot maintain session state across invocations, and Pub/Sub offers at-least-once delivery, not exactly-once. Option C is wrong because Cloud Dataproc with Spark Streaming can implement session windows but requires manual state management and does not provide built-in exactly-once semantics; Spark Streaming's checkpointing can lead to duplicate outputs in failure recovery. Option D is wrong because Cloud Bigtable with Dataflow templates is a storage and template combination, not a processing service; Dataflow templates can be used for streaming but the question asks for the service to use, and Bigtable is a NoSQL database, not a stream processing engine.

15
MCQmedium

You are using Vertex AI Feature Store to serve features for online predictions. Your model requires features from multiple sources with low latency (<10ms). Which type of serving should you use?

A.Online serving with Cloud SQL
B.Offline serving with BigQuery
C.Online serving with Bigtable
D.Offline serving with Cloud Storage
AnswerC

Online serving with Bigtable satisfies the sub-10ms latency constraint because Bigtable provides low-latency row-key lookups for feature values at prediction time. Vertex AI Feature Store uses Bigtable as its online serving store, retrieving the latest feature values for each entity directly, rather than scanning historical data as batch serving would.

Why this answer

Vertex AI Feature Store online serving is designed for low-latency lookups at prediction time, and Bigtable is the underlying low-latency store used for online serving. Bigtable provides single-digit millisecond reads at scale, which satisfies the <10ms requirement. Offline serving (BigQuery) is optimized for batch/high-throughput reads, not sub-10ms online lookups.

Exam trap

The trap is conflating 'online serving' with any low-latency database; candidates may pick Cloud SQL, but Vertex AI Feature Store's online serving is backed by Bigtable, not Cloud SQL.

How to eliminate wrong answers

Option A is wrong because Cloud SQL is a relational database not used as the online serving backend for Vertex AI Feature Store; it cannot meet the <10ms low-latency requirement at scale and is not the supported online store. Option B is wrong because offline serving with BigQuery is for batch training and bulk feature retrieval, with latencies in seconds, not <10ms. Option D is wrong because offline serving with Cloud Storage is for batch export/import of feature data, not real-time online prediction serving.

16
Multi-Selecthard

Which TWO metrics are most important to monitor for a real-time online prediction system to ensure service reliability and model performance?

Select 2 answers
A.Feature distribution skew between training and serving
B.Prediction latency (p50, p99)
C.Number of training examples used for the latest model version
D.Batch prediction job throughput
E.Prediction error rate (e.g., 4xx/5xx responses)
AnswersB, E

Prediction latency percentiles directly expose whether the real-time serving path meets its responsiveness constraint, since p99 captures tail delays that averages hide. Monitoring p50 and p99 together reveals queueing, cold starts, or overload degrading user-facing predictions, satisfying the stem's service reliability requirement for an online system.

Why this answer

Prediction latency (p50, p99) is correct because a real-time online prediction system must respond within strict time budgets, and monitoring percentile latencies (especially p99) reveals tail-latency regressions that directly impact service reliability and user experience. Prediction error rate (e.g., 4xx/5xx responses) is correct because it is the primary signal of serving failures, such as malformed requests, timeouts, or model-server crashes, and is essential for detecting availability and correctness problems in production. Feature distribution skew between training and serving, while important for model quality, is a data-quality/drift metric rather than a core real-time reliability metric, so it is not one of the two most important here.

The number of training examples used for the latest model version is a training-time artifact and does not reflect live serving health. Batch prediction job throughput applies to offline/batch inference, not to a real-time online prediction system.

Exam trap

Google Cloud often tests the distinction between offline training metrics (like feature skew or training example count) and real-time serving metrics (like latency and error rate), trapping candidates who confuse model performance monitoring with service reliability monitoring.

17
MCQmedium

A company needs to predict whether a product image contains a specific defect. They have 10,000 labeled images and want to build a model quickly without writing custom code or training from scratch. Which GCP service should they use?

A.AutoML Tables
B.AutoML Vision
C.Vertex AI custom training
D.AutoML Natural Language
AnswerB

AutoML Vision trains image classification models from labelled datasets using Google's managed infrastructure, requiring no custom code or algorithm design. It accepts the 10,000 labelled images and handles training automatically, satisfying the constraint of building a defect-detection model quickly without training from scratch.

Why this answer

AutoML Vision is designed for custom image classification tasks with minimal ML expertise. It uses transfer learning and supports up to millions of images. AutoML Tables handles tabular data, not images.

Vertex AI custom training would require more effort. AutoML NLP is for text data.

18
MCQmedium

A team trained a model on a Vertex AI custom training job and wants to deploy it to an endpoint for online predictions. They have the model artifacts stored in Cloud Storage. What steps are required?

A.Upload model to Model Registry, create endpoint, deploy model
B.Directly deploy from Cloud Storage without Model Registry
C.Create endpoint, then upload model
D.Use Vertex AI Batch Prediction only
AnswerA

Uploading the Cloud Storage artifacts to Vertex AI Model Registry creates the model resource, then creating an endpoint and deploying that model to it exposes it for online prediction. These three steps are the minimum sequence required.

Why this answer

To deploy a model for online predictions on Vertex AI, you must first upload the model artifacts from Cloud Storage to the Model Registry, which creates a versioned model resource. Then you create an endpoint (or use an existing one) and deploy the model to that endpoint, specifying machine type, traffic split, and other settings. This three-step process (upload → create endpoint → deploy) is the required workflow for online serving.

Exam trap

Google Cloud often tests the misconception that you can deploy directly from Cloud Storage without the Model Registry, or that the endpoint must be created before the model is uploaded, when in fact the model must be registered first.

How to eliminate wrong answers

Option B is wrong because Vertex AI does not allow direct deployment from Cloud Storage without first registering the model in the Model Registry; the registry is required to manage model versions and associate deployment configurations. Option C is wrong because you cannot create an endpoint before uploading the model to the Model Registry, as the endpoint deployment references a model resource that must already exist. Option D is wrong because the question explicitly asks for online predictions, and batch prediction is a separate, asynchronous process that does not involve endpoints or real-time serving.

19
Multi-Selecthard

A data science team uses Cloud Build and Vertex AI to implement CI/CD for their machine learning models. Which THREE steps are essential for a production-ready operationalization pipeline? (Choose 3.)

Select 3 answers
A.Store all training artifacts in Cloud Storage without versioning.
B.Deploy the model to a staging endpoint for manual approval before promoting to production.
C.Automatically deploy every new model version directly to the production endpoint.
D.Use Vertex AI Model Evaluation to validate the new model against the current production model metrics.
E.Include unit and integration tests for the training code in the Cloud Build pipeline.
AnswersB, D, E

A staging endpoint with manual approval inserts a human gate between model build and production release, satisfying the governance constraint that unvalidated models must not reach live traffic. Vertex AI supports this via endpoint deployment and traffic splitting, enabling controlled promotion after sign-off.

Why this answer

Option B is correct because a production-ready operationalization pipeline should gate promotion with a staging endpoint and manual approval, so a human can verify the model behaves correctly before it serves live traffic. Option D is correct because Vertex AI Model Evaluation compares the candidate model's metrics against the current production model, ensuring the new version actually improves or meets quality thresholds before promotion. Option E is correct because embedding unit and integration tests for the training code in the Cloud Build pipeline catches data, feature, and code regressions early in CI/CD, which is essential for reliable automated ML delivery.

Option A is not correct because storing training artifacts in Cloud Storage without versioning breaks reproducibility and rollback, which are required for production ML pipelines. Option C is not correct because automatically deploying every new model version straight to production bypasses validation and approval, risking unvetted or degraded models serving users.

Exam trap

Google Cloud often tests the misconception that full automation (Option C) is always better, but the trap here is that production-ready pipelines require human-in-the-loop approval for critical model changes to ensure accountability and safety.

20
MCQmedium

A data engineer needs to transfer 5 PB of historical data from an on-premises Hadoop cluster to Cloud Storage. The network bandwidth is limited to 1 Gbps, and the transfer must complete within 30 days. Which transfer method should they use?

A.gsutil rsync over the internet
B.BigQuery Data Transfer Service
C.Storage Transfer Service for on-premises
D.Transfer Appliance
AnswerD

Transfer Appliance is a physical, shippable device for offline bulk migration. At 5 PB over a 1 Gbps link, online transfer would take far longer than 30 days, so shipping appliances satisfies the bandwidth and deadline constraints.

Why this answer

The Transfer Appliance is a physical device designed for large-scale data transfers when network bandwidth is insufficient. With 5 PB of data and a 1 Gbps link, the theoretical maximum transfer time is over 500 days (5 PB × 8 bits/byte / 1 Gbps / 86400 seconds/day), far exceeding the 30-day window. The Transfer Appliance bypasses network constraints by shipping data physically to Google Cloud.

Option A (gsutil rsync over the internet) is incorrect because it relies on the same 1 Gbps bandwidth, making it impossible to transfer 5 PB within 30 days. Option B (BigQuery Data Transfer Service) is designed for transferring data into BigQuery from cloud sources, not for large on-premises data transfers. Option C (Storage Transfer Service for on-premises) is intended for smaller or incremental transfers over a network, not for a 5 PB initial load within 30 days.

Exam trap

The trap here is that candidates may overestimate network transfer speeds or assume that cloud-native services like Storage Transfer Service can handle any volume, ignoring the fundamental bandwidth math that makes physical shipping the only viable option for 5 PB within 30 days.

How to eliminate wrong answers

Option A is wrong because gsutil rsync over the internet at 1 Gbps would take approximately 500 days to transfer 5 PB, which exceeds the 30-day deadline; it also lacks reliability for such massive transfers over a public network. Option B is wrong because BigQuery Data Transfer Service is designed for scheduled imports from SaaS applications (e.g., Google Ads, Amazon S3) and does not support direct on-premises Hadoop transfers. Option C is wrong because Storage Transfer Service for on-premises requires a network connection (typically via a staging bucket or partner interconnect) and still relies on the same 1 Gbps bandwidth, making it impossible to meet the 30-day requirement.

21
Multi-Selecthard

A media company stores 900 TB of video master files in a Cloud Storage bucket in the US multi-region. Legal requires that the objects be retained for seven years and cannot be deleted or overwritten by any user, including project owners, during that period. The company also wants to minimize storage cost for objects that are rarely accessed after the first 90 days. The data engineer must implement a compliant configuration. (Choose two.)

Select 2 answers
A.Set the bucket's default storage class to Standard and rely on the multi-region location for durability.
B.Enable Uniform bucket-level access and grant the storage.objectViewer role to analysts.
C.Add an Object Lifecycle Management rule that transitions objects to Coldline storage 90 days after creation.
D.Apply a bucket lock on a retention policy that sets the retention period to seven years.
E.Add a lifecycle rule that deletes objects 90 days after creation to control cost.
AnswersC, D

Coldline is designed for data accessed less than once a quarter, so transitioning the rarely accessed video masters after 90 days lowers storage cost while still allowing reads. Lifecycle rules run automatically and do not conflict with the retention policy, so this satisfies the cost requirement.

Why this answer

A locked retention policy is the only Cloud Storage mechanism that prevents even project owners from deleting or overwriting objects, providing the required WORM guarantee for seven years. Pairing it with a lifecycle transition to Coldline after 90 days meets the cost objective because Coldline targets infrequently accessed data without conflicting with retention.

Exam trap

The trap here is believing that IAM roles or object holds alone satisfy a legal WORM requirement, when only a locked retention policy makes deletion impossible for everyone.

22
MCQeasy

You need to estimate the cost of a BigQuery query before running it. Which command or feature should you use?

A.Check the BigQuery jobs list for similar queries.
B.Use the BigQuery cache to estimate if the query is cached.
C.Run EXPLAIN on the query to see the query plan.
D.Use the bq command with the --dry_run flag.
AnswerD

The --dry_run flag validates the query and returns the bytes it would process without executing it, letting you estimate on-demand cost from that byte count. This satisfies the stem's requirement to price a query before running it.

Why this answer

The `bq` command with the `--dry_run` flag allows you to estimate the amount of data a BigQuery query will process before actually executing it. This dry run does not read any data or incur charges; it simply returns the estimated bytes to be processed, which you can use to calculate the cost based on BigQuery's pricing model.

Exam trap

A common trap is confusing the EXPLAIN command (which shows the query plan) with the `--dry_run` flag (which estimates bytes processed and cost).

How to eliminate wrong answers

Option A is wrong because checking the BigQuery jobs list for similar queries only gives you historical cost data, not an estimate for the specific query you are about to run, and it assumes a similar query exists. Option B is wrong because the BigQuery cache stores results of previously run queries, but it does not provide an estimate of cost or data processed; it only indicates whether results might be served from cache. Option C is wrong because running EXPLAIN on the query shows the query plan and execution steps, but it does not provide a cost estimate or the amount of data that will be scanned.

23
MCQmedium

Your team runs a Cloud Composer 2 environment in project `analytics-prod`. A nightly DAG loads Cloud Storage files into BigQuery, but a recent Cloud Storage outage caused several tasks to fail after exhausting their retries. You want failed task instances to automatically re-run without manual intervention once the upstream dependency recovers. What should you do?

A.Set the Airflow configuration option `[core] default_task_retries` to a higher value in the environment configuration and restart the schedulers.
B.Increase the DAG's `retry_delay` and `execution_timeout` so tasks wait longer for Cloud Storage to recover before failing.
C.Configure the DAG with `catchup=True` and set `max_active_runs=1` so missed schedule intervals are backfilled.
D.Define a task-level `on_failure_callback` that calls the Airflow REST API to clear the failed task instances after a delay.
AnswerD

Clearing failed task instances through the Airflow REST API re-queues them so the scheduler runs them again once the upstream Cloud Storage dependency is healthy. Wrapping this in an `on_failure_callback` with a delay automates recovery without manual operator intervention, which is exactly the requirement after retries are exhausted.

Why this answer

Failed task instances that have exhausted retries stay in a failed state until they are cleared, at which point the scheduler re-queues them. Triggering that clear programmatically, for example from a failure callback that waits for the dependency to recover, restores the pipeline automatically. Backfill settings, retry counts, and timeout tuning only affect future scheduling or in-flight attempts, not instances already marked failed.

Exam trap

The trap here is assuming that increasing retry counts or delays will recover tasks that have already reached the failed state, when only clearing and re-queuing those task instances makes the scheduler run them again.

24
Drag & Dropmedium

Drag and drop the steps to configure a VPC network with private Google access for on-premises connectivity using Cloud VPN into the correct order.

Drag or tap steps into the slots.

Steps
Order
1Step 1
2Step 2
3Step 3
4Step 4

Why this order

Private Google Access allows on-premises hosts to reach Google APIs via VPN without public IPs.

25
MCQmedium

A production model deployed on Vertex AI Endpoint is experiencing high latency during traffic spikes. The current configuration uses a single replica. What is the most efficient solution?

A.Set a higher min replica count (e.g., 3)
B.Enable autoscaling with minReplicaCount=1 and maxReplicaCount=10
C.Use a larger machine type (e.g., n1-highmem-8)
D.Switch to batch prediction to handle spikes
AnswerB

Autoscaling adds replicas when traffic rises, distributing load across instances so per-request latency stays low, then scales back to one replica when idle. Setting minReplicaCount=1 and maxReplicaCount=10 satisfies the spike requirement efficiently without permanent over-provisioning.

Why this answer

Enabling autoscaling with minReplicaCount=1 and maxReplicaCount=10 allows Vertex AI Endpoint to dynamically add replicas during traffic spikes, distributing inference requests across multiple instances and reducing latency. This is the most efficient solution as it scales resources up only when needed, avoiding over-provisioning and minimizing cost during low traffic periods.

Exam trap

A common mistake is thinking that manually increasing the minimum replica count or using a larger machine type (static scaling) is the best way to reduce latency during spikes. However, Vertex AI Endpoint's autoscaling (with minReplicaCount=1 and maxReplicaCount=10) dynamically adjusts replicas based on traffic, avoiding over-provisioning and reducing cost. This is a key operational excellence principle in Google Cloud: design for elasticity.

How to eliminate wrong answers

Option A is wrong because setting a higher min replica count (e.g., 3) would keep three replicas running at all times, increasing cost without addressing the root cause of latency during spikes—it does not provide dynamic scaling beyond the fixed minimum. Option C is wrong because using a larger machine type (e.g., n1-highmem-8) increases per-request throughput but does not handle concurrent request bursts; a single replica, even with more memory, can still be overwhelmed by a spike, leading to queuing and high latency. Option D is wrong because batch prediction is designed for asynchronous, offline processing of large datasets and is not suitable for real-time inference; switching to batch prediction would introduce unacceptable delays for live traffic and does not solve the latency problem during spikes.

26
Multi-Selectmedium

Which TWO configurations are required to enable online prediction for a model deployed on Vertex AI Endpoints?

Select 2 answers
A.A feature store must be attached to the endpoint.
B.The endpoint must be configured with a machine type (e.g., n1-standard-2).
C.The model must be trained on Vertex AI.
D.A model must be deployed to an endpoint.
E.Autoscaling must be enabled.
AnswersB, D

Online prediction requires each deployed model to have compute resources, so the endpoint's deployed model must specify a machine type such as n1-standard-2. Without an assigned machine type, Vertex AI cannot provision serving nodes, and the endpoint cannot return synchronous predictions.

Why this answer

Option B is correct because a Vertex AI Endpoint requires a compute resource configuration, and deploying a model to an endpoint requires specifying a machine type (such as n1-standard-2) so the endpoint has the infrastructure to serve online prediction requests. Option D is correct because online prediction is only possible after a model is deployed to an endpoint; the DeployedModel resource on the endpoint is what actually receives and serves prediction traffic. Option A is incorrect because a Feature Store is not required for online prediction; it is only relevant when the model needs to retrieve features at serving time, and it is not a mandatory endpoint configuration.

Option C is incorrect because a model can be imported and deployed to Vertex AI Endpoints even if it was trained outside Vertex AI, so training on Vertex AI is not required. Option E is incorrect because autoscaling is optional; an endpoint can serve online predictions with a fixed number of replicas without enabling autoscaling.

Exam trap

The trap here is that candidates often confuse optional features (like Feature Store or autoscaling) with mandatory configurations, or assume the model must be trained on Vertex AI, when in fact only the machine type and model deployment are strictly required for online prediction.

27
MCQmedium

You are building a Dataflow pipeline in Python that reads messages from Pub/Sub, enriches them with data from a BigQuery table, and writes the results to BigQuery. The enrichment lookup table is large and changes infrequently. Which approach minimizes cost and latency?

A.Use a CoGroupByKey transform to join the incoming stream with a stream from BigQuery.
B.Use BigQuery IO to query the table for every incoming message.
C.Use a side input that reads the BigQuery table periodically and caches it.
D.Use a stateful DoFn and store the lookup in state per key.
AnswerC

A side input reads the slowly changing BigQuery enrichment table periodically and caches it in memory across workers, avoiding a per-element BigQuery lookup. This cuts both query cost and per-message latency, satisfying the large, infrequently changing lookup constraint.

Why this answer

Using a side input that periodically reads the BigQuery table and caches it avoids querying BigQuery for every incoming message, which would be prohibitively expensive and high-latency. The side input is refreshed at a configurable interval (e.g., every 10 minutes) via a pipeline option, and the cached data is broadcast to all workers, enabling fast, in-memory lookups without per-element I/O. This approach minimizes cost by reducing BigQuery API calls and minimizes latency by avoiding synchronous queries for each message.

Exam trap

Google often tests the misconception that querying BigQuery per message is acceptable in streaming pipelines, but the trap here is that candidates overlook the cost and latency implications of per-element I/O, especially with BigQuery's pricing model and query latency.

How to eliminate wrong answers

Option A is wrong because CoGroupByKey requires both inputs to be bounded or both unbounded streams; here, the BigQuery table is a bounded dataset, and Pub/Sub is unbounded, so CoGroupByKey would not work without windowing and would introduce unnecessary complexity and latency. Option B is wrong because querying BigQuery for every incoming message would cause extremely high API costs (BigQuery charges per byte processed) and high latency (each query takes hundreds of milliseconds to seconds), making it impractical for a streaming pipeline. Option D is wrong because storing the lookup in state per key would require partitioning the lookup table across keys, which is inefficient for a large, infrequently changing table; state is per-key and not shared across keys, so each worker would need to load and maintain its own copy, leading to memory waste and complex state management.

28
MCQmedium

An online gaming company runs a Dataflow streaming pipeline that aggregates player actions into per-session metrics. Sessions are defined by a gap duration of 30 minutes of inactivity, and the pipeline must emit final session results even when a session's events arrive out of order by up to 10 minutes. Late data beyond that window can be dropped. Which combination of Beam concepts should the pipeline use?

A.Fixed windows of 30 minutes with the default trigger and allowed lateness of zero.
B.Sliding windows of 30 minutes with a one-minute period and an early trigger firing every minute.
C.Session windows with a 30-minute gap, an allowed lateness of 10 minutes, and a trigger that fires after the watermark passes the end of the window.
D.Global window with a repeated element-count trigger and a discard accumulation mode.
AnswerC

Session windows merge based on inactivity gaps, so a 30-minute gap duration models the session definition exactly. Setting allowed lateness to 10 minutes keeps window state alive so out-of-order events within that bound are still incorporated, and firing on the watermark produces final results once the window closes, which matches the emit-final-results requirement.

Why this answer

Session windows are the only window type that models inactivity gaps, and combining a 30-minute gap with 10 minutes of allowed lateness keeps state open long enough for out-of-order events while dropping anything later. A watermark-based trigger emits the final result once the window is considered complete.

Exam trap

The trap here is treating the 30-minute gap as a fixed window length, when in session windows the gap is an inactivity threshold that merges windows rather than a fixed time bucket.

29
MCQmedium

Your Dataflow streaming pipeline writes to BigQuery and occasionally fails with `QuotaExceededError` on streaming inserts during peak hours. You want to reduce insert-driven quota pressure without changing the pipeline's output schema or losing exactly-once semantics. Which change should you make?

A.Add a `GroupByKey` before the BigQuery sink to batch records into larger write requests.
B.Switch the pipeline from the BigQueryIO streaming insert method to the Storage Write API with exactly-once semantics.
C.Increase the number of Dataflow workers and set `--maxNumWorkers` higher so inserts are spread across more threads.
D.Enable `--enableStreamingEngine` and raise the pipeline's `--diskSizeGb` to buffer failed inserts locally.
AnswerB

The Storage Write API uses a different quota pool than legacy streaming inserts and supports exactly-once delivery when configured with the appropriate commit strategy. Migrating the sink reduces pressure on streaming insert quotas while preserving the schema and the exactly-once guarantee the pipeline relies on, directly addressing the peak-hour failures.

Why this answer

Legacy BigQuery streaming inserts draw on a quota pool that is easy to saturate at peak throughput. The Storage Write API is a separate ingestion path with higher and separately metered limits, and its exactly-once mode uses write streams with offsets so retries do not duplicate rows. Switching the sink preserves the output schema while relieving the quota bottleneck.

Exam trap

The trap here is assuming that scaling Dataflow workers or batching requests raises the BigQuery insert quota, when the quota is enforced server-side and only a different ingestion API changes which limit applies.

30
Multi-Selecteasy

A data engineer wants to set up automatic deletion of objects from a Cloud Storage bucket after 30 days, and transition objects older than 7 days to Nearline storage. Which THREE steps should they take? (Select three.)

Select 3 answers
A.Set the bucket's default storage class to Nearline
B.Create a lifecycle rule with action Delete and condition age 30 days
C.Create a lifecycle rule with action SetStorageClass to Nearline and condition age 7 days
D.Enable Object Versioning on the bucket
E.Apply the lifecycle configuration to the bucket using gsutil lifecycle set
AnswersB, C, E

A lifecycle rule whose action is Delete with an age condition of 30 days instructs Cloud Storage to remove objects automatically once they reach that age, directly satisfying the stated requirement for deletion after 30 days without manual intervention.

Why this answer

Option B is correct because a lifecycle rule with the Delete action and an age condition of 30 days will automatically remove objects once they reach 30 days old, satisfying the deletion requirement. Option C is correct because a lifecycle rule with the SetStorageClass action targeting Nearline and an age condition of 7 days transitions objects older than 7 days to Nearline storage, exactly matching the transition requirement. Option E is correct because lifecycle rules must be applied to the bucket by submitting the JSON configuration with the gsutil lifecycle set command (or equivalent API call); creating rules without applying them has no effect.

Option A is not needed because changing the bucket's default storage class only affects newly uploaded objects and does not perform age-based transitions or deletions. Option D is not required because Object Versioning is unrelated to time-based deletion or storage-class transitions and would actually complicate deletion by retaining noncurrent versions.

Exam trap

Google Cloud Storage lifecycle rules are applied at the bucket level and can transition storage classes or delete objects based on age. A common pitfall is confusing the bucket's default storage class (which only affects new objects) with lifecycle rules (which can act on existing objects).

31
Multi-Selectmedium

An organization is using BigQuery for analytics. They have a table that is 500 GB and is frequently queried by 'date' and 'region'. They want to optimize query performance and reduce costs. Which TWO actions should they take?

Select 2 answers
A.Use an authorized view
B.Use a wildcard table
C.Use materialized views
D.Cluster the table by region
E.Partition the table by date
AnswersD, E

Clustering by region sorts data within partitions by that column, so filters on region read only relevant blocks rather than scanning everything. This satisfies the stem's region-filtering access pattern, reducing bytes billed and improving performance.

Why this answer

Option E is correct because partitioning the table by date means BigQuery only scans the partitions that match the query's date filter, drastically reducing bytes processed and cost for a 500 GB table frequently filtered by date. Option D is correct because clustering the table by region physically sorts data within each partition by region, so queries filtering on region benefit from block pruning and scan less data. Together, partitioning by date and clustering by region match the two most common query predicates and are the standard BigQuery optimization pattern.

Option A (authorized view) only controls access to data and does not improve query performance or reduce scan costs. Option B (wildcard table) is used to query multiple similarly named tables with a UNION-like syntax and does not optimize a single table's scans. Option C (materialized views) can help some workloads but is not the primary optimization for a frequently filtered base table and adds storage and refresh overhead, so it is not one of the two best actions here.

Exam trap

PDE often tests the confusion between access control features (authorized views) and performance features (partitioning, clustering); candidates pick materialized views or wildcard tables without recognizing that partitioning and clustering directly address the query patterns.

32
MCQmedium

A company uses Workflows to orchestrate a series of Google Cloud services for data processing. They need to call an external HTTP API as part of the workflow and handle potential failures with retries. Which Workflows feature should they use?

A.Retry policy on the step
B.Subworkflows
C.Parallel steps
D.Conditional steps
AnswerA

Attaching a retry policy to the HTTP call step lets Workflows automatically re-invoke the external API on transient failures, using configurable backoff and attempt limits. This directly satisfies the stem's requirement to handle potential failures with retries, without custom error-handling logic or additional orchestration components.

Why this answer

Workflows provides a built-in retry policy that can be configured on individual steps to automatically retry an HTTP call upon transient failures (e.g., 5xx server errors or network timeouts). This allows the workflow to handle external API failures without custom code, using exponential backoff and a maximum retry count.

Exam trap

Google Cloud Workflows often tests the distinction between workflow orchestration features (retry, subworkflows, parallel, conditional) and candidates mistakenly choose parallel steps or subworkflows thinking they inherently provide fault tolerance, but only a retry policy directly addresses automatic retries on failure.

How to eliminate wrong answers

Option B is wrong because subworkflows are used to encapsulate reusable sequences of steps, not to handle retries on a single HTTP call. Option C is wrong because parallel steps execute multiple branches concurrently, which does not provide retry logic for a single failing step. Option D is wrong because conditional steps (e.g., switch/if-else) control the flow based on conditions but do not automatically retry a failed HTTP request.

33
MCQmedium

Your analytics team queries a large BigQuery table `events` that is partitioned by `event_date` and clustered by `user_id`. A new dashboard runs a query filtering on `user_id = 'abc123'` but does not include a filter on `event_date`. You want to minimize the bytes scanned by this query. What should you do?

A.Add a filter on `event_date` to the query, such as `event_date >= '2024-01-01'`.
B.Use `SELECT *` instead of selecting specific columns.
C.Add a `LIMIT` clause to the query, such as `LIMIT 1000`.
D.Create a clustered table on `user_id` instead of `event_date`.
AnswerA

Partition pruning is driven by filters on the partitioning column. Without an `event_date` predicate, BigQuery must scan all partitions, even though clustering on `user_id` helps within each partition. Adding a date filter restricts the scan to relevant partitions, dramatically reducing bytes processed and cost.

Why this answer

The table is partitioned by `event_date`; without a filter on that column, BigQuery cannot prune partitions and scans the entire table. Adding a date filter enables partition pruning, reducing bytes scanned. Clustering on `user_id` helps only within the scanned partitions, so it does not replace the need for a partition filter.

Exam trap

The trap here is assuming that clustering alone can avoid scanning all partitions when the query lacks a filter on the partitioning column.

34
MCQeasy

A mobile app needs a real-time NoSQL database that supports offline sync and automatic conflict resolution. Which Google Cloud database is best suited?

A.Cloud SQL
B.Firestore
C.Cloud Bigtable
D.Cloud Spanner
AnswerB

Firestore provides offline data persistence, automatic sync, and conflict resolution for mobile and web clients.

Why this answer

Firestore is the correct choice because it is a NoSQL, real-time database with client SDKs that provide offline data persistence and automatic conflict resolution using last-write-wins semantics. Its SDKs synchronize data in the background when connectivity is restored, making it well suited for mobile apps requiring offline-first functionality.

Exam trap

A common misconception is that any NoSQL database (like Bigtable) supports mobile offline sync, but Bigtable lacks client-side SDKs and conflict resolution mechanisms, making Firestore the best fit for real-time mobile apps with offline capabilities.

How to eliminate wrong answers

Option A (Cloud SQL) is wrong because it is a relational (SQL) database that does not natively support real-time sync or offline-first mobile clients; it requires custom backend logic for conflict resolution. Option C (Cloud Bigtable) is wrong because it is a wide-column NoSQL database designed for high-throughput analytical workloads, not for real-time mobile sync or offline support; it lacks client-side SDKs for automatic conflict resolution. Option D (Cloud Spanner) is wrong because it is a globally distributed relational database with strong consistency, but it does not provide built-in offline sync or automatic conflict resolution for mobile clients; it is optimized for OLTP workloads requiring ACID transactions across regions.

35
MCQhard

A company stores JSON-formatted application logs in Cloud Storage. They need to query these logs with SQL, but they want to avoid the cost and latency of loading them into BigQuery. The logs have a consistent schema, and queries will filter on a timestamp field and a few nested fields. Which approach should they use?

A.Create an external table in BigQuery over the Cloud Storage JSON files and enable hive partitioning.
B.Create a Dataproc cluster and run Spark SQL queries against the JSON files in Cloud Storage.
C.Load the JSON files into a BigQuery native table using schema autodetect and then delete the files.
D.Use BigQuery federated queries with a Cloud SQL connection to query the JSON files.
AnswerA

BigQuery external tables can query JSON files directly in Cloud Storage without loading. Enabling hive partitioning on a timestamp-derived folder structure allows partition pruning, reducing data scanned and improving performance. This matches the need to query with SQL while avoiding load costs and leverages the consistent schema and timestamp filter.

Why this answer

BigQuery external tables allow querying JSON files in Cloud Storage without loading, and hive partitioning enables partition pruning on the timestamp field, reducing scanned data and cost. This meets the requirement to avoid load costs and latency while supporting SQL queries. Loading into a native table, federated queries, or Dataproc all introduce extra cost or complexity and do not match the serverless, direct-query need.

Exam trap

The trap here is thinking that federated queries can access Cloud Storage files, when they are actually for external databases like Cloud SQL.

36
Multi-Selectmedium

A data engineer needs to monitor the performance of BigQuery queries to identify opportunities for optimization. Which TWO metrics should they focus on? (Choose two.)

Select 2 answers
A.Slot usage
B.Data scanned per query
C.Query execution time
D.Number of tables joined
E.Number of users
AnswersA, B

Slot usage measures the compute capacity consumed by each query, directly exposing which jobs monopolise resources and where optimisation will pay off. BigQuery allocates slots dynamically, so sustained high consumption signals inefficient queries or missing partitioning, satisfying the requirement to identify optimisation opportunities.

Why this answer

Slot usage (A) is correct because BigQuery charges and throttles based on slots, and monitoring slot utilization reveals whether queries are CPU-bound, queuing, or over-consuming resources, which directly points to optimization opportunities such as reducing complexity or using reservations. Data scanned per query (B) is correct because BigQuery bills on bytes processed, and tracking scanned data highlights inefficient patterns like SELECT *, missing partition filters, or unclustered tables that can be optimized to cut cost and improve speed. Query execution time (C) is not one of the two targeted metrics here since it is an outcome rather than a direct optimization lever, and the question asks for metrics that reveal optimization opportunities.

Number of tables joined (D) is not a standard BigQuery performance metric and does not by itself indicate inefficiency. Number of users (E) is a usage/access statistic, not a query performance metric relevant to optimization.

Exam trap

There is a common misconception that query execution time is the best indicator of performance, but in BigQuery, slot usage and data scanned are more precise metrics for identifying optimization opportunities because they isolate resource consumption from external factors like caching or concurrent workloads.

37
MCQmedium

A company uses Looker Studio to build dashboards from BigQuery data. They notice that queries take several seconds to return. They want to improve performance without changing the schema or adding materialized views. Which option should they use?

A.Enable BigQuery BI Engine on the relevant project.
B.Move the data to Cloud SQL.
C.Switch to BigQuery Omni for cross-cloud queries.
D.Use APPROX_COUNT_DISTINCT to speed up distinct counts.
AnswerA

BI Engine provides an in-memory analysis layer that caches frequently accessed data, accelerating Looker Studio queries without schema changes or materialised views. It satisfies the stem's constraint of improving dashboard performance while leaving the underlying BigQuery tables untouched, and integrates natively with Looker Studio.

Why this answer

BI Engine accelerates sub-second query response times in Looker Studio by caching data in memory within the BigQuery region.

38
MCQmedium

You are designing a data quality pipeline that must inspect PII in BigQuery tables and de-identify sensitive columns before sharing with analysts. Which GCP service should you use?

A.Dataplex
B.Cloud Data Catalog
C.Cloud DLP
D.Dataflow
AnswerC

Cloud DLP natively scans BigQuery tables, using infoType detectors to identify PII and de-identification transforms such as masking, tokenisation, and format-preserving encryption to protect sensitive columns. This directly satisfies the stem's requirement to inspect and de-identify PII in place before analysts access the shared data.

Why this answer

Cloud DLP (Data Loss Prevention) is the correct choice because it is purpose-built for inspecting, classifying, and de-identifying sensitive data such as PII. It integrates natively with BigQuery via inspection jobs and de-identification templates, allowing you to scan tables for over 150 built-in infoTypes (e.g., email, SSN) and apply transformations like masking, tokenization, or encryption before sharing data with analysts.

Exam trap

The trap here is that candidates often confuse Dataplex's data governance features (like policy tags and metadata) with actual de-identification, but Dataplex cannot transform data—it only applies access controls, whereas Cloud DLP performs the actual masking or tokenization of sensitive values.

How to eliminate wrong answers

Option A is wrong because Dataplex is a data fabric service for managing, governing, and cataloging data across lakes and warehouses, but it does not perform de-identification or PII inspection itself; it can integrate with Cloud DLP for such tasks but is not the primary tool. Option B is wrong because Cloud Data Catalog is a metadata management service for discovering and tagging assets, but it lacks native de-identification capabilities and cannot transform sensitive data. Option D is wrong because Dataflow is a stream/batch processing service that can be used to build custom de-identification pipelines, but it requires manual implementation of DLP logic and is not the out-of-the-box service for inspecting and de-identifying PII in BigQuery tables.

39
MCQmedium

A company uses Cloud Dataproc to run nightly Spark ETL jobs that process about 500 GB of data each night. The jobs currently take 4 hours to complete. The company wants to reduce the runtime to under 2 hours to meet a new SLA. The cluster is configured with 10 worker nodes (n1-standard-4) and 1 master node (n1-standard-4). The jobs are CPU-bound and use only default settings. The cluster is deleted after each job and recreated. The data is stored in Cloud Storage. The company is open to increasing cost but wants the most cost-effective solution to meet the SLA. Which approach should they take?

A.Use a regional Cloud Storage bucket to improve read throughput.
B.Replace worker nodes with n1-highmem-16 instances to increase memory.
C.Increase the number of worker nodes to 20 and use preemptible VMs for half of them.
D.Change machine type to n2-standard-8 for all nodes.
AnswerC

Doubling worker nodes to 20 directly halves CPU-bound Spark runtime, satisfying the sub-2-hour SLA. Using preemptible VMs for half the workers lowers compute cost substantially, and Dataproc tolerates their preemption by rerunning lost tasks. Since the cluster is ephemeral and data resides in Cloud Storage, no persistent state is lost.

Why this answer

Adding more worker nodes (from 10 to 20) directly increases parallelism for CPU-bound Spark jobs, and using preemptible VMs for half of them reduces cost while still meeting the SLA. Since the job is CPU-bound and uses default settings, scaling horizontally with a mix of standard and preemptible VMs is the most cost-effective way to halve runtime, as Spark can efficiently distribute the workload across more cores.

Exam trap

The trap here is that candidates may assume CPU-bound jobs require faster CPUs (Option D) or more memory (Option B), but horizontal scaling with preemptible VMs is the most cost-effective way to increase parallelism in Cloud Dataproc.

How to eliminate wrong answers

Option A is wrong because using a regional Cloud Storage bucket improves data durability and availability but does not significantly increase read throughput for a single job; the bottleneck is CPU, not I/O. Option B is wrong because the job is CPU-bound, not memory-bound; increasing memory with n1-highmem-16 instances does not address the CPU bottleneck and adds unnecessary cost. Option D is wrong because changing to n2-standard-8 (8 vCPUs per node) doubles vCPUs per node but only increases total vCPUs from 40 to 80, which may not halve runtime, and is less cost-effective than using 20 n1-standard-4 nodes (80 vCPUs) with preemptible VMs for half.

40
MCQeasy

You need to schedule a simple workflow that fetches data from an API every hour, transforms it using Cloud Functions, and writes the result to Cloud Storage. The workflow has no complex branching or retry logic beyond basic retries. Which orchestration service is the MOST cost-effective and simplest to implement?

A.Cloud Scheduler
B.Workflows
C.Cloud Composer
D.Dataflow
AnswerB

Workflows is serverless and billed per step executed, so an hourly fetch-transform-write sequence with basic retries costs far less than provisioning Cloud Composer's always-running Airflow environment, and its YAML definition is simpler for linear, non-branching orchestration.

Why this answer

Workflows is a serverless, fully managed orchestration service that charges only per step executed, making it ideal for simple linear workflows with basic retries. It natively integrates with Cloud Functions and Cloud Storage via connectors, so the hourly fetch-transform-write pipeline can be defined in YAML/JSON without managing infrastructure. Cloud Scheduler alone can trigger jobs but cannot orchestrate multi-step logic or handle the transform step.

Exam trap

PDE often tests the distinction between a scheduler (Cloud Scheduler), an orchestrator (Workflows/Composer), and a data processor (Dataflow); candidates wrongly pick Cloud Scheduler because the question mentions 'every hour' and overlook the multi-step orchestration requirement.

How to eliminate wrong answers

Option A is wrong because Cloud Scheduler is only a cron-based trigger service — it can invoke a function hourly but cannot orchestrate the multi-step fetch-transform-write sequence or manage state between steps. Option C is wrong because Cloud Composer is a managed Apache Airflow service designed for complex DAGs with branching, dependencies, and heavy retry logic; it carries significant cost (environment minimums) and operational overhead that is overkill for a simple linear workflow. Option D is wrong because Dataflow is a stream/batch data processing service based on Apache Beam, not an orchestration engine — it processes data but does not schedule or coordinate API calls and Cloud Functions.

41
MCQeasy

A media company wants to analyze clickstream data stored in Cloud Storage as JSON files. They need to run SQL queries directly on these files without loading them into BigQuery. Which BigQuery feature should they use?

A.BigQuery Data Transfer Service
B.External tables
C.BigQuery Omni
D.BigQuery BI Engine
AnswerB

External tables in BigQuery allow querying data directly from Cloud Storage in formats like JSON, CSV, Avro, Parquet, and ORC. They provide a schema and can be queried with SQL without loading data into BigQuery storage. This meets the requirement of querying JSON files in GCS.

Why this answer

External tables enable BigQuery to query data stored in Cloud Storage directly, supporting JSON and other formats. This avoids data duplication and loading costs, making it ideal for ad-hoc analysis of raw files. Other features either load data into BigQuery or serve different purposes.

Exam trap

The trap here is confusing external tables with data loading services like BigQuery Data Transfer Service, which move data rather than query in place.

42
MCQeasy

A data science team has trained a TensorFlow model for image classification and wants to deploy it to production with minimal latency. They have already exported the model as a SavedModel directory. Which service should they use to create an online prediction endpoint?

A.Cloud Functions
B.Vertex AI Endpoints
C.AI Platform Prediction (legacy)
D.Cloud Dataflow
AnswerB

Vertex AI Endpoints deploys the SavedModel directory directly as an online prediction endpoint, serving TensorFlow models with low latency. It satisfies the minimal-latency constraint by hosting the model on managed infrastructure with autoscaling, avoiding the overhead of custom serving code or batch-only pipelines.

Why this answer

Vertex AI Endpoints is the correct service for deploying a TensorFlow SavedModel to an online prediction endpoint with minimal latency. It provides managed, autoscaling infrastructure optimized for real-time inference, including GPU/TPU support, request batching, and automatic health checking, which are essential for production deployment.

Exam trap

The trap here is that candidates may confuse Vertex AI Endpoints with AI Platform Prediction (legacy) or think Cloud Functions can serve models, but the Google Professional Data Engineer exam tests that Vertex AI is the modern, fully managed service for online prediction with minimal latency, while the others are either deprecated or designed for different workloads.

How to eliminate wrong answers

Option A is wrong because Cloud Functions is a serverless compute service for event-driven, short-lived functions, not designed for hosting persistent ML models with low-latency prediction endpoints; it lacks built-in model serving, batching, and autoscaling for inference workloads. Option C is wrong because AI Platform Prediction (legacy) is the older, deprecated service that has been replaced by Vertex AI; while it could serve models, it is no longer the recommended or supported path for new deployments, and Vertex AI offers superior latency optimization and integration. Option D is wrong because Cloud Dataflow is a batch and stream data processing service based on Apache Beam, intended for ETL and data pipelines, not for hosting online prediction endpoints; it cannot serve real-time inference requests with sub-second latency.

43
MCQmedium

A company runs Apache Spark jobs on Dataproc. They want to reduce costs by using preemptible instances for worker nodes. The jobs are fault-tolerant and can handle occasional node loss. However, the cluster must remain available for interactive querying during business hours. Which Dataproc cluster configuration meets these requirements?

A.Use a single-node cluster that automatically scales with preemptible instances
B.Use a standard cluster with preemptible instances as secondary workers
C.Use standard cluster with master and worker nodes as preemptible instances
D.Use a high-availability cluster with preemptible instances for primary workers
AnswerB

Secondary workers (preemptible workers) are ideal for fault-tolerant batch jobs. They do not store HDFS data, so losing them does not affect data durability. The cluster remains available because primary workers and master nodes are regular instances.

Why this answer

Dataproc supports preemptible VMs as secondary workers, which are added to a cluster alongside non-preemptible primary workers. Secondary workers handle the bulk of shuffle and task execution, so if a preemptible node is reclaimed, the job can recover while the primary workers and master keep the cluster alive and available for interactive queries. This gives the cost benefit of preemptible pricing without risking cluster or master availability.

Exam trap

The trap here is conflating 'fault-tolerant job' with 'fault-tolerant cluster' — candidates assume preemptible nodes can be used anywhere because the job can retry, forgetting that master and primary worker roles must remain non-preemptible to preserve cluster availability.

How to eliminate wrong answers

Option A is wrong because a single-node cluster has no separate master/worker separation and cannot use preemptible secondary workers — losing the node would take down the entire cluster, breaking interactive availability. Option C is wrong because making master and worker nodes preemptible risks losing the master (which terminates the cluster) and the primary workers, violating the availability requirement. Option D is wrong because high-availability mode protects the master via multiple masters, but using preemptible instances for primary workers still risks losing the workers that must remain stable for interactive querying.

44
MCQmedium

A team wants to ingest streaming data from millions of IoT devices and store historical data in BigQuery for analysis. They need near real-time analytics on the most recent data, with sub-second latency. Which architecture should they use?

A.Use Pub/Sub to receive data, then stream directly into BigQuery using the streaming API, and use standard SQL queries for real-time analytics.
B.Use Pub/Sub, then a Dataflow pipeline that filters and transforms data, writing to Cloud Bigtable for real-time queries and to Cloud Storage for periodic BigQuery loads.
C.Use Pub/Sub to ingest data into a Dataproc Spark Streaming job that writes to both Bigtable and BigQuery.
D.Use Cloud SQL to store the latest data and periodically move historical data to BigQuery via cron jobs.
AnswerB

Cloud Bigtable provides single-digit-millisecond row-key lookups, satisfying the sub-second latency requirement for recent data, while Dataflow writes the same stream to Cloud Storage for periodic BigQuery loads that serve historical analytics without straining the real-time path.

Why this answer

It uses Cloud Bigtable for sub-second latency on recent data, which is ideal for near real-time analytics on streaming IoT data. Dataflow provides the necessary stream processing, filtering, and transformation before writing to Bigtable for low-latency queries and to Cloud Storage for periodic batch loads into BigQuery for historical analysis. This architecture decouples real-time and historical paths, meeting both latency and storage requirements.

Exam trap

Google Cloud often tests the misconception that BigQuery's streaming API can provide sub-second query latency, but in reality, BigQuery is a columnar analytics engine optimized for large scans, not for low-latency point reads, which is why a separate low-latency store like Bigtable is required for real-time access.

How to eliminate wrong answers

Option A is wrong because streaming directly into BigQuery via the streaming API does not guarantee sub-second latency for queries; BigQuery is optimized for analytical queries on large datasets, not for real-time point lookups or low-latency access to the most recent data. Option C is wrong because Dataproc Spark Streaming adds unnecessary operational overhead and latency compared to a managed service like Dataflow, and writing directly to both Bigtable and BigQuery from Spark can cause contention and complexity without the built-in exactly-once semantics and auto-scaling of Dataflow. Option D is wrong because Cloud SQL is not designed for high-throughput streaming ingestion from millions of devices and cannot handle the scale; also, periodic cron jobs to move data to BigQuery introduce latency that violates the sub-second requirement for near real-time analytics.

45
Multi-Selecteasy

A company uses Cloud Logging to monitor application errors. They want to set up real-time notifications for critical errors. Which two actions are essential? (Choose two.)

Select 2 answers
A.Create a log-based metric for critical errors.
B.Export logs to BigQuery for later analysis.
C.Create a Cloud Pub/Sub notification directly on the log sink.
D.Enable VPC Flow Logs to capture network traffic.
E.Set up a Cloud Monitoring alert policy based on the log-based metric.
AnswersA, E

Cloud Monitoring cannot alert directly on raw log entries, so a log-based metric must first extract and count critical-error entries. This converts log data into a time series the alerting policy can evaluate, enabling real-time notification.

Why this answer

Option A is correct because a log-based metric is the mechanism that converts matching log entries (e.g., a filter for severity=ERROR or specific critical error text) into a numeric time series that Cloud Monitoring can evaluate; without it there is no metric for alerting. Option E is correct because a Cloud Monitoring alert policy is what actually triggers real-time notifications, and it must be based on the log-based metric created in A to fire when critical errors occur. Option B is not essential since exporting to BigQuery is for retrospective analysis, not real-time alerting.

Option C is not essential because a Pub/Sub notification on a log sink is a separate log-routing feature and does not by itself create a Monitoring alert with notification channels. Option D is not essential because VPC Flow Logs capture network traffic metadata and are unrelated to application error alerting.

Exam trap

A common trap is to think that creating a Cloud Pub/Sub notification directly on a log sink (option C) provides real-time notifications for critical errors. However, a log sink exports log entries to a destination (like Pub/Sub) for storage or processing, not for immediate alerting. Real-time notifications require a log-based metric and an alerting policy in Cloud Monitoring (options A and E).

46
MCQeasy

A team wants to retrain a model weekly using new data stored in BigQuery. They want to minimize manual effort. Which approach should they use?

A.Use Cloud Scheduler to trigger a Cloud Function that retrains
B.Retrain manually in a notebook each week
C.Use Cloud Composer to orchestrate retraining
D.Create a Vertex AI Pipeline scheduled via Cloud Scheduler
AnswerD

Scheduling a Vertex AI Pipeline with Cloud Scheduler automates the weekly retrain, satisfying the minimal-manual-effort constraint. The pipeline reads the fresh BigQuery data, retrains, and redeploys without intervention, whereas manual notebook runs or ad-hoc jobs would require recurring human triggers.

Why this answer

Vertex AI Pipelines allow you to define a repeatable, automated ML workflow that can be triggered on a schedule via Cloud Scheduler. This minimizes manual effort by handling data extraction from BigQuery, model retraining, and deployment without human intervention, while also providing versioning and monitoring capabilities.

Exam trap

Google Cloud often tests the distinction between simple scheduling (Cloud Scheduler + Cloud Function) and full ML orchestration (Vertex AI Pipelines), where candidates mistakenly choose the simpler option without considering the need for a managed, scalable ML workflow.

How to eliminate wrong answers

Option A is wrong because Cloud Scheduler triggering a Cloud Function is suitable for lightweight tasks, but retraining a model typically requires more complex orchestration, dependency management, and resource handling that a Cloud Function alone cannot efficiently provide. Option B is wrong because manual retraining in a notebook each week introduces significant manual effort and is error-prone, directly contradicting the goal of minimizing manual effort. Option C is wrong because Cloud Composer (based on Apache Airflow) is a powerful orchestration tool but is overkill for a simple weekly retraining schedule; it adds unnecessary complexity and cost compared to a Vertex AI Pipeline scheduled via Cloud Scheduler.

47
MCQeasy

You need to load a 500 GB CSV file from Cloud Storage into BigQuery. The file has a header row and uses comma delimiters. You want to load it as quickly as possible without transforming the data. Which approach should you use?

A.Create an external table pointing to the CSV and then run `CREATE TABLE AS SELECT * FROM external_table`.
B.Use `bq query` with an `EXTERNAL_QUERY` function to read the CSV from Cloud Storage.
C.Use the `gcloud storage cp` command to copy the CSV into BigQuery.
D.Use the `bq load` command with `--source_format=CSV` and `--skip_leading_rows=1`.
AnswerD

The `bq load` command with `--source_format=CSV` and `--skip_leading_rows=1` directly loads the CSV, skipping the header. This is the fastest method for bulk loading without transformation, leveraging BigQuery's native load job which can parallelize reading from Cloud Storage.

Why this answer

A direct load job using `bq load` with CSV format and skipping the header is the fastest and most straightforward way to load a large CSV into BigQuery without transformation. It uses BigQuery's native loading capabilities, which are optimized for bulk data.

Exam trap

The trap here is confusing external tables or copy commands with a direct load job, which is specifically designed for efficient bulk loading.

48
MCQhard

A team uses dbt on BigQuery to transform data in their data warehouse. They have a large table with nested and repeated fields (arrays and structs). The transformation needs to normalize this data into a star schema. Which dbt feature and BigQuery SQL feature should they use together?

A.dbt hooks with BigQuery STRUCT access
B.dbt models with BigQuery UNNEST and CROSS JOIN
C.dbt snapshots with BigQuery JSON functions
D.dbt seeds with BigQuery ARRAY_AGG
AnswerB

UNNEST flattens BigQuery arrays and structs into individual rows, while CROSS JOIN multiplies each parent row across its unnested elements. Together inside dbt models they convert nested, repeated source data into the normalised fact and dimension tables of a star schema.

Why this answer

To normalize nested and repeated fields (arrays and structs) into a star schema, you need to flatten the arrays into separate rows. BigQuery's UNNEST operator, when used with CROSS JOIN, expands each array element into its own row, effectively denormalizing the nested structure. dbt models (SQL SELECT statements) are the correct dbt feature to define these transformations as version-controlled, reusable SQL files. Together, they allow you to write a dbt model that uses CROSS JOIN UNNEST to produce dimension and fact tables from a single nested table.

Exam trap

A common pitfall in this question is confusing the purpose of dbt features: hooks (automation), snapshots (SCD type 2), and seeds (CSV loading) are not designed for flattening nested structures. The correct approach is to use dbt models with BigQuery's UNNEST and CROSS JOIN to expand arrays into rows.

How to eliminate wrong answers

Option A is wrong because dbt hooks are SQL or shell commands executed at specific points in the dbt run (e.g., before/after model builds) and are not designed for transforming nested data into a star schema; BigQuery STRUCT access alone cannot flatten arrays. Option C is wrong because dbt snapshots are used for slowly changing dimension (SCD) tracking over time, not for normalizing nested data; BigQuery JSON functions are for parsing JSON strings, not for unnesting native arrays and structs. Option D is wrong because dbt seeds are CSV files loaded into the warehouse as static lookup tables, not for transforming existing data; ARRAY_AGG is an aggregation function that creates arrays, the opposite of the flattening needed here.

49
MCQhard

A healthcare company stores patient documents in a Cloud Storage bucket with a retention policy. Auditors require that once an object is written, it cannot be overwritten or deleted by any user, including project owners, for 7 years. The company also wants to minimize storage cost after the first year while preserving immutability. What should the data engineer configure?

A.Enable uniform bucket-level access and create a retention policy without locking it, then transition objects to Coldline after 365 days.
B.Apply a bucket lock on a retention policy with a 7-year retention period, then add a lifecycle rule to transition objects to Archive after 365 days.
C.Set an IAM deny policy on storage.objects.delete for all principals and use a lifecycle rule to move objects to Nearline after 365 days.
D.Enable Object Versioning and set a lifecycle rule to move objects to Coldline after 365 days.
AnswerB

A locked retention policy enforces WORM semantics: objects cannot be deleted or overwritten until the retention period expires, and even project owners cannot remove the lock. A lifecycle rule can still transition the storage class to Archive after one year, reducing cost while the retention policy continues to protect the data for the full seven years.

Why this answer

A locked retention policy is the only Cloud Storage mechanism that enforces immutability against all users, including project owners, for the specified duration. Combining it with a lifecycle rule that transitions objects to Archive after one year satisfies the retention requirement while lowering storage cost for the remaining six years.

Exam trap

The trap here is confusing permission-based deletion prevention, such as IAM deny policies, with a locked retention policy that enforces WORM even against project owners.

50
MCQhard

A logistics company runs Cloud SQL for MySQL for its order management system. The database is 2 TB and growing 100 GB per month. They need to run analytical queries without impacting transactional performance, and they want to minimize operational overhead. They also require the analytics data to be no more than 15 minutes stale. Which storage approach should they use?

A.Export the Cloud SQL database to a Cloud Storage bucket nightly with mysqldump, then load the files into BigQuery each morning.
B.Enable the Cloud SQL federated query feature in BigQuery and query the Cloud SQL instance directly from BigQuery.
C.Create a read replica of the Cloud SQL instance and run analytical queries directly against the replica.
D.Configure a Datastream stream from Cloud SQL to BigQuery with a change data capture policy and query the replicated tables in BigQuery.
AnswerD

Datastream replicates changes from Cloud SQL for MySQL to BigQuery using change data capture, keeping the analytics copy fresh within minutes. BigQuery handles large analytical scans without affecting the transactional database and requires minimal operational management. This satisfies the 15-minute staleness requirement and the need to avoid impacting order management performance, while minimizing overhead compared to managing replicas.

Why this answer

Datastream provides serverless change data capture from Cloud SQL for MySQL into BigQuery, keeping the analytics copy within minutes of the source. BigQuery isolates analytical scans from the transactional instance, so order management performance is unaffected. The managed service minimizes operational overhead compared to maintaining replicas or custom export jobs, and it satisfies the 15-minute freshness requirement through continuous replication.

Exam trap

The trap here is assuming a read replica is sufficient for analytics, when heavy scans can still cause replication lag and resource contention.

51
Multi-Selectmedium

Which THREE steps are required to set up a continuous training pipeline on Google Cloud using Vertex AI?

Select 3 answers
A.Run training on a single Compute Engine VM with a cron job.
B.Create a Vertex AI Pipeline to orchestrate data preprocessing, training, and model evaluation.
C.Set up a trigger (e.g., Cloud Scheduler or Cloud Build) to start training on a schedule or new data.
D.Manually upload the model to Vertex AI Model Registry after each training run.
E.Configure model evaluation and promotion rules (e.g., if accuracy > threshold, deploy to endpoint).
AnswersB, C, E

A Vertex AI Pipeline orchestrates the preprocessing, training, and evaluation components as a repeatable directed acyclic graph, which is what makes the pipeline continuous rather than a one-off run. This directly satisfies the stem's requirement to automate the full training workflow on Google Cloud.

Why this answer

Option B is correct because a Vertex AI Pipeline is the native orchestration mechanism that chains the preprocessing, training, and evaluation components into a repeatable, versioned workflow, which is the backbone of a continuous training pipeline. Option C is correct because continuous training requires an automated trigger — Cloud Scheduler for time-based runs or Cloud Build (often via Eventarc/pub-sub on new data) — to invoke the pipeline without manual intervention. Option E is correct because continuous training must include automated model evaluation and promotion logic, such as comparing accuracy against a threshold and deploying to a Vertex AI endpoint only when the gate passes, enabling CI/CD-style model rollout.

Option A is not appropriate because a single Compute Engine VM with cron is a manual, non-managed approach that bypasses Vertex AI's managed pipelines, scaling, and metadata tracking. Option D is not appropriate because manually uploading models to the Vertex AI Model Registry after each run defeats automation; the pipeline should register models programmatically as a step.

Exam trap

Google Cloud often tests the distinction between manual, ad-hoc automation (like cron jobs) and fully managed, integrated orchestration services (like Vertex AI Pipelines), leading candidates to incorrectly select simpler but non-scalable options.

52
Multi-Selecthard

Which TWO are common causes of prediction bias in a deployed machine learning model in production?

Select 2 answers
A.Model accuracy is too high.
B.Data drift between training and serving data distributions.
C.Model is overfitted to training data.
D.Low latency predictions.
E.Training-serving skew due to differences in feature engineering.
AnswersB, E

Data drift means serving inputs diverge statistically from the training distribution, so learned relationships no longer hold and predictions skew systematically. This satisfies the bias-cause constraint because the model extrapolates beyond its training support, producing skewed outputs without any code change.

Why this answer

Option B is correct because data drift occurs when the statistical distribution of the input features or target changes between the training environment and the live serving environment, causing the model's learned mappings to become stale and its predictions to be systematically biased. Option E is correct because training-serving skew arises when feature engineering logic differs between the training pipeline and the production inference path (for example, different imputation, normalization, or aggregation code), so the model receives inputs at serving time that do not match what it learned from, producing biased predictions. Option A is not a cause of bias — high accuracy is generally desirable and does not by itself indicate or produce prediction bias.

Option C is not correct in this context because overfitting primarily harms generalization and variance rather than being a canonical cause of prediction bias in production, and it is a training-time issue rather than the drift/skew mechanisms described. Option D is not correct because low latency is a performance characteristic of the serving system and has no direct causal relationship with prediction bias.

Exam trap

Google Cloud often tests the distinction between training-time issues (like overfitting) and production-time causes (like data drift and training-serving skew), so candidates mistakenly select overfitting as a production bias cause.

53
MCQmedium

An organization wants to integrate BigQuery Omni to query data stored in AWS S3. They have set up the necessary connections. What is the primary benefit of using BigQuery Omni over simply copying the data to BigQuery?

A.Ability to use BigQuery ML models on data in S3 without moving data.
B.Automatic encryption of data at rest in S3.
C.Lower latency queries due to in-memory caching.
D.Support for real-time streaming inserts into S3.
AnswerA

BigQuery Omni supports BigQuery ML, allowing you to train and run models on cross-cloud data.

Why this answer

BigQuery Omni allows you to query data across clouds without moving it, providing a unified analytics experience. It reduces data egress costs and avoids duplication.

54
Multi-Selectmedium

A company uses Pub/Sub to ingest events from multiple sources. They need to ensure that messages from a specific source are processed in order (per source partition). They also need to deduplicate messages. Which TWO features should they use?

Select 2 answers
A.Set a message schema to enforce ordering
B.Use a dead letter topic to handle out-of-order messages
C.Enable exactly-once delivery on the subscription
D.Use a pull subscription with a large ack deadline
E.Enable message ordering by setting an ordering key
AnswersC, E

Exactly-once delivery on the subscription removes duplicate messages by tracking acknowledgement state, satisfying the deduplication requirement. However, it does not guarantee ordering; that needs the ordering key or message ordering property set on the publisher side, so this feature alone covers only one of the two stated constraints.

Why this answer

Option E is correct because Pub/Sub message ordering is enabled by setting an ordering key on messages; messages sharing the same ordering key are delivered to subscribers in the order they were published, which satisfies the per-source-partition ordering requirement. Option C is correct because enabling exactly-once delivery on the subscription ensures that messages are not redelivered after successful acknowledgment, providing the required deduplication behavior. Option A is incorrect because schemas validate message format and structure, not ordering.

Option B is incorrect because a dead letter topic only captures messages that cannot be successfully processed after retries, and does not reorder or deduplicate messages. Option D is incorrect because a large ack deadline only extends the processing window to reduce premature redelivery; it does not guarantee ordering or exactly-once deduplication.

Exam trap

The trap is confusing schema enforcement or dead letter topics with ordering and deduplication — candidates pick schema (A) thinking it 'enforces order' or DLQ (B) thinking it 'handles out-of-order messages,' when only ordering keys and exactly-once delivery address the actual requirements.

55
MCQmedium

A data engineer is building a Looker Studio dashboard that requires a calculated field to compute the running total of sales per day per store. Which Looker Studio function should they use?

A.RANK()
B.TOTAL()
C.RUNNING_SUM()
D.PERCENTILE()
AnswerC

RUNNING_SUM computes cumulative sums.

Why this answer

Looker Studio's RUNNING_SUM function computes a running total within a group, exactly what is needed for a running total per store partitioned by date.

56
MCQhard

A financial services firm stores trade documents in a Cloud Storage bucket. Auditors require that each object be retained for exactly seven years and that no user, including project owners, be able to delete or overwrite the objects during that period. The firm also needs to prove compliance to auditors. Which combination of controls should the data engineer implement?

A.Create a bucket lock on a retention policy of seven years and grant auditors the Storage Object Viewer role.
B.Apply IAM conditions that deny storage.objects.delete unless the request comes from an auditor's service account.
C.Set an Object Lifecycle Management rule to delete objects after 2555 days and enable uniform bucket-level access.
D.Enable Cloud Storage versioning and set a lifecycle rule to keep noncurrent versions for seven years.
AnswerA

A bucket retention policy prevents deletion or replacement of objects until the retention period elapses, and locking the policy makes it permanent and irreversible even for project owners. Granting auditors Storage Object Viewer lets them inspect objects and the bucket metadata to verify the lock. This combination enforces seven-year immutability and provides the evidence auditors need.

Why this answer

A locked bucket retention policy enforces object immutability for the specified duration and cannot be removed or shortened, satisfying the requirement that even project owners cannot delete data. Because the lock is a permanent bucket property, auditors can verify it directly. Lifecycle rules, versioning, and IAM conditions are all modifiable controls and therefore cannot guarantee the required immutability.

Exam trap

The trap here is treating Cloud Storage versioning or a lifecycle rule as equivalent to WORM retention, when only a locked retention policy prevents deletion by privileged users.

57
Multi-Selectmedium

A media company stores 50 TB of video metadata in Cloud Storage buckets in the Standard class. Legal requires that records for a subset of titles be retained for exactly seven years and that they cannot be altered or deleted by any user, including project owners, during that period. The company also wants to minimize storage cost for a second bucket holding derived thumbnails that are accessed only a few times per year but must be retrievable within seconds. Which two configurations should the data engineer implement? (Choose two.)

Select 2 answers
A.Enable Object Versioning on the metadata bucket to preserve prior versions of each object.
B.Set an Object Lifecycle Management rule to move the metadata objects to Archive storage after one year.
C.Create the thumbnails bucket in the Coldline storage class and enable Turbo Replication.
D.Create a bucket-level retention policy with a seven-year retention period and lock it.
E.Create the thumbnails bucket in the Nearline storage class.
AnswersD, E

A locked bucket retention policy prevents any user, including project owners, from deleting or overwriting objects until the retention period elapses, which is exactly the immutability legal requires. Locking the policy makes it permanent and irreversible, so the seven-year period cannot be shortened, satisfying the WORM-style guarantee for the retained titles while keeping objects readable.

Why this answer

A locked bucket retention policy delivers the immutability legal demands because it cannot be removed once locked and blocks deletion or overwrite by any principal for the full seven years. For the thumbnails, Nearline matches an access pattern of a few reads per year with second-level retrieval and a lower price than Standard, avoiding the longer minimum-duration and slower-access profile that Coldline or Archive would impose.

Exam trap

The trap here is confusing versioning or lifecycle tiering with true immutability, when only a locked retention policy blocks even project owners from deleting or altering objects.

58
Multi-Selectmedium

A company has a data lake on Cloud Storage with raw data in the 'raw' bucket, curated data in 'curated', and processed data in 'processed'. They want to implement lifecycle management to reduce costs. Which TWO actions should they take? (Choose 2)

Select 2 answers
A.Set a lifecycle rule to change storage class from Standard to Nearline after 30 days for the 'raw' bucket.
B.Enable object versioning on all buckets to automatically delete older versions.
C.Set a partition expiration on BigQuery tables that reference data in the 'processed' bucket.
D.Set a lifecycle rule to delete objects older than 365 days in the 'curated' and 'processed' buckets.
E.Set a lifecycle rule to change storage class from Standard to Archive after 30 days for the 'raw' bucket.
AnswersA, D

Raw data is typically accessed rarely after initial ingestion, so transitioning it from Standard to Nearline at 30 days cuts storage cost while preserving availability. This satisfies the stem's cost-reduction goal for the 'raw' bucket without affecting curated or processed data.

Why this answer

Option A is correct because a lifecycle rule that transitions objects in the 'raw' bucket from Standard to Nearline after 30 days is a valid cost-reduction strategy for raw data that is accessed infrequently but may still be needed; Nearline is designed for data accessed less than once a month. Option D is correct because setting a lifecycle rule to delete objects older than 365 days in the 'curated' and 'processed' buckets removes stale data that is no longer needed, directly reducing storage costs. Option B is incorrect because enabling object versioning does not automatically delete older versions; versioning preserves them and typically increases storage costs unless combined with a separate lifecycle deletion rule.

Option C is incorrect because partition expiration applies to BigQuery table partitions, not to objects in a Cloud Storage bucket, and the scenario is about Cloud Storage lifecycle management. Option E is incorrect because transitioning raw data to Archive after only 30 days is overly aggressive and would incur early-deletion charges and retrieval costs if the data is still needed, making it a poor fit compared with Nearline.

Exam trap

PDE often tests the misconception that enabling object versioning deletes old versions automatically, and the trap of choosing Archive too early without accounting for its 365-day minimum duration and retrieval costs.

59
Multi-Selectmedium

A company is designing a Cloud Bigtable row key for a time-series dataset of device readings. They want to avoid hotspotting (uneven load across tablets). Which TWO row key design patterns are effective? (Choose 2)

Select 2 answers
A.Use a monotonically increasing counter as row key
B.Use timestamp directly as the first part of the row key
C.Reverse the timestamp string
D.Prepend a hash of the device ID to the timestamp
E.Use a secondary index on the timestamp column
AnswersC, D

Reversing the timestamp string makes the least significant digits the row key prefix, spreading sequential writes across many row ranges rather than concentrating them on one tablet. This satisfies the anti-hotspotting constraint for time-series ingestion.

Why this answer

Option C is correct because reversing the timestamp string (e.g., turning a lexicographically increasing timestamp into a decreasing one) prevents new writes from always landing on the last tablet, spreading sequential writes across the keyspace. Option D is correct because prepending a hash of the device ID to the timestamp distributes writes across many distinct key prefixes, so consecutive readings from different devices do not concentrate on a single tablet. Options A and B are incorrect because a monotonically increasing counter and a timestamp as the leading key component both produce sequential, ever-increasing keys that funnel all new writes to the last tablet, causing hotspotting.

Option E is incorrect because a secondary index on the timestamp column is not a row key design pattern and Bigtable does not support secondary indexes natively; it would not address hotspotting.

60
MCQmedium

Refer to the exhibit. A Dataflow streaming pipeline subscribes to this Pub/Sub subscription. The pipeline occasionally takes more than 10 seconds to process a message. Which behavior will occur?

A.The message will be sent to the dead letter topic immediately.
B.The message will be retried with exponential backoff as per retry policy.
C.The message will be redelivered after 10 seconds if not acknowledged.
D.The message will be dropped after 10 seconds due to expiration policy.
AnswerC

The subscription's acknowledgement deadline is 10 seconds, so a message still unacknowledged when that window expires becomes eligible for redelivery. Because processing occasionally exceeds 10 seconds, the pipeline may receive duplicates, requiring idempotent handling or an extended deadline to avoid repeated processing.

Why this answer

Pub/Sub delivery requires an acknowledgment within the configurable `ackDeadlineSeconds` (default 10 seconds). If the pipeline takes longer than the ack deadline to process a message, Pub/Sub considers the message unacknowledged and redelivers it. This is the standard behavior for at-least-once delivery in Google Cloud Pub/Sub.

Exam trap

Google Cloud often tests the distinction between ack deadline expiration and dead letter topics, trapping candidates who assume any processing delay immediately triggers a dead letter or that Pub/Sub uses exponential backoff like some other messaging systems.

How to eliminate wrong answers

Option A is wrong because a dead letter topic is only triggered after a message has been retried the maximum number of times (configurable via `maxDeliveryAttempts`), not immediately upon exceeding the ack deadline. Option B is wrong because Pub/Sub does not use exponential backoff for redelivery; it uses a fixed or configurable `ackDeadlineSeconds` and redelivers after that deadline expires, with no built-in exponential backoff retry policy. Option D is wrong because the expiration policy (`messageRetentionDuration`) controls how long unacknowledged messages are retained in the subscription, not a 10-second drop; messages are retained for up to 7 days by default.

61
MCQeasy

A company wants to version its ML models and track lineage from training data to deployed model. Which Google Cloud service should they use?

A.Cloud Storage with object versioning
B.Data Catalog
C.Artifact Registry
D.Vertex AI ML Metadata
AnswerD

Vertex AI ML Metadata stores and queries artefacts, executions and contexts, giving the lineage graph linking training datasets to models and deployments. It is the purpose-built tracking store, unlike generic Cloud Storage or BigQuery, satisfying the versioning and lineage requirement.

Why this answer

Vertex AI ML Metadata is the Google Cloud service designed to track ML artifacts, executions, and contexts, providing model versioning and lineage from training data through to deployed models. It is the native metadata store for Vertex AI pipelines and experiments.

Exam trap

PDE often tests whether candidates confuse storage versioning (Cloud Storage), data cataloging (Data Catalog), and artifact storage (Artifact Registry) with ML-specific lineage tracking (Vertex AI ML Metadata) — the trap is picking a general-purpose service for an ML-specific requirement.

How to eliminate wrong answers

Option A is wrong because Cloud Storage object versioning only versions individual objects (files) and provides no ML-specific lineage tracking between datasets, training runs, and models. Option B is wrong because Data Catalog is a metadata management service for discovering and tagging data assets, not for tracking ML model lineage or versioning models. Option C is wrong because Artifact Registry stores and manages container images and language packages, not ML metadata or lineage relationships.

62
MCQhard

You are building a Dataflow pipeline that reads from Cloud Storage and writes to BigQuery. The pipeline must handle files that are compressed with gzip and contain JSON data. You need to ensure that the pipeline can process these files efficiently and write to BigQuery with minimal errors. Which approach should you take?

A.Use TextIO to read the gzip files, parse the JSON using a DoFn, and write to BigQuery using BigQueryIO with the STORAGE_WRITE_API method.
B.Use TextIO to read the gzip files, but disable compression detection and manually decompress each file in a DoFn before parsing JSON.
C.Use AvroIO to read the gzip files, convert the Avro records to JSON, and write to BigQuery using BigQueryIO with the FILE_LOADS method.
D.Use BigQueryIO to read the gzip files directly, parse the JSON, and write to BigQuery using the default write method.
AnswerA

TextIO can read gzip-compressed files transparently, and a DoFn can parse the JSON into a TableRow. Writing with BigQueryIO using the STORAGE_WRITE_API method provides high-performance, exactly-once writes. This combination efficiently processes compressed JSON files and handles errors through dead-letter patterns if configured, meeting the requirements.

Why this answer

TextIO natively supports reading gzip-compressed files, simplifying ingestion. Parsing JSON in a DoFn allows transformation into BigQuery-compatible rows. Using BigQueryIO with the STORAGE_WRITE_API method ensures high-throughput, exactly-once writes, which is critical for efficient and reliable loading.

This combination addresses the compressed format and minimizes errors through robust write semantics.

Exam trap

The trap here is assuming that BigQueryIO can read files or that manual decompression is necessary, when TextIO already handles gzip seamlessly.

63
MCQmedium

A company uses Cloud Spanner and needs to store a parent-child relationship where the child table is frequently queried together with the parent. The parent has millions of rows and the child billions. Which Spanner feature optimizes performance for this pattern?

A.Partitioned tables
B.Interleaved tables
C.Secondary indexes
D.Change streams
AnswerB

Interleaved tables physically co-locate child rows with their parent row in the same storage split, so parent-child joins avoid network hops. This satisfies the pattern where the child table is queried alongside its parent, since Spanner reads both from one locality rather than performing distributed lookups.

Why this answer

Interleaved tables in Cloud Spanner physically co-locate child rows with their parent rows on the same split, so a parent-child join can be satisfied with a single local lookup rather than a distributed join across nodes. This is ideal when the child table is frequently queried together with the parent, as in this scenario with millions of parents and billions of children. The interleaving declaration (INTERLEAVE IN PARENT) makes the child's primary key prefix the parent's key, enabling efficient prefix scans and reduced network hops.

Exam trap

The trap here is confusing interleaved tables with secondary indexes or partitioned tables; candidates often pick secondary indexes thinking they optimize joins, but only interleaving provides physical co-location for parent-child access patterns.

How to eliminate wrong answers

Option A is wrong because partitioned tables (partitioned DML or table partitioning concepts) are about managing large-scale data operations and DML, not about co-locating related parent-child rows for join performance. Option C is wrong because secondary indexes improve query performance for specific access patterns but do not co-locate child rows with parents; they add storage and write overhead without solving the parent-child locality problem. Option D is wrong because change streams capture data modifications for downstream processing (e.g., analytics, event-driven apps) and have nothing to do with optimizing parent-child query performance.

64
MCQeasy

Which BigQuery SQL function returns the rank of a row within a window, with gaps in the ranking for ties?

A.RANK()
B.NTILE()
C.DENSE_RANK()
D.ROW_NUMBER()
AnswerA

RANK() assigns positions ordered by the window's ORDER BY, and tied values receive the same rank, leaving gaps before the next rank. It satisfies the explicit gap requirement, unlike DENSE_RANK(), which numbers ties consecutively without gaps.

Why this answer

The RANK() function assigns a rank to each row within a window, and when there are ties, it leaves gaps in the ranking sequence. For example, if two rows tie for rank 1, the next row receives rank 3. This behavior distinguishes it from DENSE_RANK(), which does not leave gaps.

RANK() is part of BigQuery's window functions and is used with an OVER clause specifying the partitioning and ordering.

Exam trap

PDE often tests the subtle differences between RANK(), DENSE_RANK(), and ROW_NUMBER(); candidates may pick DENSE_RANK() because they forget that RANK() leaves gaps, or pick ROW_NUMBER() because they overlook the tie-handling requirement.

How to eliminate wrong answers

Option B is wrong because NTILE() divides the rows into a specified number of roughly equal buckets and assigns a bucket number, not a rank with gaps. Option C is wrong because DENSE_RANK() assigns ranks without gaps for ties — if two rows tie for rank 1, the next row gets rank 2. Option D is wrong because ROW_NUMBER() assigns a unique sequential number to each row, regardless of ties, so it never produces gaps or duplicate ranks.

65
MCQeasy

Which Dataflow feature automatically scales the number of workers based on the pipeline's current workload, and also selects the optimal machine type for each worker based on the pipeline's resource requirements?

A.Dataflow Shuffle
B.Dataflow Prime
C.Dataflow Streaming Engine
D.Dataflow Flex Templates
AnswerB

Dataflow Prime provides vertical autoscaling, dynamically resizing worker machine types to match each stage's resource demands, alongside horizontal worker-count scaling. This satisfies the stem's dual requirement: automatic worker scaling plus optimal machine-type selection based on the pipeline's resource requirements.

Why this answer

Dataflow Prime is the correct answer because it is the only Dataflow feature that provides both automatic worker scaling (horizontal autoscaling) and intelligent machine type selection (vertical autoscaling). It dynamically adjusts the number of workers based on the pipeline's current workload and selects the optimal machine type (e.g., CPU, memory, or accelerator-optimized) for each worker based on the pipeline's resource requirements, such as CPU utilization, memory pressure, or shuffle throughput.

Exam trap

Google often tests the distinction between horizontal autoscaling (adding/removing workers) and vertical autoscaling (changing machine type), and the trap here is that candidates assume Dataflow Shuffle or Streaming Engine handle scaling, when in fact they only optimize specific pipeline phases (shuffle or state management) without affecting worker count or machine type.

How to eliminate wrong answers

Option A is wrong because Dataflow Shuffle is a service that separates the shuffle operation from worker VMs, improving scalability and reliability, but it does not handle worker scaling or machine type selection. Option C is wrong because Dataflow Streaming Engine moves state storage and computation away from worker VMs for streaming pipelines, reducing resource overhead, but it does not automatically scale workers or select machine types. Option D is wrong because Dataflow Flex Templates allow you to package and reuse pipeline code with custom container images, but they do not provide any autoscaling or machine type optimization; scaling is handled separately by the Dataflow service.

66
MCQeasy

A data pipeline ingests streaming data from Pub/Sub into BigQuery via Dataflow. Recently, the pipeline has been failing with 'deadline exceeded' errors. What is the most likely cause?

A.The BigQuery streaming quota is exceeded.
B.Dataflow workers are underutilized due to batch size settings.
C.Dataflow autoscaling is disabled.
D.The Pub/Sub subscription's acknowledgement deadline is too short for the processing time.
AnswerD

Pub/Sub redelivers a message when its acknowledgement deadline expires before the subscriber acknowledges it; Dataflow then surfaces deadline exceeded errors. If processing consistently takes longer than the configured deadline, the subscription's acknowledgement window is too short for the actual processing time.

Why this answer

'deadline exceeded' errors in a Dataflow pipeline reading from Pub/Sub indicate that the subscriber is taking longer to process messages than the acknowledgement deadline allows. When the deadline expires, Pub/Sub redelivers the message, causing duplicate processing and eventual pipeline failure. This is a common issue when processing time exceeds the default 10-second acknowledgement deadline.

Exam trap

Google Cloud often tests the distinction between resource quota errors (like BigQuery streaming quota) and Pub/Sub-specific timeout errors, trapping candidates who confuse 'deadline exceeded' with general quota exhaustion.

How to eliminate wrong answers

Option A is wrong because BigQuery streaming quota exceeded would produce 'quota exceeded' or 'rate limit exceeded' errors, not 'deadline exceeded' errors. Option B is wrong because underutilized workers due to batch size settings would cause poor performance or backpressure, not 'deadline exceeded' errors; the error is about processing time vs. acknowledgement deadline, not worker utilization. Option C is wrong because disabled autoscaling would lead to resource exhaustion or latency, but the specific 'deadline exceeded' error is tied to Pub/Sub's acknowledgement mechanism, not Dataflow's scaling behavior.

67
MCQhard

Your team is processing a large dataset with Apache Beam on Dataflow. The pipeline sometimes fails due to transient errors when writing to a BigQuery sink. You need to ensure that failed records are not lost and can be reprocessed later without blocking the pipeline. What is the best approach?

A.Configure the pipeline to use at-least-once semantics and rely on Dataflow to retry the entire bundle.
B.Increase the number of workers to reduce the chance of transient errors.
C.Use a try-catch block in the DoFn and log the error; continue processing other elements.
D.Use a side output (e.g., via TupleTag) to write failed records to a dead letter sink (e.g., GCS or Pub/Sub) and continue processing the main output.
AnswerD

A TupleTag side output diverts records that fail the BigQuery write into a dead letter sink while the main pipeline continues. This satisfies both constraints: failed records are retained for later reprocessing, and the pipeline is not blocked by transient sink errors.

Why this answer

Using a dead letter pattern with a side output to write failed records to a GCS bucket (or Pub/Sub) allows the pipeline to continue processing healthy records while failed records are stored for later analysis and reprocessing.

68
MCQeasy

A retail analytics team must move 30 TB of Parquet files from an on-premises Hadoop cluster into BigQuery once, then run standard SQL dashboards. The transfer window is 48 hours and the source cluster has limited outbound bandwidth. Which approach should the data engineer choose?

A.Export the Parquet files to encrypted portable drives and use the Transfer Appliance to ship them to a Google ingest location, then load the objects into BigQuery.
B.Use the Storage Transfer Service with a transfer job configured against the on-premises source, then load the resulting Cloud Storage objects into BigQuery using a load job.
C.Create a Dataproc cluster in the same region and use a Spark job to read from the on-premises Hadoop cluster over the network and write directly into BigQuery.
D.Run gcloud storage cp from a Compute Engine VM to upload the files into a BigQuery-managed bucket, then query them with an external table.
AnswerA

Transfer Appliance is purpose-built for bulk offline migration when network bandwidth makes online transfer impractical. Shipping encrypted appliances avoids the limited outbound link entirely and comfortably moves 30 TB within the window. After Google uploads the appliance contents to Cloud Storage, a BigQuery load job reads the Parquet objects into native tables, after which dashboards query managed storage with full columnar performance.

Why this answer

When the data volume is large and the source's outbound bandwidth is the constraint, an offline transfer avoids the network bottleneck. Transfer Appliance ships encrypted storage to Google, where the data lands in Cloud Storage, and a BigQuery load job then ingests the Parquet files into native managed tables. Dashboards subsequently query columnar managed storage rather than external objects, so both the migration deadline and query performance requirements are met.

Exam trap

The trap here is defaulting to a network-based copy or a Spark job when the stated constraint is limited outbound bandwidth within a fixed time window, which points to offline transfer.

69
MCQeasy

What is the primary purpose of Vertex AI Feature Store?

A.To manage and track ML experiments
B.To train machine learning models using AutoML
C.To transform raw data into features using SQL
D.To store and serve features for machine learning models at scale
AnswerD

Vertex AI Feature Store provides a centralised repository that stores feature values and serves them online at low latency or in bulk for training, keeping features consistent between training and serving. This satisfies the need to store and serve features at scale.

Why this answer

Vertex AI Feature Store is a managed service for storing, serving, and sharing ML features at scale, providing low-latency online serving and high-throughput batch serving with point-in-time correctness. It centralizes feature definitions so training and serving use consistent feature values, reducing training-serving skew.

Exam trap

PDE often tests the distinction between Feature Store (feature storage/serving) and other Vertex AI components like Experiments, AutoML, and Pipelines — candidates pick a component that sounds related but serves a different purpose.

How to eliminate wrong answers

Option A is wrong because managing and tracking ML experiments is the role of Vertex AI Experiments (and ML Metadata), not Feature Store. Option B is wrong because training models with AutoML is handled by Vertex AI AutoML (Tabular, Vision, etc.), not Feature Store. Option C is wrong because transforming raw data into features using SQL is done in BigQuery or Dataflow, not Feature Store — Feature Store ingests already-computed features.

70
MCQhard

A company uses Dataplex to manage data quality across multiple BigQuery datasets. They want to define a data quality rule that checks if a column 'email' contains a valid email format. Which Dataplex feature should they use?

A.Use Cloud DLP to classify and validate emails.
B.Use the built-in 'email' rule type in Dataplex.
C.Create a custom Data Quality rule using the 'regex' type.
D.Create a Dataflow pipeline to validate emails and write results to a separate table.
AnswerC

Dataplex data quality rules support a regex rule type, letting you define a pattern that each value in the email column must match. This satisfies the requirement to validate email format without writing custom SQL assertions.

Why this answer

Dataplex Data Quality tasks support a set of built-in rule types (range, non-null, uniqueness, set, regex, sql_assertion, row_condition), and email format validation is not one of the built-in types. The 'regex' rule type lets you supply a regular expression that each value in the 'email' column must match, so a pattern like ^[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Za-z]{2,}$ enforces valid email formatting. This is the intended Dataplex mechanism for pattern-based validation.

Exam trap

PDE often tests the assumption that Dataplex ships a rich library of semantic built-in rules (email, phone, SSN); in reality only generic rule types exist, so pattern validation must be expressed as a regex rule.

How to eliminate wrong answers

Option A is wrong because Cloud DLP is a data discovery, classification, and de-identification service (infoTypes, masking, tokenization) — it does not define or execute Dataplex data quality rules on BigQuery columns. Option B is wrong because Dataplex Data Quality has no built-in 'email' rule type; the built-ins are range, non-null, uniqueness, set, regex, sql_assertion, and row_condition. Option D is wrong because building a Dataflow pipeline is a custom, out-of-band solution that bypasses Dataplex's native Data Quality task framework and does not integrate with Dataplex scan results or dashboards.

71
MCQeasy

A company needs to process real-time clickstream data and store it in a data warehouse for SQL-based analytics. The data volume is moderate. Which combination of Google Cloud services is most cost-effective?

A.Cloud Pub/Sub, Cloud Dataproc, Cloud Storage
B.Cloud Pub/Sub, Cloud Dataflow, Cloud Spanner
C.Cloud Pub/Sub, Cloud Dataflow, BigQuery
D.Cloud Pub/Sub, Cloud Dataflow, Cloud Storage
AnswerC

Pub/Sub ingests the clickstream, Dataflow performs streaming transformation, and BigQuery stores the results for SQL analytics. This serverless trio scales automatically and bills per use, suiting moderate volume without provisioning idle capacity, unlike Dataproc or Bigtable alternatives.

Why this answer

Cloud Pub/Sub ingests real-time clickstream data, Cloud Dataflow processes it with low latency, and BigQuery provides a serverless, SQL-based data warehouse that is cost-effective for moderate data volumes due to its pay-per-query pricing and automatic scaling. This combination avoids the overhead of managing clusters (Dataproc) or expensive storage (Cloud Spanner) while directly supporting SQL analytics.

Exam trap

Google Cloud often tests the misconception that Cloud Storage is a suitable destination for analytics-ready data, but it lacks native SQL querying, forcing candidates to overlook BigQuery's direct integration with Dataflow for real-time analytics.

How to eliminate wrong answers

Option A is wrong because Cloud Dataproc requires a running cluster (even with preemptible VMs) and is optimized for batch processing, not real-time streaming, and Cloud Storage is not a SQL-queryable data warehouse, forcing additional ETL steps. Option B is wrong because Cloud Spanner is a globally distributed, strongly consistent relational database designed for transactional workloads, not cost-effective for analytics at moderate data volumes; its per-node pricing makes it expensive compared to BigQuery's serverless model. Option D is wrong because Cloud Storage is an object store, not a data warehouse; storing processed data there would require additional services (e.g., BigQuery external tables or Dataproc) to run SQL analytics, increasing complexity and cost.

72
MCQmedium

You are building a forecasting model to predict daily sales for the next 90 days using historical sales data with clear seasonality and trend. You want to use BigQuery ML with minimal manual tuning. Which model type should you choose?

A.ARIMA
B.Boosted tree (XGBoost)
C.ARIMA_PLUS
D.Linear regression
AnswerC

ARIMA_PLUS automatically handles seasonality, trend and holiday effects, and performs automatic model selection and hyperparameter tuning within BigQuery ML. This satisfies the minimal manual tuning constraint while producing accurate 90-day daily sales forecasts from historical data.

Why this answer

ARIMA_PLUS is the correct choice because it is a Google-developed, automated time-series model in BigQuery ML that handles seasonality, trend, holidays, and outliers with minimal manual tuning. It automatically performs preprocessing such as missing value imputation, holiday effect detection, and seasonality decomposition, making it ideal for forecasting daily sales over a 90-day horizon. Unlike standard ARIMA, ARIMA_PLUS requires no manual specification of p, d, q parameters or seasonal orders, aligning with the 'minimal manual tuning' requirement.

Exam trap

PDE often tests the misconception that any time-series model (like ARIMA) automatically handles seasonality and trend, but only ARIMA_PLUS in BigQuery ML offers automated preprocessing and multiple seasonality detection, so candidates may incorrectly choose standard ARIMA or a non-time-series model like XGBoost.

How to eliminate wrong answers

Option A is wrong because standard ARIMA in BigQuery ML requires manual selection of non-seasonal and seasonal parameters (p, d, q, P, D, Q) and does not automatically handle multiple seasonalities, holidays, or outliers, demanding significant tuning. Option B is wrong because boosted trees (XGBoost) are not designed for time-series forecasting; they treat data as independent observations, ignore temporal ordering, and cannot extrapolate trends beyond the training range, making them unsuitable for forecasting future values. Option D is wrong because linear regression assumes a linear relationship and cannot capture complex seasonality or non-linear trends without extensive feature engineering, and it also fails to model temporal dependencies.

73
Multi-Selecthard

A healthcare analytics group must build a pipeline that ingests HL7 messages from an on-premises interface engine, must retain raw messages for seven years for compliance, and must expose de-identified aggregates to analysts. The security team requires that protected health information never be written to a dataset analysts can query, and that all data be encrypted with keys the organization manages and can revoke. Which two design choices satisfy these requirements? (Choose two.)

Select 2 answers
A.Load the raw messages directly into a BigQuery table that analysts can query, and rely on column-level security to mask the PHI columns.
B.De-identify the messages in Dataflow, write only the de-identified aggregates into a separate BigQuery dataset secured with CMEK and policy tags, and grant analysts access only to that dataset.
C.Land raw HL7 messages in a Cloud Storage bucket configured with a customer-managed encryption key (CMEK) and a retention lock, and restrict access with IAM to the ingestion service account only.
D.Store the raw messages in BigQuery and use authorized views so analysts query a view that filters out PHI columns instead of the base table.
E.Encrypt the raw messages with a customer-supplied key in the on-premises interface engine and upload only the ciphertext to Cloud Storage, decrypting in Dataflow for de-identification and discarding the key afterward.
AnswersB, C

Performing de-identification in Dataflow before the data reaches BigQuery ensures PHI never enters an analyst-queryable dataset, satisfying the hard boundary. Writing aggregates to a distinct dataset protected by CMEK gives the organization revocable key control, and policy tags add fine-grained access restrictions. Analysts are granted access only to this curated dataset, keeping the raw zone entirely separate.

Why this answer

The requirements demand a hard separation between raw PHI and anything analysts can query, plus organization-managed revocable encryption and seven-year retention. Landing raw messages in Cloud Storage with CMEK and a retention lock, accessible only to the ingestion service account, satisfies retention and key control. De-identifying in Dataflow and publishing only aggregates to a separate CMEK-protected BigQuery dataset keeps PHI out of analyst-facing storage entirely, so both boundaries hold independently.

Exam trap

The trap here is treating masking or authorized views over a raw BigQuery table as equivalent to never writing PHI into an analyst-queryable dataset, when the requirement demands physical separation, not query-time filtering.

74
Multi-Selecthard

A media company processes video metadata using a Dataflow pipeline. They need to join two streaming sources: user activity (Pub/Sub) and video catalog updates (Pub/Sub). Which THREE transforms should be used in the pipeline?

Select 3 answers
A.Flatten to combine the two PCollections
B.ParDo to process each element individually
C.CoGroupByKey to join the two PCollections on a common key (e.g., video_id)
D.Window both PCollections into a common window (e.g., fixed 1-minute)
E.GroupByKey on each PCollection separately before joining
AnswersB, C, D

ParDo is needed to process each element and extract the common key (e.g., video_id) before joining.

Why this answer

To join two streaming Pub/Sub sources in Dataflow, you need to: (1) Use ParDo to extract the key from each element (e.g., video_id), (2) Window both PCollections into a common window (e.g., fixed 1-minute) to align the data, and (3) Use CoGroupByKey to join on the common key. GroupByKey separately is not required because CoGroupByKey internally groups by key. Flatten is used to combine PCollections of the same type, which is not applicable here.

Exam trap

Candidates often think GroupByKey is needed before CoGroupByKey, but CoGroupByKey does the grouping internally.

75
Multi-Selecthard

Which THREE considerations are important when designing a batch prediction pipeline for a large dataset on Vertex AI?

Select 3 answers
A.Batch prediction automatically uses GPUs if the model framework requires them
B.Batch prediction requires a dedicated real-time endpoint
C.Choosing the appropriate machine type (e.g., n1-standard-16) balances cost and throughput
D.Large input files can be split into multiple smaller files to improve parallelism
E.Input data should be in Cloud Storage in a format supported by Vertex AI (e.g., JSONL, TFRecord)
AnswersC, D, E

Batch prediction throughput and cost scale with the machine type chosen for the prediction nodes; larger types process more rows concurrently but bill higher. Selecting an appropriate machine type therefore directly satisfies the stem's requirement to balance cost against throughput for a large dataset.

Why this answer

Option C is correct because the machine type selected for a Vertex AI batch prediction job (for example n1-standard-16) directly determines the CPU, memory, and accelerator resources available, so it must be sized to balance cost against throughput for a large dataset. Option D is correct because Vertex AI batch prediction shards input across worker nodes, and splitting very large input files into multiple smaller files (e.g., many shards in Cloud Storage) increases parallelism and reduces per-file processing bottlenecks. Option E is correct because batch prediction reads input from Cloud Storage and requires a supported format such as JSONL, CSV, or TFRecord with the appropriate instance keys, so the data must be staged there in a compatible layout.

Option A is not correct because batch prediction does not automatically attach GPUs based on the model framework; you must explicitly configure accelerators, and many frameworks run fine on CPU. Option B is not correct because batch prediction is an offline, asynchronous job that does not require or use a dedicated real-time endpoint, which is only needed for online prediction.

Exam trap

Google Cloud often tests the misconception that batch prediction requires a real-time endpoint or automatically uses GPUs, when in fact batch prediction is a serverless, endpoint-free process that requires explicit machine type and GPU configuration.

Page 1 of 10

Page 2

All pages