Courseiva

Google Professional Data Engineer (PDE) — Questions 301–375

747 questions total · 10pages · All types, answers revealed

Page 4

Page 5 of 10

Page 6
301
MCQhard

Your team is migrating a batch ETL job from an on-premises Hadoop cluster to Dataproc. The job reads CSV files from Cloud Storage, joins them with a slowly changing dimension table in BigQuery, and writes aggregated results back to BigQuery. The on-premises job used Hive on Tez and took six hours. You need to reduce runtime on Dataproc while minimizing cost. Which approach should you take?

A.Run the job as a Dataproc Serverless for Spark batch, reading the CSV files from Cloud Storage and using the BigQuery connector to read the dimension table, with autoscaling enabled.
B.Create a Dataproc cluster with a fixed number of standard workers, install Hive, and run the same Hive on Tez script to preserve the existing logic.
C.Create a long-running Dataproc cluster with many preemptible workers and run the job with Spark SQL, caching the BigQuery dimension table as a Spark DataFrame.
D.Export the CSV files to a temporary BigQuery table, perform the join and aggregation entirely in BigQuery SQL, and schedule the query with Cloud Scheduler.
AnswerA

Dataproc Serverless for Spark provisions resources on demand, scales with the workload, and shuts down when the batch completes, so you pay only for the execution time. The BigQuery connector reads the dimension table efficiently, and Spark's optimizer can broadcast the dimension if it is small. This combination reduces runtime versus the legacy Tez job and minimizes cost by avoiding an idle cluster.

Why this answer

Dataproc Serverless for Spark runs batch workloads on ephemeral, autoscaling infrastructure that is released when the batch finishes, so cost tracks actual execution rather than idle cluster time. Reading CSV from Cloud Storage and using the BigQuery connector for the dimension table lets Spark optimize the join, often broadcasting the small dimension. This reduces runtime compared with the legacy Tez engine and avoids paying for a cluster that sits idle between runs.

Exam trap

The trap here is defaulting to a persistent cluster with preemptible workers for cost savings, when preemption can extend runtime and an always-on cluster bills for idle time.

302
MCQmedium

A data pipeline ingests streaming events into Pub/Sub and needs to join them with a slowly updating reference table (few thousand rows) from a Cloud Storage CSV file. The pipeline runs on Dataflow with Apache Beam. Which approach is most cost-effective and operationally simple?

A.Read the CSV in a DoFn and perform a BigQuery query each time an event is processed
B.Use a side input that reads the CSV once and broadcasts it to all workers
C.Implement a custom sink that writes events to Cloud SQL and performs a SQL JOIN there
D.Use CoGroupByKey to join the stream and batch PCollections by a common key after reading the CSV into a batch PCollection each window
AnswerB

A side input reads the CSV once and broadcasts the small reference table to every worker, so each streaming element joins locally without per-element Cloud Storage reads or external lookups. This satisfies the cost-effectiveness and operational-simplicity constraints for a few-thousand-row table, avoiding the latency and expense of querying an external store per event.

Why this answer

Option B is correct because a Beam side input reads the small Cloud Storage CSV once and broadcasts the reference table to all workers, letting each streaming element be enriched in memory without per-event external calls, which is both cheap and operationally simple. Since the table has only a few thousand rows, it easily fits in worker memory as a side input. Option A is wrong because issuing a BigQuery query per event is slow, expensive, and adds an external dependency.

Option C is wrong because routing events through Cloud SQL and doing SQL JOINs adds infrastructure and cost. Option D is wrong because CoGroupByKey requires reading the CSV into a batch PCollection per window and shuffling both sides, which is more complex and less efficient than a side input for a small, slowly changing table.

303
MCQmedium

You need to create a time-series forecast for inventory demand using BigQuery ML. The data includes daily sales for 5 years. Which model type should you use?

A.K-means
B.Linear regression
C.ARIMA_PLUS
D.Matrix factorization
AnswerC

ARIMA_PLUS handles seasonality, trends and holiday effects automatically, which suits five years of daily sales data containing weekly and yearly patterns. It also supports forecasting directly in BigQuery ML without exporting data, satisfying the requirement to build a time-series forecast for inventory demand.

Why this answer

BigQuery ML supports ARIMA_PLUS for time-series forecasting. Linear regression, k-means, and matrix factorization are not appropriate for time-series forecasting.

304
MCQeasy

A streaming Dataflow pipeline needs to be updated without draining the existing pipeline. Which update strategy should be used?

A.Drain the pipeline first, then start a new one
B.Replace the job with a new job using the same pipeline name
C.Use a different pipeline name and cancel the old one
D.Stop the job, update the code, and restart
AnswerB

A replacement job cannot be updated in place, so the pipeline is relaunched under the same name, letting the new job start while the old one is cancelled. This avoids the drain step that would stop ingestion, satisfying the no-drain constraint.

Why this answer

Dataflow supports replacing a running streaming job with an updated one using the same pipeline name via the 'Replace job' (or update) mechanism, which performs an in-place upgrade that preserves the pipeline's state and does not require draining. This is the supported path for updating streaming pipelines without stopping data ingestion. Draining, cancelling, or stopping the job all interrupt processing and are not the recommended no-drain update strategy.

Exam trap

The trap is assuming any update requires stopping or draining the pipeline; candidates pick the 'safe-sounding' drain option without realizing Dataflow's replace-job feature is specifically designed to avoid draining.

How to eliminate wrong answers

Option A is wrong because draining stops the pipeline from ingesting new data and waits for in-flight data to finish, which is the opposite of updating without draining. Option C is wrong because using a different pipeline name and cancelling the old one creates a brand-new job with no state continuity and causes a gap in processing. Option D is wrong because stopping the job, editing code, and restarting is a manual stop/start cycle that halts the pipeline and loses the seamless update semantics Dataflow provides.

305
Multi-Selectmedium

A data engineering team is building a CI/CD pipeline for machine learning models using Cloud Build and AI Platform. Which TWO practices are essential for ensuring reproducible and safe model deployments?

Select 2 answers
A.Use Cloud Functions to trigger retraining on new data arrival.
B.Tag each model version with the Git commit hash of the training code.
C.Run integration tests against the model on a staging endpoint before promoting to production.
D.Use the same environment for training and serving, possibly via custom containers.
E.Directly deploy from the development environment using gcloud commands.
AnswersB, C

Tagging each model version with the training code's Git commit hash creates a traceable link between artefact and source, so any deployment can be reproduced exactly. This satisfies the reproducibility requirement of the CI/CD pipeline by pinning the code revision that produced the model.

Why this answer

Option B is correct because tagging each model version with the Git commit hash of the training code creates an immutable link between the deployed artifact and the exact source code that produced it, which is fundamental for reproducibility and auditability in ML CI/CD. Option C is correct because running integration tests against the model on a staging endpoint validates the model's behavior, input/output schema, and serving performance in an environment that mirrors production before promotion, catching regressions safely. Option A is not essential here since Cloud Functions triggering retraining addresses automation of retraining, not reproducibility or safe deployment of models.

Option D is not essential and can even be risky, because sharing the same environment for training and serving is not required for reproducibility and may introduce dependency conflicts; custom containers are a separate concern. Option E is incorrect because deploying directly from the development environment with gcloud commands bypasses CI/CD controls, testing, and versioning, which undermines safe and reproducible deployments.

Exam trap

A common trap in this question is confusing best practices for consistency (such as using the same environment for training and serving) with essential practices for reproducibility and safety (such as version tagging and staged testing) in Google Cloud's AI Platform CI/CD pipelines. Candidates often select Option D because it is a good practice, but it is not explicitly required for reproducibility and safety as defined by the question.

306
MCQeasy

A company stores raw data files in Cloud Storage in a bucket named 'raw-data'. After processing, the files are moved to a 'processed' bucket. To reduce costs, they want to automatically delete raw data older than 30 days. What should they do?

A.Enable object versioning on the 'raw-data' bucket and configure a lifecycle rule to delete noncurrent versions.
B.Configure a lifecycle rule on the 'raw-data' bucket to delete objects older than 30 days.
C.Set a retention policy on the 'raw-data' bucket to expire objects after 30 days.
D.Use a bucket policy that denies read access to objects older than 30 days.
AnswerB

A Cloud Storage lifecycle rule with an Age condition of 30 days automatically deletes objects in the 'raw-data' bucket, matching the stated retention requirement. This is server-side, bucket-level management, so no pipeline code, cron job or manual intervention is needed.

Why this answer

Cloud Storage lifecycle management allows you to set a rule that automatically deletes objects after a specified number of days from their creation time. By configuring a lifecycle rule on the 'raw-data' bucket to delete objects older than 30 days, the company can achieve cost reduction without manual intervention. This directly addresses the requirement to remove raw data files that have been processed and are no longer needed.

Exam trap

The trap here is confusing lifecycle deletion rules with retention policies or versioning: candidates often think retention policies delete data after a period, but they actually prevent deletion, while versioning with noncurrent deletion only removes old versions, not the current object.

How to eliminate wrong answers

Option A is wrong because enabling object versioning and deleting noncurrent versions does not delete the current (original) objects; it only removes older versions, so raw data files would remain in the bucket indefinitely. Option C is wrong because a retention policy (e.g., using Object Hold or Retention Policy) prevents deletion or modification of objects for a specified duration, which would keep the data for at least 30 days, not delete it after 30 days. Option D is wrong because a bucket policy that denies read access does not delete the objects; the files would still exist and incur storage costs, failing to meet the cost-reduction goal.

307
MCQeasy

A media company stores 400 TB of compressed JSON clickstream logs in a Cloud Storage bucket. Analysts need to run ad hoc SQL over this data several times a day, and the team wants to avoid the cost and delay of loading it into BigQuery. Which BigQuery capability should you use?

A.Load the JSON files into a BigQuery native table partitioned by ingestion time.
B.Create an external table over the Cloud Storage files using the BigLake connection, and query it directly.
C.Use the Storage Transfer Service to copy the bucket into a second bucket in the same region as the dataset.
D.Create a BigQuery dataset and use the bq command-line tool to stream the JSON records with the insertAll API.
AnswerB

BigQuery external tables backed by a BigLake connection let you run standard SQL over data that remains in Cloud Storage, so the clickstream logs are queried in place without an ingestion job. This satisfies the ad hoc querying need and avoids duplication of the 400 TB, while the connection provides fine-grained access control and metadata caching.

Why this answer

External tables over Cloud Storage, accessed through a BigLake connection, keep the data in its original bucket while exposing it to BigQuery SQL. This removes both the load latency and the duplicate storage cost that the team is trying to avoid. Native loading, bucket-to-bucket copying, and streaming inserts all either duplicate the data or provide no query surface at all.

Exam trap

The trap here is treating the BigLake connection as a data movement mechanism, when it is actually a governance and metadata layer that lets BigQuery read the Cloud Storage objects in place without copying them.

308
Multi-Selecthard

A company uses BigQuery for analytics on petabyte-scale data. They want to improve query performance by denormalizing schemas and reducing joins. Which TWO BigQuery features should they use? (Choose 2)

Select 2 answers
A.Clustering on frequently filtered columns
B.Using subqueries instead of JOINs
C.External tables reading from Cloud Storage
D.Table partitioning by date
E.Nested and repeated fields (ARRAY<STRUCT<...>>)
AnswersA, E

Clustering physically co-locates rows sharing the clustered column values, so filters on those columns scan far fewer blocks. This directly supports denormalisation by pruning data before joins, cutting bytes read and improving performance on petabyte-scale tables.

Why this answer

Option E is correct because BigQuery's native support for nested and repeated fields (ARRAY<STRUCT<...>>) lets you store parent-child relationships inside a single table, which is the standard way to denormalize schemas and eliminate joins while preserving relational structure. Option A is correct because clustering on frequently filtered columns physically co-locates related rows within partitions, so BigQuery scans far less data and returns results faster for filtered queries on the denormalized table. Option B is not a BigQuery feature for denormalization; subqueries still execute as separate query blocks and do not remove join costs.

Option C is wrong because external tables read data from Cloud Storage and generally offer worse performance than native tables, not better denormalized performance. Option D is wrong because partitioning by date improves pruning on time filters but does not denormalize schemas or reduce joins.

Exam trap

A common mistake is to think that only clustering or partitioning can replace denormalization, but nested/repeated fields are needed for schema denormalization. Clustering is a complementary physical optimization.

309
MCQeasy

A startup is building a data lake on Google Cloud. They need to store raw JSON, CSV, and Parquet files from various sources. The files will be accessed by multiple analytics tools, including BigQuery and Dataproc. The startup wants a cost-effective, durable, and highly available storage solution that integrates natively with these services. Which Google Cloud service should they use?

A.Cloud SQL for MySQL with a large SSD and binary large object (BLOB) storage.
B.Cloud Bigtable with a column family for each data format.
C.Persistent Disk attached to a Compute Engine instance, shared via NFS.
D.Cloud Storage with a Standard storage class.
AnswerD

Cloud Storage is the foundational object store for data lakes on Google Cloud. It provides durable, highly available, and cost-effective storage with native integration into BigQuery, Dataproc, and other analytics services. The Standard class is suitable for frequently accessed data, and lifecycle policies can transition older data to colder classes to optimize cost.

Why this answer

Cloud Storage is the correct choice for a data lake because it is a durable, highly available, and cost-effective object store that integrates natively with BigQuery, Dataproc, and other Google Cloud analytics services. It supports all file formats and can be organized into buckets with lifecycle policies for cost optimization. The Standard storage class is appropriate for frequently accessed data.

Exam trap

The trap here is considering Bigtable or Cloud SQL because they are managed services, but they are not designed for storing raw files in a data lake architecture and lack native integration with analytics tools for file-based access.

310
MCQmedium

You are building a Dataflow pipeline that reads messages from Pub/Sub, performs windowed aggregations, and writes results to BigQuery. The pipeline must handle late-arriving data, but after a certain point, you want to stop accepting late data to allow the pipeline to emit final results. Which Apache Beam concept should you use to define when to stop accepting late data?

A.Side input
B.Allowed lateness
C.Watermark
D.Trigger
AnswerB

Allowed lateness specifies how long after the watermark passes the end of a window the pipeline will still accept and process late data. After this period, late data is dropped. This is exactly what is needed to bound the waiting period for late-arriving events, enabling final results to be emitted. It is a key parameter in windowing strategies.

Why this answer

Allowed lateness defines the maximum time after the watermark that late data will still be processed. Once this period elapses, the pipeline drops late data and can finalize results. This directly addresses the requirement to stop accepting late data after a certain point, while still accommodating some delay.

Other concepts like watermarks and triggers manage timing but do not enforce a cutoff for late data.

Exam trap

The trap here is confusing the watermark (which estimates completeness) with allowed lateness (which explicitly sets a deadline for late data).

311
MCQhard

Your company ingests streaming data into BigQuery using the Storage Write API. You need to ensure that duplicate records are not inserted when the streaming job retries due to transient errors. Which feature should you use?

A.Use the Storage Write API's exactly-once semantics by setting a stream name and offset.
B.Create a unique constraint on the destination table to reject duplicate rows.
C.Enable BigQuery's streaming inserts with insertId to deduplicate records.
D.Use a MERGE statement to upsert records after each streaming batch.
AnswerA

The Storage Write API provides exactly-once delivery semantics when you use a stream name and specify offsets for each record. This ensures that retried records are not duplicated. It is designed for high-throughput streaming and supports deduplication natively, making it the correct choice for preventing duplicates on retries.

Why this answer

The Storage Write API's exactly-once semantics, enabled by specifying a stream name and offsets, guarantee that records are not duplicated even if the writer retries. This is the only option that provides native deduplication for streaming inserts. Other methods either do not apply to the Storage Write API or are not supported in BigQuery.

Exam trap

The trap here is confusing the legacy streaming API's insertId deduplication with the Storage Write API's exactly-once semantics.

312
MCQmedium

Your company uses Vertex AI Pipelines to automate model retraining. The pipeline has three steps: data extraction from BigQuery, feature engineering using Dataflow, and model training using a custom container on Vertex AI Training. Recently, the pipeline has been failing intermittently at the Dataflow step with a 'The job encountered a transient error. Please retry.' message. You have enabled pipeline retries with 3 attempts. However, the pipeline still fails after 3 retries. You check the logs and find that the Dataflow job requires more resources than the default worker configuration provides. Which change should you make to reduce the failure rate?

A.Increase the number of Dataflow workers to improve parallelism
B.Increase the number of retries in the pipeline to 5
C.Replace Dataflow with Dataproc to run the feature engineering step
D.Increase the Dataflow worker machine type to have more memory and CPU in the pipeline step configuration
AnswerD

The Dataflow job fails because default workers lack memory and CPU, and retries cannot fix a resource shortfall. Increasing the worker machine type gives the step sufficient capacity, addressing the logged root cause and reducing the intermittent failures.

Why this answer

The pipeline fails due to insufficient resources (memory and CPU) in the default Dataflow worker configuration. By increasing the worker machine type (e.g., using a custom machine type with more vCPUs and memory), the Dataflow job can handle the feature engineering workload without hitting resource limits, reducing transient failures. This directly addresses the root cause identified in the logs, unlike retries or parallelism changes.

Exam trap

Google Cloud often tests the misconception that increasing parallelism (more workers) or retries will fix resource exhaustion errors, when the actual fix is to increase per-worker resources by selecting a larger machine type.

How to eliminate wrong answers

Option A is wrong because increasing the number of workers improves parallelism but does not address the root cause of insufficient per-worker resources (memory/CPU); it may even increase resource contention. Option B is wrong because increasing retries from 3 to 5 does not fix the underlying resource constraint; the job will continue to fail on each retry if the worker configuration remains inadequate. Option C is wrong because replacing Dataflow with Dataproc is an unnecessary architectural change that introduces new operational complexity and does not solve the specific resource issue; the problem is with worker sizing, not the service itself.

313
MCQeasy

A data science team has built a model using scikit-learn. They want to operationalize it on Google Cloud without rewriting the code. Which approach should they take?

A.Export the model as a PMML file and use BigQuery ML
B.Use AI Platform Training to host the model directly
C.Package the model in a custom container and deploy to Vertex AI Endpoints
D.Convert the scikit-learn model to TensorFlow SavedModel format
AnswerC

Custom containers preserve the existing scikit-learn code and dependencies unchanged, satisfying the no-rewrite constraint. Vertex AI Endpoints accepts any container implementing the prediction server contract, so the team packages their model and serves it without porting to a native framework.

Why this answer

Vertex AI Endpoints support custom containers, allowing you to package your scikit-learn model with its dependencies (e.g., a Flask or FastAPI inference server) and deploy it without rewriting any code. This approach directly meets the requirement to operationalize the existing model on Google Cloud without modification.

Exam trap

Google Cloud often tests the misconception that AI Platform Training can host models directly, but it is strictly for training jobs, not serving endpoints; candidates confuse the training service with the prediction service.

How to eliminate wrong answers

Option A is wrong because PMML (Predictive Model Markup Language) is not natively supported by BigQuery ML; BigQuery ML uses SQL-based model creation and does not import PMML files for inference. Option B is wrong because AI Platform Training is designed for training jobs, not for hosting models as endpoints; hosting is done via AI Platform Prediction (now part of Vertex AI), but even then, scikit-learn models require a custom prediction routine or container, not direct hosting. Option D is wrong because converting a scikit-learn model to TensorFlow SavedModel format would require rewriting the model's inference logic and dependencies, contradicting the requirement to avoid code changes.

314
MCQhard

A data engineer is migrating a Composer 1 environment to Cloud Composer 2 and notices that a DAG relying on a legacy operator for a deprecated service no longer works. The engineer wants a durable fix that keeps the pipeline running and avoids repeating this problem in future upgrades. Which action should be taken?

A.Downgrade the environment's Python version to match the requirements of the legacy operator.
B.Move the deprecated logic into a PythonOperator that calls the service's current client library, and add unit tests to validate the behaviour.
C.Pin the Composer 2 environment to the exact Airflow version that still includes the legacy operator and avoid upgrades.
D.Copy the legacy operator's source into the DAGs folder so it is imported locally instead of from the provider package.
AnswerB

Replacing the removed operator with a PythonOperator that uses the service's supported client library removes the dependency on deprecated code and keeps the DAG working. Adding unit tests guards against regressions during future upgrades, giving a durable, maintainable solution rather than a temporary workaround.

Why this answer

Replacing deprecated operators with a PythonOperator that calls the service's current client library removes reliance on unsupported code and keeps the DAG functional on Composer 2. Adding tests ensures the replacement behaves correctly and catches regressions during future upgrades, which is the durable, maintainable outcome the scenario asks for.

Exam trap

The trap here is treating version pinning or vendoring deprecated code as a fix, when it only postpones the incompatibility.

315
MCQhard

You are using Dataproc to run a Spark job that reads data from Cloud Storage, performs aggregations, and writes results back to Cloud Storage. The job is failing with out-of-memory errors on the shuffle. Which optimization should you apply?

A.Increase spark.sql.shuffle.partitions
B.Use RDDs instead of DataFrames
C.Increase spark.executor.memory
D.Decrease the number of executors
AnswerA

Raising `spark.sql.shuffle.partitions` splits shuffle data into more, smaller partitions, so each executor task holds less in memory during aggregation. This directly relieves the out-of-memory condition on the shuffle stage, satisfying the stem's constraint without changing cluster size or storage layout.

Why this answer

Out-of-memory errors during a shuffle in Spark are usually caused by too few shuffle partitions, which makes each partition (and its in-memory sort/aggregation buffer) too large. Increasing spark.sql.shuffle.partitions splits the shuffle into more, smaller partitions, reducing per-task memory pressure and allowing the job to complete.

Exam trap

PDE often tests the reflex to 'add more memory' for any OOM — candidates pick C, but shuffle OOM is a partitioning/skew problem, and the correct first lever is spark.sql.shuffle.partitions, not executor memory.

How to eliminate wrong answers

Option B is wrong because RDDs are lower-level than DataFrames and lack Catalyst/Tungsten optimizations — switching to RDDs would generally increase memory usage and slow the job, not fix shuffle OOM. Option C is wrong because increasing spark.executor.memory can help only if the executor itself is the bottleneck; shuffle OOM is typically caused by skewed or oversized partitions, and simply adding memory often just delays the failure or causes GC pressure. Option D is wrong because decreasing executors reduces total cluster parallelism and memory available for the shuffle, making OOM more likely, not less.

316
MCQmedium

A company is designing a data lake on Cloud Storage with three zones: raw, curated, and processed. They need to enforce data governance by restricting access to each zone using IAM. Which approach should they take?

A.Create a single bucket with folders for each zone, and use IAM conditions to restrict access
B.Use Cloud Storage lifecycle rules to move objects between zones
C.Use a single bucket and rely on object ACLs
D.Create three separate buckets, one per zone, and assign IAM roles per bucket
AnswerD

Separate buckets per zone let IAM policies and roles be scoped at bucket level, giving clean, enforceable isolation between raw, curated and processed data. Bucket-level IAM is the granularity Cloud Storage supports, so this directly satisfies the zone-restriction governance requirement.

Why this answer

Cloud Storage buckets are the fundamental access boundary for IAM policies. By creating three separate buckets (raw, curated, processed), you can assign distinct IAM roles (e.g., roles/storage.objectViewer, roles/storage.objectAdmin) per bucket, ensuring that users or service accounts only have access to the specific zone they are authorized for. This approach aligns with the principle of least privilege and avoids the complexity and limitations of IAM conditions or object-level ACLs.

Exam trap

A common misconception is that folders within a single Cloud Storage bucket can serve as effective security boundaries. However, folders are just a naming convention (prefixes) and do not provide native access control isolation without complex IAM conditions or ACLs.

How to eliminate wrong answers

Option A is wrong because IAM conditions on a single bucket with folders can restrict access based on object name prefixes, but they are complex to manage, prone to misconfiguration, and do not provide the same clear security boundary as separate buckets; also, IAM conditions are not supported for all roles and can lead to unintended access if not carefully crafted. Option B is wrong because Cloud Storage lifecycle rules are used for automating object transitions (e.g., moving to Nearline or deleting) based on age or other conditions, not for enforcing access control or governance between zones. Option C is wrong because relying on object ACLs is a legacy approach that is harder to audit and maintain at scale; ACLs provide per-object permissions but do not offer the centralized, hierarchical control of IAM roles, and they are not recommended for data lake architectures where consistent governance is required.

317
MCQmedium

A retail analytics team loads point-of-sale transactions into BigQuery every 15 minutes using a load job that appends records. A nightly deduplication job currently removes duplicate transaction IDs, but the team wants duplicates eliminated at ingestion time without adding a separate job. The source system can emit a stable transaction ID for each record. What should you do?

A.Change the load job to write to a staging table, then run a MERGE statement to upsert into the final table.
B.Add a unique constraint to the destination table's transaction ID column so duplicate inserts are rejected.
C.Partition the table by transaction date and cluster it by transaction ID, then rely on the storage engine to drop duplicates.
D.Use a MERGE statement in the load pipeline that matches on transaction ID, or write through the Storage Write API with the transaction ID as the primary key.
AnswerD

Both a MERGE keyed on transaction ID and the Storage Write API with a primary key collapse duplicates as part of the write itself. The Storage Write API in particular applies the primary key at commit time, so a re-delivered transaction ID is discarded without any separate cleanup job, satisfying the ingestion-time deduplication requirement.

Why this answer

Deduplication at write time requires the write path itself to enforce identity: either a MERGE keyed on transaction ID or the Storage Write API with a primary key. Physical layout choices such as partitioning and clustering improve query performance but never remove rows, and BigQuery does not enforce unique constraints. A staging-plus-MERGE design works but reintroduces the extra job the team is trying to eliminate.

Exam trap

The trap here is assuming that a unique constraint or a clustering key on transaction ID prevents duplicate rows from being stored, when BigQuery enforces neither for deduplication purposes.

318
Multi-Selecthard

A company uses Cloud Dataproc for large-scale Spark jobs. They notice that some jobs are failing due to insufficient memory on the worker nodes. They want to improve memory management without over-provisioning. Which three configurations should they apply? (Choose 3)

Select 3 answers
A.Set spark.executor.memory to a value that fits within the node memory
B.Enable Spark dynamic allocation
C.Use custom machine types with high memory ratios
D.Use local SSDs for temporary storage
E.Use preemptible worker nodes for volatile tasks
AnswersA, B, C

Prevents out-of-memory errors by ensuring executor memory fits worker capacity.

Why this answer

Setting spark.executor.memory to a value that fits within the node memory ensures that each executor does not exceed the available RAM on a worker node, preventing out-of-memory (OOM) errors. This configuration directly controls the heap size allocated to each executor, and when combined with spark.executor.cores and spark.executor.instances, it allows precise memory budgeting per node. Over-provisioning is avoided by calculating the maximum safe executor memory as (node memory - OS overhead - HDFS cache) / number of executors per node.

Exam trap

Google Cloud often tests the distinction between memory management and storage optimization, so candidates mistakenly choose local SSDs (option D) thinking they help with memory, when in fact they only improve disk I/O for shuffle operations.

319
MCQmedium

A data science team wants to deploy a model that requires a custom container with specific NVIDIA CUDA version. They build the image and push to Artifact Registry. When deploying to Vertex AI, the model fails to load with an error: 'Failed to start container: invalid ELF header'. What is the most likely cause?

A.The container image was built for a different CPU architecture (e.g., ARM64) than the Vertex AI machine (x86_64)
B.The model file (saved as .pkl) is corrupted
C.The CUDA version in the container is incompatible with the GPU on the machine
D.The container does not have the necessary permissions to access the model file in Cloud Storage
AnswerA

An ELF header mismatch means the binary's architecture does not match the host's. Building on ARM64 (for example, an Apple silicon laptop) produces ARM64 executables, which Vertex AI's x86_64 machines cannot load. Rebuilding the image with `--platform linux/amd64` resolves the failure.

Why this answer

The error 'invalid ELF header' indicates that the container image's executable format is not compatible with the host architecture. Vertex AI uses x86_64 machines, so if the image was built for ARM64 (e.g., on an Apple M1), it will fail to start. This is the most likely cause.

Other options would produce different errors (e.g., CUDA incompatibility, permission denied, corrupted model file).

Exam trap

The trap is focusing on CUDA or model file issues, but the ELF header error specifically points to architecture mismatch.

How to eliminate wrong answers

Option B is wrong because a corrupted model file would typically cause a different error during model loading, not an ELF header error. Option C is wrong because CUDA incompatibility would result in runtime errors related to GPU libraries, not an ELF header error. Option D is wrong because permission issues would cause access denied errors, not an ELF header error.

320
MCQhard

A logistics company uses Cloud Dataflow to process a continuous stream of GPS events from delivery trucks. The pipeline must compute the distance traveled per truck per hour and write the results to BigQuery. The events are keyed by truck ID, and the pipeline uses windowing with a one-hour fixed window. The team notices that some trucks report events with timestamps that are several minutes late due to network delays. They want to ensure that late events are still included in the correct window and that the results are emitted only after a reasonable wait. What should they configure?

A.Use a session window with a gap duration equal to the expected network delay.
B.Set a trigger that fires when the watermark passes the end of the window, and set allowed lateness to zero.
C.Set a trigger that fires when the watermark passes the end of the window, and set allowed lateness to a duration longer than the expected delay.
D.Set a trigger that fires every minute, and set allowed lateness to zero.
AnswerC

Configuring allowed lateness to a duration greater than the maximum expected delay ensures that late events are still accepted and processed into the correct window. The trigger firing on the watermark passing the window end provides timely results, and the allowed lateness extends the window's lifetime to accommodate late data. This balances completeness and latency.

Why this answer

To include late events in the correct window, you must set allowed lateness to a value greater than the expected delay. The trigger on the watermark passing the window end ensures results are emitted after the window closes, and the allowed lateness keeps the window state alive to accept late data. This configuration balances timely results with completeness.

Exam trap

The trap here is focusing only on triggers and forgetting that allowed lateness controls whether late data is accepted; without it, late events are discarded regardless of the trigger.

321
MCQhard

A retail company runs a batch pipeline in Cloud Dataflow that reads from Cloud Storage and writes to BigQuery. The pipeline uses a GroupByKey operation on a key that is heavily skewed: one customer ID accounts for 40% of all transactions. This causes a single worker to process a massive amount of data, and the job takes hours longer than expected. The team wants to reduce the skew without changing the pipeline's logic or output. What should they do?

A.Increase the number of workers and use a larger machine type for the Dataflow job.
B.Replace GroupByKey with CombinePerKey using an associative and commutative CombineFn.
C.Apply a hot key fanout by adding a random secondary key, perform the GroupByKey, then remove the secondary key and re-group.
D.Increase the disk size of the worker VMs to allow more data to be spilled to disk.
AnswerC

Hot key fanout splits a single hot key into multiple keys by appending a random suffix, so the GroupByKey distributes the load across many workers. After the first grouping, a second grouping removes the suffix and combines the partial results. This reduces skew without altering the pipeline's logic or output, and it is a standard Dataflow pattern for hot keys.

Why this answer

Hot key fanout is the correct technique for mitigating skew in a GroupByKey. By adding a random secondary key, the hot key's records are spread across multiple workers, and a subsequent grouping removes the secondary key to produce the final result. This preserves the pipeline's semantics while dramatically improving parallelism for the skewed key.

Exam trap

The trap here is thinking that adding more workers or larger machines can solve a hot key bottleneck, when the issue is that a single key must be processed by one worker regardless of cluster size.

322
MCQmedium

A data pipeline reading from Cloud Storage and writing to BigQuery using Dataflow is experiencing high cost. The data is CSV and needs schema inference. What change reduces cost?

A.Use Dataproc instead of Dataflow
B.Use Cloud Functions to transform data
C.Use BigQuery load jobs with schema auto-detection
D.Use BigQuery Data Transfer Service
AnswerC

Load jobs are free for data ingestion (only storage cost) and support auto-detection.

Why this answer

BigQuery load jobs with schema auto-detection can directly ingest CSV files from Cloud Storage without the need for a Dataflow pipeline, eliminating the compute cost associated with Dataflow. Schema auto-detection infers column names and types from the CSV header and data, matching the requirement for schema inference while being a serverless, no-cost-for-compute operation (you only pay for storage and querying). This reduces cost by removing the Dataflow processing step entirely.

Exam trap

Google Cloud often tests the misconception that any data transformation or schema inference requires a processing framework like Dataflow or Dataproc, when in fact BigQuery's native load jobs with auto-detection can handle many CSV ingestion scenarios at zero compute cost.

How to eliminate wrong answers

Option A is wrong because Dataproc is a managed Spark/Hadoop service that incurs compute costs for cluster VMs, and using it instead of Dataflow would not reduce cost—it would likely increase cost due to cluster overhead and the need to manage schema inference manually. Option B is wrong because Cloud Functions are event-driven compute that would still require processing each CSV file, incurring invocation and execution costs, and they lack native schema inference for BigQuery, requiring custom code that adds complexity and potential cost. Option D is wrong because BigQuery Data Transfer Service is designed for scheduled transfers from sources like Google Ads, Amazon S3, or SaaS applications, not for ad-hoc CSV files in Cloud Storage; it does not support schema auto-detection for arbitrary CSV files and would not replace the need for a pipeline.

323
MCQmedium

You need to split a time-series dataset into training and evaluation sets for a forecasting model. The data is ordered by timestamp. Which splitting technique should you use?

A.Sequential split where training data precedes evaluation data in time.
B.Use k-fold cross-validation with random folds.
C.Stratified split based on the target variable.
D.Random split with 80% training, 20% evaluation.
AnswerA

A sequential split keeps all training observations earlier in time than evaluation observations, preserving temporal order. This satisfies the forecasting constraint, since random splitting would leak future information into training and produce misleadingly optimistic evaluation metrics.

Why this answer

For time-series forecasting, the training set must only contain data that occurs before the evaluation set to prevent data leakage from the future into the past. A sequential split preserves the temporal order, ensuring the model is trained on historical data and evaluated on subsequent, unseen future data. This mimics real-world forecasting where you predict future values based on past observations.

Random or stratified splits violate this principle by allowing future data points to influence training, leading to overly optimistic performance estimates.

Exam trap

PDE often tests the misconception that standard cross-validation techniques (like k-fold or random split) are universally applicable, but for time-series data, they cause data leakage and invalidate the evaluation.

How to eliminate wrong answers

Option B is wrong because k-fold cross-validation with random folds shuffles the data, breaking temporal dependencies and causing data leakage; it is suitable for i.i.d. data, not time series. Option C is wrong because stratified splitting based on the target variable is designed for classification tasks with imbalanced classes, not for time-series forecasting where the target is continuous and order matters. Option D is wrong because a random split ignores the temporal ordering, allowing the model to train on future data and evaluate on past data, which is invalid for forecasting.

324
MCQmedium

Your team stores raw JSON event logs in Cloud Storage. Analysts need to run ad-hoc SQL queries on this data with minimal setup and cost, but they also require the ability to join it with existing BigQuery tables. Which approach should you recommend?

A.Create an external table in BigQuery that points to the Cloud Storage JSON files, and query it using standard SQL.
B.Load the JSON files into a new BigQuery table using the bq load command with --autodetect, then query the table directly.
C.Use Cloud Dataflow to read the JSON files, transform them, and write the results to BigQuery for querying.
D.Import the JSON files into Google Sheets and use the BigQuery data connector to join with BigQuery tables.
AnswerA

BigQuery external tables allow querying data directly from Cloud Storage without loading, providing minimal setup and no storage cost in BigQuery. They support standard SQL and can be joined with native BigQuery tables. This matches the requirement for ad-hoc queries with minimal cost and setup.

Why this answer

Querying external data in Cloud Storage via BigQuery external tables provides a serverless, cost-effective way to run ad-hoc SQL without loading data. It supports joins with native BigQuery tables, fulfilling the requirement. Other options involve data movement or unsuitable tools, adding cost and complexity.

Exam trap

The trap here is assuming that data must be loaded into BigQuery before it can be joined with existing tables.

325
MCQhard

A financial services firm ingests trade events into Cloud Storage and must load them into BigQuery. Compliance requires that each event be processed exactly once and that the load be idempotent across retries, even if a Dataflow job restarts mid-batch. The destination table must also be queryable immediately after each successful load. Which loading approach best satisfies these requirements?

A.Use BigQuery's LOAD DATA statement triggered by Cloud Storage object notifications, appending each new file into the destination table.
B.Use a Dataflow pipeline with the BigQueryIO Storage Write API in exactly-once mode, writing to the table with deterministic record identifiers.
C.Use a Dataproc Spark job that writes Parquet files to Cloud Storage, then run a scheduled bq load command every fifteen minutes to append into BigQuery.
D.Use a Dataflow pipeline that writes to BigQuery with STREAMING_INSERTS and relies on the pipeline's exactly-once semantics to deduplicate on retry.
AnswerB

The Storage Write API with exactly-once semantics assigns stream offsets and only commits data when checkpoints succeed, so retries do not duplicate rows. Combined with deterministic record identifiers and Dataflow checkpointing, a restarted job resumes without re-inserting committed records. Data written through committed streams becomes immediately queryable, satisfying the requirement that the table be readable right after each successful load.

Why this answer

Exactly-once loading into BigQuery is achieved with the Storage Write API's exactly-once streams, where offsets are committed only on checkpoint and deterministic record IDs prevent duplicates on restart. This gives idempotent loads across retries and makes committed rows immediately queryable. Streaming inserts and periodic batch appends both risk duplicates and cannot meet the compliance constraint.

Exam trap

The trap here is believing that Dataflow's exactly-once processing guarantee automatically extends to any BigQuery sink, when the sink itself must support offset-based commits to avoid duplicates.

326
Multi-Selecteasy

A company is designing a CI/CD pipeline for their ML models using Cloud Build and Vertex AI. Which TWO practices should they adopt to ensure reliable and reproducible deployments?

Select 2 answers
A.Require manual approval for every model change before deployment
B.Store all model artifacts in a single Cloud Storage bucket without versioning
C.Use immutable container images with version tags for each model deployment
D.Include unit tests for data preprocessing and feature engineering code in the pipeline
E.Deploy every model version directly to production for immediate use
AnswersC, D

Immutable, version-tagged container images pin the exact code and dependencies used at training and serving time. This eliminates drift between pipeline runs, giving the reproducible deployments the stem requires, since a tag always resolves to identical, unmodified image content.

Why this answer

Option C is correct because using immutable container images with version tags ensures that each model deployment is tied to a specific, unchangeable artifact, which is essential for reproducibility and rollback in a CI/CD pipeline. Option D is correct because including unit tests for data preprocessing and feature engineering code in the pipeline catches errors early and verifies that the transformation logic behaves consistently, which directly supports reliable and reproducible ML deployments. Option A is not required for reproducibility and would slow the pipeline; manual approval is a governance choice, not a technical practice for reliable, reproducible deployment.

Option B is wrong because storing artifacts in a single bucket without versioning prevents tracking and reproducing specific model versions. Option E is wrong because deploying every model version directly to production bypasses validation and testing, undermining reliability and reproducibility.

Exam trap

Google often tests the misconception that manual approval gates or single-bucket storage without versioning are acceptable for reproducibility, when in fact they undermine automation and traceability in CI/CD pipelines.

327
MCQeasy

Which BigQuery feature allows you to estimate the cost of a query before running it, by returning the number of bytes that would be processed?

A.EXPLAIN statement
B.INFORMATION_SCHEMA.JOBS
C.--dry_run flag
D.Slot estimator
AnswerC

The `--dry_run` flag validates a query and returns the bytes it would process without executing it, so no compute charges are incurred. This directly satisfies the stem's requirement to estimate cost before running, since BigQuery on-demand pricing is calculated from bytes processed.

Why this answer

The --dry_run flag in BigQuery allows you to validate a query and estimate the number of bytes it will process without actually running it, thus providing a cost estimate. This is a standard feature in the bq command-line tool and the BigQuery API. It returns the total bytes processed, which can be used to calculate the cost based on current pricing.

Exam trap

The trap is confusing the dry run flag with other features like EXPLAIN or INFORMATION_SCHEMA, which serve different purposes.

How to eliminate wrong answers

Option A is wrong because EXPLAIN is not a BigQuery feature; it is used in other databases to show query plans. Option B is wrong because INFORMATION_SCHEMA.JOBS provides metadata about completed jobs, not a pre-execution cost estimate. Option D is wrong because the slot estimator is not a feature for estimating query cost; slots are compute resources, and there is no direct 'slot estimator' for cost estimation.

328
MCQmedium

A team runs a Dataflow streaming pipeline that reads from Pub/Sub, windows events by processing time, and writes to BigQuery. Some late-arriving events are being dropped. The requirement is to include all events that arrive within 10 minutes of the watermark. Which pipeline configuration should be used?

A.Use sliding windows with no allowed lateness
B.Use fixed windows with .withAllowedLateness(Duration.standardMinutes(10))
C.Use fixed windows with withAllowedLateness(Duration.standardSeconds(10))
D.Switch from processing time to event time and use default triggers
AnswerB

Allows late data up to 10 minutes after watermark.

Why this answer

`withAllowedLateness(Duration.standardMinutes(10))` on a fixed window allows late-arriving events to be included up to 10 minutes after the watermark passes the window's end. This directly meets the requirement to retain events arriving within 10 minutes of the watermark, while still using processing-time windows as specified.

Exam trap

Google Cloud often tests the distinction between processing time and event time, and the exact value of allowed lateness, tricking candidates into choosing a shorter duration or the wrong window type.

How to eliminate wrong answers

Option A is wrong because sliding windows with no allowed lateness will drop all late events, failing the requirement to include events within 10 minutes of the watermark. Option C is wrong because `withAllowedLateness(Duration.standardSeconds(10))` only allows 10 seconds of lateness, not the required 10 minutes. Option D is wrong because switching to event time would change the windowing basis from processing time, which is not requested, and default triggers alone do not provide the explicit 10-minute lateness allowance needed.

329
Multi-Selectmedium

A healthcare company stores patient records as JSON files in Cloud Storage for analysis. They want to design a data lake that enables querying the data with BigQuery while minimizing storage costs and maintaining data security. Which two actions should they take? (Choose two.)

Select 2 answers
A.Partition the data by date and store in separate directories for each partition.
B.Configure object lifecycle management to transition files older than 90 days to Nearline storage.
C.Convert all JSON files to CSV to reduce storage size.
D.Use BigLake to create external tables with row-level security and access delegation.
E.Enable Cloud KMS to encrypt the data with customer-managed encryption keys.
AnswersB, D

Object lifecycle management automatically transitions JSON objects to Nearline after 90 days, cutting storage cost while keeping data queryable. This satisfies the cost-minimisation constraint without deleting records, and Nearline's 30-day minimum retention suits the healthcare retention profile.

Why this answer

Option B is correct because configuring Object Lifecycle Management to transition objects older than 90 days to Nearline storage lowers storage costs for infrequently accessed historical patient records while keeping them queryable, since Nearline has a 30-day minimum storage duration and is cheaper than Standard. Option D is correct because BigLake external tables let BigQuery query JSON files directly in Cloud Storage without ingestion, and BigLake supports fine-grained security such as row-level security and access delegation (via BigQuery delegated access to the underlying Cloud Storage objects), satisfying the data security requirement. Option A is not required: partitioning by date in separate directories is a performance/cost optimization for query pruning, but it does not by itself minimize storage cost or provide security, and BigLake can query the data without this layout.

Option C is not appropriate: converting JSON to CSV does not reliably reduce storage size and loses the nested schema, and it is not needed for BigQuery querying. Option E is not the best fit: Cloud KMS with CMEK encrypts data at rest but does not enable BigQuery querying or reduce storage costs, and Cloud Storage is already encrypted by default with Google-managed keys.

Exam trap

A common misconception is that converting JSON to CSV always reduces storage size, but the primary cost-saving mechanism for infrequently accessed data is lifecycle management to colder storage tiers like Nearline. Additionally, BigLake provides security and access delegation without needing to transform the data format.

330
MCQmedium

A company needs to store petabytes of time-series IoT sensor data and query it with single-digit millisecond latency at millions of reads per second. The data has a simple key-value structure with timestamps. Which Google Cloud database is MOST appropriate?

A.Firestore
B.Cloud Spanner
C.Cloud Bigtable
D.BigQuery
AnswerC

Cloud Bigtable is a wide-column NoSQL store engineered for petabyte scale and consistent single-digit-millisecond reads at millions of operations per second on row-key lookups, matching the key-value timestamped workload. Spanner and BigQuery cannot meet that latency.

Why this answer

Cloud Bigtable is a fully managed, scalable NoSQL database designed for large analytical and operational workloads, handling petabytes of data with consistent single-digit millisecond latency for high-throughput reads and writes. Its key-value model with timestamp-based versioning is ideal for time-series IoT sensor data, and it supports millions of reads per second via its HBase API and Bigtable's underlying tablet-based architecture.

Exam trap

A common trap in Google exams is to choose Cloud Spanner for structured data, but Spanner is optimized for strong consistency and transactional workloads, not for high-throughput time-series data at petabyte scale with millions of reads per second. BigQuery is for analytical queries, not real-time key-value access. Bigtable's key-value model and low-latency high-throughput design make it the correct choice.

How to eliminate wrong answers

Option A is wrong because Firestore is a document-oriented NoSQL database optimized for mobile and web app real-time sync, not for petabyte-scale time-series workloads with millions of reads per second; it has throughput limits (e.g., 10,000 writes/second per database) and does not natively handle high-throughput time-series data. Option B is wrong because Cloud Spanner is a globally distributed relational database with strong consistency and SQL support, but it is designed for transactional workloads (OLTP) with moderate throughput, not for petabyte-scale key-value time-series data at millions of reads per second; its latency and cost profile are not optimal for this use case. Option D is wrong because BigQuery is a serverless data warehouse for analytical SQL queries on large datasets, but it is not designed for single-digit millisecond latency at millions of reads per second; it is optimized for batch and interactive analytics, not real-time key-value lookups.

331
MCQeasy

A data engineer has a BigQuery SQL script that must run every day at 06:00, load its results into a reporting table, and retry automatically if the query fails due to transient errors. The team has no existing orchestration tooling, wants the lowest operational overhead, and needs the schedule and the SQL to be managed entirely inside Google Cloud. Which approach should the engineer use?

A.Use Dataflow with a batch pipeline that reads from the source tables, applies the transformation, and writes to the reporting table on a schedule.
B.Create a Cloud Composer environment and author a single-task DAG that runs the SQL daily with Airflow retries configured.
C.Create a scheduled query in BigQuery with the desired SQL and a daily schedule, targeting the reporting table as the destination.
D.Deploy the SQL as a Cloud Run service and create a Cloud Scheduler HTTP job that invokes it daily at 06:00.
AnswerC

BigQuery scheduled queries let you store SQL, define a schedule, and write results to a destination table directly in the BigQuery UI or API. The service manages execution and retries failed runs automatically, requiring no containers, schedulers, or external orchestration, which is exactly the lowest-overhead option described.

Why this answer

BigQuery scheduled queries are the native, serverless way to run SQL on a recurring schedule and write output to a destination table, with automatic retry of failed runs. Because the schedule, SQL, and destination all live inside BigQuery, no additional infrastructure or orchestration tooling is needed, matching the low-overhead requirement precisely.

Exam trap

The trap here is reaching for external schedulers or orchestration platforms when BigQuery's own scheduled query feature already covers recurring SQL with retries and destination tables.

332
MCQeasy

A startup uploads JSON log files to a Cloud Storage bucket throughout the day and wants BigQuery to query them within minutes with no ETL code and no data duplication. The files use newline-delimited JSON and the schema is stable. The team wants the lowest-effort option that still lets queries see newly arrived files automatically. Which approach should they use?

A.Schedule a nightly bq load job to append the JSON files into a native BigQuery table
B.Build a Dataflow streaming job that parses the JSON and writes rows to BigQuery
C.Use the Storage Write API to stream each uploaded file's contents directly into a BigQuery table
D.Create a BigQuery external table over the Cloud Storage bucket using the newline-delimited JSON format
AnswerD

An external table over Cloud Storage lets BigQuery query the JSON files in place without loading or duplicating data, and newly added files that match the URI pattern are visible to subsequent queries automatically. No ETL code or pipeline is required, so this is the lowest-effort option that still reflects new arrivals within minutes.

Why this answer

A BigQuery external table reads newline-delimited JSON directly from Cloud Storage, so no data is copied and no pipeline code is needed. Because the table definition is a URI pattern over the bucket, files added later are picked up by subsequent queries, which satisfies the fast-visibility and low-effort requirements.

Exam trap

The trap here is assuming data must be loaded into BigQuery storage before it can be queried, when external tables query Cloud Storage in place.

333
MCQmedium

Your team runs a weekly batch ETL pipeline using Cloud Dataproc. The pipeline reads raw data from Cloud Storage, transforms it with Apache Spark, and writes results to BigQuery. Recently, the pipeline has been failing with the error 'Out of Memory' during the shuffle phase. The cluster uses standard worker nodes (n1-standard-4). What is the most effective way to resolve this without increasing total cost?

A.Increase the number of Spark partitions by setting spark.sql.shuffle.partitions to a higher value.
B.Increase the number of worker nodes by adding more n1-standard-4 instances.
C.Enable dynamic allocation and use preemptible VMs for some workers.
D.Switch worker nodes to n1-highmem-4 instances to provide more memory.
AnswerA

Raising `spark.sql.shuffle.partitions` splits shuffle data into more, smaller partitions, reducing per-task memory pressure during the shuffle phase. This directly addresses the Out of Memory failures on the existing n1-standard-4 workers without adding nodes, so total cost stays unchanged.

Why this answer

The 'Out of Memory' error during the shuffle phase indicates that individual executor tasks are processing too much data per partition. Increasing `spark.sql.shuffle.partitions` reduces the amount of data each task handles, lowering memory pressure per executor without adding more nodes or upgrading hardware. This directly addresses the shuffle memory bottleneck while keeping the total cluster cost unchanged.

Exam trap

The trap here is that candidates often assume memory errors must be solved by adding more memory (Option D) or more nodes (Option B), ignoring the cost constraint and the fact that repartitioning can resolve the issue without additional resources.

How to eliminate wrong answers

Option B is wrong because adding more worker nodes increases total cost, which violates the constraint of not increasing cost. Option C is wrong because enabling dynamic allocation and using preemptible VMs does not resolve the per-executor memory shortage; it only changes cluster scaling and cost structure, but the shuffle memory issue persists on the existing nodes. Option D is wrong because switching to n1-highmem-4 instances increases per-node memory but also increases cost per node, raising total cost unless the number of nodes is reduced, which is not specified and may not be feasible without losing parallelism.

334
MCQeasy

A team deploys a new version of a Cloud Function. After deployment, error rates increase significantly. What is the most efficient way to diagnose the cause?

A.Deploy a debug version with additional logging.
B.Check Cloud Logging for error stacks and exceptions.
C.Increase the function timeout and retry settings.
D.Immediately rollback to the previous version.
AnswerB

Cloud Logging captures the stack traces and exception details emitted by the failing function, revealing the exact error and code path. Reviewing these logs is faster than redeploying or guessing, directly identifying the cause of the increased error rate.

Why this answer

Cloud Logging automatically captures error stacks and exceptions from Cloud Functions without requiring code changes. Checking these logs is the most efficient first step because it provides immediate visibility into the root cause of errors, such as unhandled exceptions, timeouts, or dependency failures, without incurring additional deployment overhead.

Exam trap

Google tests the principle of 'most efficient diagnostic step' by tempting candidates to choose a reactive action (like rollback or timeout increase) or a time-consuming code change, rather than leveraging existing observability tools like Cloud Logging that provide immediate, detailed error context.

How to eliminate wrong answers

Option A is wrong because deploying a debug version with additional logging is inefficient and time-consuming; it requires modifying code, redeploying, and potentially introducing new issues, whereas Cloud Logging already captures detailed error information. Option C is wrong because increasing function timeout and retry settings does not diagnose the cause of errors; it only masks symptoms by allowing more time for execution or retrying failed invocations, which could exacerbate resource consumption and latency. Option D is wrong because immediately rolling back to the previous version is a reactive mitigation step, not a diagnostic one; it may restore service but fails to identify the root cause, leaving the team without insight into what went wrong in the new version.

335
MCQmedium

A company processes real-time clickstream data from websites. They need to aggregate user sessions that may span multiple hours and handle events that arrive late due to network delays. The pipeline must avoid discarding late data. Which Dataflow feature should they configure?

A.Use fixed windows with a trigger that fires after every element
B.Use session windows with a gap duration and allow late data with a suitable allowed_lateness
C.Use the GlobalWindow with a watermark
D.Use sliding windows with no allowed lateness
AnswerB

Session windows group events separated by a gap duration, so multi-hour user sessions merge correctly. Setting allowed_lateness retains events arriving after the watermark, preventing discarding of network-delayed clickstream data, which directly satisfies the no-late-data-loss constraint.

Why this answer

Session windows are ideal for aggregating user sessions that span multiple hours, as they group events based on a gap duration of inactivity. By configuring `allowed_lateness`, the pipeline can handle late-arriving events without discarding them, ensuring completeness. This directly addresses the requirement to avoid discarding late data while aggregating sessions.

Exam trap

Google Cloud often tests the distinction between window types and late-data handling; the trap here is that candidates might choose fixed or sliding windows without realizing they lack the session-gap logic needed for variable-length user sessions, or they might overlook the `allowed_lateness` parameter as the key to preserving late data.

How to eliminate wrong answers

Option A is wrong because fixed windows with a trigger after every element would create a new window per event, failing to aggregate sessions that span hours and not handling late data properly. Option C is wrong because GlobalWindow with a watermark is used for global aggregations (e.g., counting all events) but does not naturally group events into sessions based on inactivity gaps; it would require complex triggers and does not inherently support sessionization. Option D is wrong because sliding windows with no allowed lateness would discard any late-arriving events, violating the requirement to avoid discarding late data.

336
MCQhard

A retail company ingests point-of-sale clickstream events into Cloud Pub/Sub at roughly 200,000 messages per second during flash sales. Analysts need near-real-time dashboards that aggregate revenue by product category over sliding 5-minute windows, with results visible in BigQuery within 30 seconds of the event. The pipeline must handle occasional bursts up to 3x the normal rate without dropping messages, and the team wants to minimize operational overhead. Which design should the data engineer use?

A.Use Cloud Pub/Sub push subscriptions delivering events directly to a Cloud Run service that aggregates in memory and calls the BigQuery streaming insert API once per minute.
B.Deploy a Dataproc cluster with Spark Structured Streaming reading from Pub/Sub and writing micro-batches to BigQuery every 10 seconds, sizing the cluster for the 3x burst rate.
C.Create a Cloud Dataflow streaming pipeline in the Apache Beam Java SDK using Pub/Sub as the source, apply windowing with a 5-minute sliding window and a 30-second allowed lateness, and write aggregated results to BigQuery using the Storage Write API.
D.Load raw events into Cloud Storage using a Pub/Sub subscription with a Cloud Storage sink, then run a BigQuery scheduled query every 5 minutes that reads the new files and computes the sliding window aggregates.
AnswerC

Dataflow provides exactly-once streaming semantics, autoscaling that absorbs the 3x burst, and native Pub/Sub and BigQuery connectors. Sliding windows with allowed lateness handle out-of-order events, and the Storage Write API gives low-latency, high-throughput BigQuery inserts, satisfying the 30-second visibility requirement without managing clusters.

Why this answer

The requirement for sliding 5-minute windows with sub-30-second BigQuery visibility under bursty load points to a managed stream processor with native windowing and autoscaling. Dataflow's sliding windows, allowed lateness, and BigQuery Storage Write API together meet latency, correctness, and operational-overhead constraints. Cluster-based and file-based approaches add latency or management burden that the scenario rules out.

Exam trap

The trap here is assuming that any Pub/Sub consumer with a BigQuery sink is equivalent, when only a stateful stream processor with true windowing semantics can meet the latency and correctness requirements.

337
Multi-Selecthard

A company uses Cloud Dataproc for ephemeral clusters to run batch jobs. They want to ensure job reliability and data quality. Which two configuration options should they use? (Choose two.)

Select 2 answers
A.Enable preemptible VMs for cost savings.
B.Use initialization actions for cluster setup.
C.Enable idle timeout to automatically delete clusters.
D.Use custom machine types for better performance.
E.Use graceful decommissioning of workers.
AnswersB, E

Initialization actions run scripts on every node during cluster creation, installing agents, libraries or configuration needed for consistent job execution. This satisfies the reliability requirement by ensuring each ephemeral cluster starts identically, preventing job failures caused by missing dependencies or inconsistent setup.

Why this answer

Option B is correct because initialization actions let you run scripts on every node at cluster creation, so you can install monitoring/validation agents, configure logging, and enforce consistent setup that supports job reliability and data quality across ephemeral clusters. Option E is correct because graceful decommissioning (set via yarn:yarn.resourcemanager.decommissioning.timeout or Dataproc's gracefulDecommissionTimeout) lets running tasks finish before workers are removed, preventing partial job failures and data corruption when nodes are scaled down or preempted. Option A is not appropriate here because preemptible VMs reduce cost but can be reclaimed at any time, which undermines reliability rather than improving it.

Option C is not appropriate because idle timeout only deletes idle clusters to save cost; it does not improve job reliability or data quality. Option D is not appropriate because custom machine types tune CPU/memory performance but do not by themselves ensure reliable job completion or data quality.

Exam trap

The trap here is that candidates might confuse cost-saving or performance features with reliability and data quality mechanisms. For example, enabling preemptible VMs (which are spot instances in Google Cloud) reduces cost but can cause job failures if workers are reclaimed. Idle timeout only deletes clusters after inactivity, not ensuring reliable job execution.

Custom machine types improve performance but not reliability. In contrast, initialization actions for Google Cloud Dataproc ensure every ephemeral cluster node has the correct software and data sources, directly supporting job reliability and data quality. Graceful decommissioning allows workers to complete their tasks before being removed, preventing data loss during scaling down or cluster deletion.

These two options directly address consistency and fault tolerance for Dataproc batch jobs.

338
MCQmedium

A data engineer is building a batch pipeline that runs daily using Cloud Composer. The pipeline has three tasks: extract data from Cloud Storage, transform data using Dataflow, and load the transformed data into BigQuery. The engineer wants to ensure that the Dataflow job only starts after the extraction task completes successfully, and the load task only starts after the Dataflow job finishes. How should the engineer define the task dependencies in the Airflow DAG?

A.extract >> [transform, load]
B.transform >> extract >> load
C.extract >> transform >> load
D.extract >> load >> transform
AnswerC

Airflow's bitshift operator sets explicit downstream dependencies, so transform runs only after extract succeeds and load only after transform completes. This directly enforces the sequential ordering the stem requires, preventing the Dataflow job from starting prematurely.

Why this answer

Airflow uses the bitshift operator (>>) to define task dependencies in a linear sequence. The DAG must ensure that the extract task completes before the transform task starts, and the transform task completes before the load task starts. This is achieved by chaining the tasks in order: extract >> transform >> load, which enforces the required sequential execution.

Exam trap

A common misconception is that multiple tasks can be chained in parallel with a single bitshift operator, leading candidates to choose Option A, which incorrectly allows the load task to start before the Dataflow job completes. In Airflow, sequential dependencies are defined by chaining tasks with '>>' in order.

How to eliminate wrong answers

Option A is wrong because it sets transform and load as parallel downstream tasks of extract, meaning load could start before transform finishes, violating the requirement that load waits for Dataflow. Option B is wrong because it places transform before extract, which would attempt to run the Dataflow job before the extraction completes, breaking the dependency chain. Option D is wrong because it places load before transform, which would attempt to load data into BigQuery before the Dataflow transformation is done, leading to incorrect or missing data.

339
MCQhard

You are designing a BigQuery schema for a table that will store user profile data. The data includes a unique user ID, a list of email addresses (each with a type and address), and a timestamp of last update. You need to support efficient queries that retrieve all email addresses for a given user and also filter users by email type. Which schema design should you use?

A.Store emails as a single STRING column with a delimiter, and use SPLIT and REGEXP_CONTAINS to query.
B.Store emails as a repeated RECORD with fields 'type' and 'address' within the user table, using dot notation and UNNEST to query.
C.Store emails as a JSON STRING column and use JSON functions to extract values.
D.Create a separate table for emails with a foreign key to the user table, and join when needed.
AnswerB

A repeated RECORD allows storing multiple emails per user while preserving the nested structure. Queries can use UNNEST to flatten the array and filter by email type, and dot notation to access fields. This denormalized approach is efficient in BigQuery, avoiding joins and enabling fast retrieval of all emails for a user.

Why this answer

Using a repeated RECORD for emails allows BigQuery to store multiple emails per user with type and address fields. Queries can UNNEST the array to filter by email type and use dot notation to access fields. This denormalized design is efficient, avoids joins, and aligns with BigQuery's strengths for nested data.

Exam trap

The trap here is normalizing into separate tables, which is common in relational databases but can degrade performance in BigQuery due to joins and increased data shuffling.

340
MCQeasy

A startup needs a fully managed, serverless Spark service to run occasional data processing jobs without managing clusters. They want to pay only for the resources used during job execution. Which Google Cloud service should they use?

A.Dataproc Serverless
B.Dataflow
C.Cloud Data Fusion
D.Dataproc
AnswerA

Dataproc Serverless runs Spark workloads without provisioning or managing a cluster, allocating resources only while the job executes and charging for that consumption. This matches the stem's requirements for occasional jobs, no cluster management and pay-per-use billing, unlike a standard Dataproc cluster.

Why this answer

Dataproc Serverless provides a serverless Spark environment where you pay per job execution. Cloud Data Fusion is for visual ETL. Dataproc is managed but not serverless.

Dataflow is serverless for Beam, not Spark.

341
MCQhard

A Dataflow pipeline as described in the exhibit has increasing lag. Which optimization is most likely to reduce the lag?

A.Use FileLoads instead of StreamingInserts for BigQuery output
B.Increase the number of workers
C.Use global windows instead of fixed windows
D.Add additional ParDo transforms
AnswerA

FileLoads (batch loads) are more efficient for high throughput and reduce lag.

Why this answer

The exhibit shows increasing lag in a Dataflow pipeline writing to BigQuery. StreamingInserts (the default) use the BigQuery Storage Write API, which can throttle under high throughput, causing backpressure and lag. Switching to FileLoads writes data to temporary files in Cloud Storage and then loads them into BigQuery via batch load jobs, which decouples the write path from the streaming insert quota and reduces lag by avoiding per-row insert limits.

Exam trap

Google Cloud often tests the misconception that scaling workers or changing windowing fixes all performance issues, but the trap here is that the lag is specifically caused by the BigQuery sink's streaming insert throttling, which requires a sink-level optimization like FileLoads.

How to eliminate wrong answers

Option B is wrong because increasing the number of workers can help with parallel processing but does not address the root cause of lag from BigQuery streaming insert quota exhaustion or throttling; it may even increase the rate of inserts and worsen the problem. Option C is wrong because using global windows instead of fixed windows does not affect the write path to BigQuery; windowing changes how data is grouped for aggregation but does not reduce lag caused by the sink's throughput limitations. Option D is wrong because adding additional ParDo transforms increases the processing steps and can introduce more latency, making the lag worse rather than reducing it.

342
MCQhard

You have two versions of a classification model (v1 and v2) deployed on a Vertex AI Endpoint. You want to gradually roll out v2 to 10% of traffic, monitor performance, and if metrics are better, increase traffic to 100%. You have set up model monitoring for skew and drift. Which configuration should you use?

A.Use the Vertex AI Endpoint 'traffic_split' parameter to assign 10% of traffic to v2 and 90% to v1.
B.Deploy v2 to a separate endpoint and use a load balancer to route 10% of traffic.
C.Create a new deployment with v2 on the same endpoint and set the 'min_replica_count' to 1 for both versions.
D.Enable Vertex AI Model Monitoring on the endpoint and set up alerting for performance drop.
AnswerA

The traffic_split parameter on a Vertex AI Endpoint distributes prediction requests across deployed model versions by percentage. Setting 10% to v2 and 90% to v1 enables gradual canary rollout, letting monitoring metrics inform whether to increase v2 traffic to 100%.

Why this answer

The Vertex AI Endpoint 'traffic_split' parameter allows you to direct a percentage of inference requests to different model versions deployed on the same endpoint. Setting 10% to v2 and 90% to v1 enables a gradual rollout while monitoring skew and drift, and you can adjust the split as needed. This is the native, supported method for canary deployments in Vertex AI, avoiding the complexity and latency of external load balancers.

Exam trap

The trap here is that candidates confuse infrastructure-level load balancing (Option B) with Vertex AI's built-in traffic splitting, or think that replica counts (Option C) control traffic distribution, when in fact traffic_split is the only parameter that directly controls request routing percentages.

How to eliminate wrong answers

Option B is wrong because deploying v2 to a separate endpoint and using an external load balancer adds unnecessary complexity, latency, and cost; Vertex AI Endpoints natively support traffic splitting without additional infrastructure. Option C is wrong because setting 'min_replica_count' to 1 for both versions does not control traffic distribution; it only ensures minimum instance availability, not the percentage of requests routed to each model. Option D is wrong because enabling Model Monitoring and alerting for performance drop is a monitoring step, not a configuration for traffic splitting; it does not direct 10% of traffic to v2.

343
MCQmedium

You are building a data pipeline that ingests JSON files from Cloud Storage into BigQuery. The JSON files contain deeply nested arrays and objects. You need to load these files with minimal transformation so that analysts can query individual nested fields using dot notation and also unnest arrays when needed. Which approach should you use?

A.First load the JSON files into a Cloud SQL PostgreSQL instance, then use federated queries to access the data from BigQuery.
B.Load the JSON files into BigQuery with a schema that defines nested fields as RECORD and arrays as REPEATED, allowing queries to reference nested fields with dot notation and use UNNEST for arrays.
C.Convert the JSON files to CSV using a Dataflow pipeline that flattens all arrays, then load the CSV into BigQuery.
D.Load the JSON files into BigQuery using the autodetect schema option, which will automatically flatten all nested structures into separate columns.
AnswerB

BigQuery natively supports nested and repeated fields via RECORD and REPEATED types. When JSON is loaded with such a schema, nested objects become RECORDs and arrays become REPEATED fields. Analysts can then use dot notation to access nested fields and UNNEST to flatten arrays in queries, meeting the requirement without transformation.

Why this answer

BigQuery supports semi-structured data through nested and repeated fields. Defining nested objects as RECORD and arrays as REPEATED preserves the original structure, enabling dot notation for nested fields and UNNEST for arrays. This avoids flattening during load, aligning with minimal transformation and providing flexible query capabilities.

Exam trap

The trap here is assuming that BigQuery automatically flattens nested JSON during load, when in fact it preserves nested structures as RECORD and REPEATED types if the schema is defined accordingly.

344
MCQmedium

A data scientist wants to train a custom TensorFlow model on Vertex AI using a managed Jupyter notebook. Which Vertex AI service should they use to set up a notebook environment with pre-installed deep learning frameworks?

A.Compute Engine with Deep Learning VM
B.Vertex AI Training via custom job
C.Vertex AI Workbench
D.Vertex AI Pipelines
AnswerC

Vertex AI Workbench provides managed JupyterLab notebook instances with pre-installed deep learning frameworks such as TensorFlow and PyTorch, plus optional GPU accelerators. This satisfies the requirement for a managed notebook environment ready for custom TensorFlow training without manual framework installation.

Why this answer

Vertex AI Workbench provides managed Jupyter notebooks with pre-installed deep learning frameworks (TensorFlow, PyTorch, etc.) and easy scaling options. Notebooks on Compute Engine would require manual setup. AI Platform Training is for training jobs, not interactive notebooks.

Vertex AI Pipelines is for orchestrating ML workflows.

345
MCQeasy

Which BigQuery SQL function can be used to get an approximate count of distinct values in a large column faster than COUNT(DISTINCT) with lower accuracy?

A.APPROX_QUANTILES
B.COUNT(DISTINCT)
C.APPROX_COUNT_DISTINCT
D.DISTINCT_COUNT
AnswerC

APPROX_COUNT_DISTINCT uses HyperLogLog++ sketches to estimate cardinality, scanning far less data than COUNT(DISTINCT)'s exact deduplication. This satisfies the stem's demand for faster approximate distinct counts on large columns, trading a small, bounded error rate for substantially reduced query cost and latency.

Why this answer

APPROX_COUNT_DISTINCT is a BigQuery function that returns an approximate count of distinct values using a HyperLogLog++ algorithm. It is significantly faster and uses far fewer resources than COUNT(DISTINCT) on large datasets, at the cost of a small statistical error (typically under 1%). This matches the requirement for a faster, lower-accuracy distinct count.

Exam trap

The trap is a fabricated-sounding option (DISTINCT_COUNT) that mimics the correct function name, plus the presence of the exact COUNT(DISTINCT) the question says to avoid — candidates must recognize the real BigQuery function name.

How to eliminate wrong answers

Option A is wrong because APPROX_QUANTILES returns approximate quantile boundaries (e.g., median, percentiles) for a column, not a count of distinct values. Option B is wrong because COUNT(DISTINCT) is the exact, resource-intensive method the question explicitly wants to avoid due to speed and cost. Option D is wrong because DISTINCT_COUNT is not a valid BigQuery function — it is a fabricated name that resembles the correct answer but does not exist in the BigQuery SQL reference.

346
MCQeasy

A company wants to trigger a Cloud Run service whenever a new file is uploaded to a specific Cloud Storage bucket. Which event-driven solution should they use?

A.Eventarc with Cloud Storage trigger and Cloud Run destination
B.Cloud Scheduler to periodically poll the bucket
C.Cloud Functions triggered by Cloud Storage
D.Pub/Sub with a push subscription to Cloud Run
AnswerA

Eventarc receives Cloud Storage object-finalise events and routes them to Cloud Run, providing native event-driven invocation without polling. This satisfies the requirement to trigger the service whenever a new file lands in the specified bucket.

Why this answer

Eventarc is the recommended service for routing events from Cloud Storage to Cloud Run because it provides a fully managed, event-driven architecture with built-in filtering and retry logic. When a new file is uploaded, Cloud Storage emits a notification that Eventarc captures and delivers directly to the Cloud Run service as an HTTP request, enabling serverless processing without polling or additional infrastructure.

Exam trap

The trap here is that candidates confuse Cloud Functions (option C) as the only serverless compute option for Cloud Storage events, overlooking that Eventarc is the modern, preferred service for routing events to Cloud Run, and that Pub/Sub (option D) requires manual setup not shown in the question.

How to eliminate wrong answers

Option B is wrong because Cloud Scheduler is a cron job service for scheduled, not event-driven, tasks; periodically polling a bucket introduces latency and inefficiency, and it cannot react instantly to uploads. Option C is wrong because Cloud Functions triggered by Cloud Storage is a valid event-driven approach, but the question specifically asks for a Cloud Run destination, and Cloud Functions cannot directly invoke Cloud Run without additional integration. Option D is wrong because Pub/Sub with a push subscription to Cloud Run requires manually configuring Cloud Storage to publish to Pub/Sub, which adds complexity and is not the native, recommended pattern for Cloud Storage events; Eventarc abstracts this by directly managing the event flow from Cloud Storage to Cloud Run.

347
Multi-Selecthard

A financial services firm is designing a data processing system on Google Cloud that must ingest change data capture (CDC) streams from an on-premises PostgreSQL database into BigQuery with sub-minute latency, preserve the ordering of changes per primary key, and apply updates and deletes so that BigQuery reflects the current state of each row. The source database cannot be modified to add triggers. Which two design elements should you include? (Choose two.)

Select 2 answers
A.Add database triggers to the PostgreSQL source that publish change events to a Pub/Sub topic.
B.Use Datastream to read the PostgreSQL write-ahead log and stream changes into Cloud Storage or BigQuery.
C.Use a federated BigQuery external table that queries the on-premises PostgreSQL instance directly at read time.
D.Enable BigQuery streaming inserts with insertId deduplication to apply deletes directly to the target table.
E.Configure the BigQuery destination to use the staging and merge pattern with a primary key so deletes and updates are applied.
AnswersB, E

Datastream performs log-based CDC by reading the PostgreSQL write-ahead log, so no triggers or schema changes are needed on the source, satisfying the constraint that the database cannot be modified. It delivers low-latency change records that can land in Cloud Storage or be written directly into BigQuery, providing the ingestion mechanism the design requires.

Why this answer

Log-based CDC through Datastream reads the PostgreSQL write-ahead log without touching the source schema, and the staging-plus-merge destination pattern keyed on the primary key is what turns a change stream into correct current-state rows in BigQuery, including deletes. Together they meet the latency, ordering, and mutation requirements.

Exam trap

The trap here is assuming that streaming inserts into BigQuery can apply updates and deletes, when the streaming API is append-only and mutations require a merge pattern.

348
MCQeasy

A company deploys a machine learning model on Vertex AI for online predictions. The model experiences intermittent spikes in traffic, causing latency increases. Which strategy should the company use to ensure consistent low latency during traffic spikes?

A.Enable autoscaling on the Vertex AI endpoint with appropriate min and max nodes.
B.Manually scale the deployed model to a larger machine type during peak hours.
C.Reduce the number of prediction nodes to minimize overhead.
D.Switch to batch prediction to handle all requests asynchronously.
AnswerA

Autoscaling adds or removes endpoint nodes in response to traffic, so capacity grows during spikes rather than queuing requests on fixed replicas. Setting min nodes preserves a warm baseline and max nodes caps cost, directly addressing the intermittent latency increases described.

Why this answer

Vertex AI endpoints support autoscaling, which dynamically adjusts the number of prediction nodes based on incoming traffic. By setting appropriate min and max nodes, the endpoint can scale up during traffic spikes to maintain low latency and scale down during low traffic to reduce costs. This ensures consistent performance without manual intervention.

Exam trap

Google Cloud often tests the misconception that manual scaling or switching to batch prediction is a valid solution for real-time latency spikes, when in fact autoscaling is the only automated, cost-effective method for handling intermittent traffic on Vertex AI endpoints.

How to eliminate wrong answers

Option B is wrong because manually scaling to a larger machine type during peak hours is reactive, not proactive, and cannot respond instantly to intermittent spikes; it also incurs higher costs during all peak hours rather than scaling only when needed. Option C is wrong because reducing the number of prediction nodes would decrease capacity, worsening latency during traffic spikes rather than improving it. Option D is wrong because batch prediction is designed for asynchronous, offline processing of large datasets and does not provide real-time, low-latency responses required for online predictions.

349
MCQmedium

A media analytics team needs to copy 200 TB of historical log files from an on-premises NFS server into Cloud Storage, then transform them with a Dataflow batch job. The NFS server is reachable only from the corporate network, and the team wants to minimize transfer time and cost. Which approach should they use?

A.Create a Cloud VPN tunnel and use gcloud storage cp from an on-premises machine to upload the files.
B.Mount the NFS share on a Compute Engine VM and run gsutil -m rsync to copy the data to Cloud Storage.
C.Use Storage Transfer Service with an agent pool deployed in the corporate network to transfer directly from the NFS server to Cloud Storage.
D.Use Transfer Appliance to ship the data to Google, then load it into Cloud Storage.
AnswerC

Storage Transfer Service supports on-premises sources through agent pools, which run in the corporate network and pull data from NFS to Cloud Storage. This avoids staging through a VPN and is designed for large-scale, reliable transfers with scheduling and integrity checks, making it the right fit for 200 TB from an NFS source.

Why this answer

Storage Transfer Service with an agent pool is purpose-built for moving large datasets from on-premises sources like NFS to Cloud Storage. Agents run inside the corporate network, so no inbound firewall changes are needed, and the service handles parallelism, retries, and integrity checks. Manual rsync, Transfer Appliance, and single-machine VPN copies are slower or inappropriate at this scale.

Exam trap

The trap here is assuming a VPN plus command-line copy is sufficient for large on-premises transfers, when agent-based Storage Transfer Service is the scalable, managed option.

350
MCQeasy

You want to train a custom TensorFlow model on Vertex AI using a managed Jupyter notebook environment. Which service should you use?

A.Vertex AI Workbench
B.Cloud Datalab
C.Vertex AI Training
D.AI Platform Notebooks
AnswerA

Vertex AI Workbench provides managed Jupyter notebook instances with native integration into Vertex AI training, letting you author and launch custom TensorFlow training jobs from the same environment. It satisfies the managed notebook constraint without provisioning or patching your own compute.

Why this answer

Vertex AI Workbench is Google Cloud's managed Jupyter notebook environment integrated with Vertex AI, allowing you to develop and train custom TensorFlow models directly. It provides pre-built containers, GPU support, and seamless integration with Vertex AI Training and Pipelines.

Exam trap

PDE often tests the rebranding of AI Platform Notebooks to Vertex AI Workbench; candidates may pick the outdated name or confuse the notebook environment with the training service.

How to eliminate wrong answers

Option B is wrong because Cloud Datalab is a deprecated, older notebook service not integrated with Vertex AI's modern training workflows. Option C is wrong because Vertex AI Training is a service for running training jobs, not a managed notebook environment for interactive development. Option D is wrong because AI Platform Notebooks was the predecessor to Vertex AI Workbench and has been rebranded/replaced; it is no longer the current service name.

351
MCQmedium

A company wants to transform data using dbt (data build tool) on BigQuery. They have a CI/CD pipeline and need to version-control their transformations. Which setup is recommended?

A.Create Dataflow pipelines for each transformation
B.Deploy dbt models in a Cloud Build pipeline that runs dbt run
C.Use Cloud Composer to orchestrate dbt jobs
D.Run dbt directly on BigQuery using scripting
AnswerB

Running dbt models inside a Cloud Build pipeline executes the transformations on BigQuery while the dbt project files remain in source control, giving versioned, repeatable deployments. This satisfies the CI/CD and version-control constraints without manual execution.

Why this answer

Dbt is designed for version-controlled, SQL-based transformations, and integrating it with Cloud Build allows you to run `dbt run` as part of a CI/CD pipeline. This setup ensures that every change to dbt models is automatically tested and deployed, which aligns with the requirement for version control and automated deployment on BigQuery.

Exam trap

This question tests the distinction between orchestration (Cloud Composer) and CI/CD (Cloud Build), so candidates mistakenly choose Cloud Composer because they think scheduling equals version control, but the question explicitly requires version control and CI/CD, not just scheduling.

How to eliminate wrong answers

Option A is wrong because Dataflow pipelines are intended for stream or batch data processing using Apache Beam, not for version-controlled SQL transformations; they add unnecessary complexity and cost for simple transformation logic. Option C is wrong because Cloud Composer (Apache Airflow) is an orchestration tool for scheduling and monitoring workflows, not a CI/CD pipeline for version-controlled dbt models; while it can run dbt, it is not the recommended setup for version control and automated deployment. Option D is wrong because running dbt directly on BigQuery using scripting bypasses version control, CI/CD integration, and proper environment management, leading to manual, error-prone processes.

352
MCQmedium

A company is migrating their on-premises Apache Spark jobs to Dataproc. They want to minimize code changes and take advantage of serverless infrastructure. Which Dataproc feature should they use?

A.Dataproc clusters with preemptible VMs
B.Dataproc Workflow Templates
C.Dataproc Serverless Spark
D.Dataproc Jobs API with custom machine types
AnswerC

Dataproc Serverless Spark runs PySpark and Spark SQL workloads without provisioning clusters, so existing Apache Spark code executes with minimal modification. It directly satisfies both stem constraints: reducing code changes and consuming serverless infrastructure, since compute scales automatically and no cluster management is required.

Why this answer

Dataproc Serverless Spark is the correct choice because it allows the company to run Spark workloads without provisioning or managing clusters, minimizing code changes by using the same Spark APIs and libraries. This serverless infrastructure automatically scales resources and handles failures, aligning with the goal of reducing operational overhead while maintaining compatibility with existing Spark jobs.

Exam trap

Google Cloud often tests the distinction between 'serverless' and 'managed' services; the trap here is that candidates may confuse Dataproc Workflow Templates or Jobs API with serverless capabilities, but those still require cluster management, whereas Dataproc Serverless Spark truly abstracts the infrastructure.

How to eliminate wrong answers

Option A is wrong because preemptible VMs are cost-effective but still require managing a cluster and do not provide serverless infrastructure; they are prone to termination, which can disrupt jobs without proper checkpointing. Option B is wrong because Workflow Templates orchestrate job sequences on existing clusters but do not eliminate cluster management or provide serverless execution. Option D is wrong because the Dataproc Jobs API with custom machine types still requires a running cluster to submit jobs, thus not achieving serverless infrastructure or minimizing cluster management.

353
MCQeasy

A data engineer wants to quickly estimate the cost of running a BigQuery query before executing it. Which command-line tool or command should they use?

A.gcloud logging read
B.bq query --use_cache=false
C.gcloud bigtable queries run
D.bq query --dry_run
AnswerD

The bq query --dry_run flag validates the query and returns the bytes it would process without executing it or incurring charges. Multiplying those bytes by the on-demand rate gives a cost estimate before running the job.

Why this answer

The bq query --dry_run flag validates a query and returns the amount of data it would process without actually executing it, which is the standard way to estimate BigQuery on-demand query cost. Since BigQuery on-demand pricing is based on bytes scanned (currently $6.25 per TiB), the dry-run's reported bytes give a direct cost estimate. It also validates syntax and permissions without incurring charges.

Exam trap

PDE often tests the distinction between flags that change execution behavior (--use_cache=false) and flags that only estimate without executing (--dry_run), so candidates who don't know the dry-run semantics pick a plausible-sounding but wrong flag.

How to eliminate wrong answers

Option A is wrong because gcloud logging read retrieves log entries and has nothing to do with query cost estimation. Option B is wrong because --use_cache=false disables result caching, which would if anything increase cost by forcing a full scan — it does not estimate anything. Option C is wrong because gcloud bigtable queries run is not a valid BigQuery cost tool and targets Bigtable, a different NoSQL service entirely.

354
MCQeasy

You need to schedule a BigQuery query to run every day at 6:00 AM and write the results to a new table. The query is simple and does not require complex dependencies. You want a low-maintenance, serverless solution with minimal configuration. What should you do?

A.Use BigQuery scheduled queries to run the query daily and write to a destination table.
B.Deploy a Cloud Function triggered by Cloud Scheduler to execute the query via the BigQuery API.
C.Create a Dataflow pipeline with a scheduled trigger to run the query and write results.
D.Create a Cloud Composer DAG with a BigQueryInsertJobOperator scheduled with a cron expression.
AnswerA

BigQuery scheduled queries are a native, serverless feature that allows you to schedule SQL queries with a specified frequency. They require minimal configuration and no infrastructure management. You can set the schedule, destination table, and even notifications. This is the simplest solution for a single daily query.

Why this answer

BigQuery scheduled queries are a built-in, serverless feature that lets you schedule SQL queries to run at specified intervals. They require no infrastructure management and can write results to a destination table. For a simple daily query, this is the most straightforward and low-maintenance solution, avoiding the overhead of Cloud Composer, Cloud Functions, or Dataflow.

Exam trap

The trap here is over-engineering a simple scheduling requirement by choosing orchestration tools like Cloud Composer or custom code with Cloud Functions, when BigQuery's native scheduling suffices.

355
Multi-Selectmedium

A retail company is designing a Dataflow pipeline to process point-of-sale transactions from Cloud Pub/Sub and write to BigQuery. The pipeline must handle late-arriving data up to 24 hours and ensure that all data is written to BigQuery exactly once, even in the event of worker failures. Which two features should the engineer implement to meet these requirements? (Choose two.)

Select 2 answers
A.Use the Dataflow ExactlyOnce processing mode and enable the BigQueryIO.Write with FILE_LOADS and a trigger that fires after 24 hours.
B.Set the allowed lateness to 24 hours on the windowing strategy.
C.Enable the Dataflow ExactlyOnce processing mode and use the BigQueryIO.Write with STORAGE_WRITE_API and a deterministic deduplication key.
D.Use BigQueryIO.Write with the UseBeamSchema option and set the write disposition to WRITE_APPEND.
E.Configure the pipeline to use the BigQueryIO.Write with STREAMING_INSERTS and set the insertId based on a unique transaction identifier.
AnswersB, C

Setting allowed lateness to 24 hours allows the pipeline to accept and process data that arrives up to 24 hours after the window ends. This directly addresses the late-arriving data requirement. Combined with a proper trigger, it ensures that late data is included in the results and written to BigQuery.

Why this answer

To handle late-arriving data up to 24 hours, the pipeline must set allowed lateness to 24 hours on the windowing strategy. To achieve exactly-once writes to BigQuery, the engineer should use the Storage Write API with a deterministic deduplication key and enable Dataflow's ExactlyOnce mode. Together, these features satisfy both requirements.

Exam trap

The trap here is assuming that streaming inserts with insertId provide exactly-once semantics, but they only offer best-effort deduplication and can still produce duplicates.

356
MCQmedium

Your company is migrating an on-premises Apache Hadoop cluster to Google Cloud. The cluster runs Hive for SQL-like queries and stores data in HDFS. You want a managed service that minimizes operational overhead while supporting existing Hive scripts. Which Google Cloud service should you choose?

A.Cloud Dataflow
B.Cloud Bigtable
C.Cloud Dataproc
D.BigQuery
AnswerC

Cloud Dataproc is a managed Spark and Hadoop service that supports Hive, HDFS, and other Hadoop ecosystem tools. It allows you to run existing Hive scripts with minimal changes and provides cluster management, autoscaling, and integration with Cloud Storage. It is the ideal choice for migrating on-premises Hadoop workloads.

Why this answer

Cloud Dataproc is the correct choice because it is a managed Hadoop and Spark service that supports Hive, HDFS, and other ecosystem tools, allowing you to run existing Hive scripts with minimal changes. It reduces operational overhead by handling cluster provisioning, configuration, and scaling.

Exam trap

The trap here is assuming that a serverless data warehouse like BigQuery can run existing Hive scripts, overlooking that Dataproc is designed for Hadoop compatibility.

357
MCQmedium

A data engineer needs to design a schema in BigQuery for a dataset that contains customer orders. Each order has a header and multiple line items. Queries frequently need to retrieve the entire order including line items. Which schema design is MOST performant and cost-effective?

A.Store all data in a flat table with repeated order info per line item
B.Use nested and repeated fields (orders table with line items as REPEATED RECORD)
C.Normalize into separate orders and line_items tables, join on order_id
D.Use a partitioned table on order date
AnswerB

Nested repeated RECORDs store line items inside the parent order row, so retrieving a full order needs one read rather than a join. This eliminates shuffle and join costs, satisfying the frequent whole-order retrieval requirement while cutting bytes scanned.

Why this answer

BigQuery is optimized for denormalized schemas using nested and repeated fields (REPEATED RECORD). Storing line items as a repeated record within the orders table avoids expensive JOIN operations, reduces data shuffling, and allows BigQuery to scan only the necessary columns, making queries that retrieve entire orders with line items both faster and more cost-effective.

Exam trap

Google often tests the misconception that normalization (Option C) is always the best practice for relational databases, but in BigQuery's distributed, columnar architecture, denormalization with nested and repeated fields is the recommended pattern for performance and cost efficiency.

How to eliminate wrong answers

Option A is wrong because storing all data in a flat table with repeated order info per line item leads to massive data duplication (each line item repeats all order header fields), increasing storage costs and query scan size without leveraging BigQuery's native nested structure. Option C is wrong because normalizing into separate orders and line_items tables and joining on order_id introduces expensive JOIN operations that require shuffling and sorting large datasets, which is inefficient in BigQuery's distributed architecture and incurs higher slot usage and cost. Option D is wrong because partitioning on order date alone does not address the structural inefficiency of storing line items separately; while partitioning can improve query performance for date-range filters, it does not eliminate the need for JOINs or duplication, and the question specifically asks about retrieving entire orders with line items.

358
MCQmedium

An MLOps team wants to implement continuous deployment of ML models using Cloud Build and Vertex AI. They have a GitHub repository with training code. What should they use?

A.Deploy using Cloud Run
B.Vertex AI Pipelines integrated with Cloud Build
C.Cloud Functions to monitor GitHub
D.Cloud Build trigger with a custom step to run Vertex AI Training job and deploy
AnswerD

A Cloud Build trigger on the GitHub repository fires on commits, and a custom build step invokes a Vertex AI Training job then deploys the resulting model. This wires source control to training and deployment, satisfying the continuous deployment requirement.

Why this answer

It directly addresses the requirement for continuous deployment of ML models using Cloud Build and Vertex AI. A Cloud Build trigger can be configured to fire on GitHub commits, and a custom step in the Cloud Build pipeline can invoke a Vertex AI Training job, followed by deploying the trained model to Vertex AI Endpoints. This provides a fully automated CI/CD pipeline for ML models without additional orchestration overhead.

Exam trap

The trap here is that candidates may overthink the solution and choose Vertex AI Pipelines (Option B) because it is a dedicated ML orchestration tool, but the question specifically asks for integration with Cloud Build, and a simple Cloud Build trigger with custom steps is the most direct and efficient approach for continuous deployment.

How to eliminate wrong answers

Option A is wrong because Cloud Run is a serverless compute platform for containerized applications, not a service for training or deploying ML models in a Vertex AI context; it lacks native support for model versioning, evaluation, and endpoint management. Option B is wrong because Vertex AI Pipelines is an orchestration service for ML workflows, but integrating it with Cloud Build would add unnecessary complexity and is not the standard approach for a simple continuous deployment trigger; Cloud Build can directly invoke Vertex AI services without requiring a separate pipeline. Option C is wrong because Cloud Functions to monitor GitHub would require custom code to detect changes and trigger actions, which is less efficient and more error-prone than using Cloud Build's native GitHub trigger; Cloud Build already provides built-in event-driven triggers for GitHub repositories.

359
MCQeasy

A data engineer needs to load a 10 GB CSV file from GCS into BigQuery. The file contains some malformed rows that should be skipped. Which approach is most efficient?

A.Use Dataproc to run a Spark job that cleans the data and writes to BigQuery
B.Use the Storage Write API to stream each row, skipping bad ones in code
C.Use a Dataflow pipeline to read CSV, filter bad rows, and write to BigQuery
D.Use the bq command-line tool with the --max_bad_records flag
AnswerD

bq load with --max_bad_records skips malformed rows efficiently.

Why this answer

The `bq` command-line tool's `--max_bad_records` flag allows BigQuery's native CSV loader to skip malformed rows up to a specified limit during a load job. This is the most efficient approach for a one-time batch load of a 10 GB file, as it avoids the overhead of spinning up separate processing clusters (Dataproc, Dataflow) or streaming each row individually, leveraging BigQuery's optimized ingestion pipeline.

Exam trap

Google often tests the misconception that complex ETL pipelines (Spark, Dataflow) are always required for data cleaning, when in fact BigQuery's native load options like `--max_bad_records` can handle common malformed row scenarios directly and more efficiently.

How to eliminate wrong answers

Option A is wrong because using Dataproc to run a Spark job introduces unnecessary complexity and cost; for a simple load with malformed row skipping, a native BigQuery load job is far more efficient without needing a separate cluster. Option B is wrong because the Storage Write API is designed for real-time streaming, not batch loading a 10 GB file; streaming each row would be slower, more expensive, and less reliable than a single batch load with `--max_bad_records`. Option C is wrong because a Dataflow pipeline adds unnecessary processing overhead and cost; while it can filter bad rows, BigQuery's native load job with `--max_bad_records` achieves the same result more directly without requiring a separate data processing service.

360
MCQmedium

A data engineer is designing a BigQuery table for a clickstream dataset with frequent queries aggregating over user sessions. Each user session has multiple events, and the engineer wants to avoid joins for performance. Which schema design pattern should they use?

A.Use a normalized schema with separate tables for sessions and events, then join on session ID
B.Store each event as a separate row with session key and use clustering on session ID
C.Use partitioning on event timestamp and clustering on user ID
D.Use nested and repeated fields to store events within each session row
AnswerD

Nested and repeated fields let BigQuery store each session's events inside one row, eliminating the join between sessions and events that the stem explicitly wants to avoid. Aggregations over sessions then scan a single denormalised table, and BigQuery's columnar storage reads only the referenced nested attributes.

Why this answer

BigQuery's nested and repeated fields (using STRUCT and ARRAY) allow you to store multiple events within a single session row, eliminating the need for joins when aggregating over sessions. This denormalized pattern is ideal for clickstream data and improves query performance by keeping related data together.

Exam trap

PDE often tests the trade-off between normalization and denormalization in BigQuery; candidates may default to normalized designs or clustering/partitioning without recognizing that nested repeated fields eliminate joins.

How to eliminate wrong answers

Option A is wrong because a normalized schema with separate tables requires joins, which the engineer explicitly wants to avoid for performance. Option B is wrong because storing each event as a separate row with clustering on session ID still requires aggregation across rows and does not eliminate joins if session-level attributes are needed. Option C is wrong because partitioning on event timestamp and clustering on user ID helps with filtering but does not avoid joins for session-level aggregation.

361
MCQhard

A Dataproc cluster uses preemptible worker nodes to reduce costs. The cluster runs a long-running Spark job that occasionally experiences worker failures. How should the job be configured to handle preemptible worker failures gracefully?

A.Set spark.task.maxFailures to a high number to allow retries.
B.Disable preemptible workers for the job.
C.Use persistent disks for preemptible workers.
D.Enable automatic restart of the Spark driver on failure.
AnswerA

Raising `spark.task.maxFailures` lets individual tasks retry after a preemptible worker is reclaimed, so the job survives transient node loss without failing the stage. However, this only addresses task-level retries; Dataproc's enhanced flexibility mode, which pairs a standard primary pool with preemptible secondary workers, is what actually enables graceful recovery.

Why this answer

Preemptible VMs on Dataproc can be reclaimed at any time, causing executor loss mid-task. Setting spark.task.maxFailures to a higher value (default is 4) allows Spark to retry failed tasks on remaining executors instead of failing the entire job. This is the standard, low-cost way to tolerate preemptible worker churn without abandoning the cost savings.

Exam trap

PDE often tests whether candidates confuse driver-level resilience (automatic restart) with executor/task-level resilience (maxFailures), or mistakenly believe persistent disks prevent preemption.

How to eliminate wrong answers

Option B is wrong because disabling preemptible workers eliminates the cost benefit and defeats the purpose of the scenario; the question asks how to handle failures gracefully, not avoid them. Option C is wrong because persistent disks do not prevent preemption — preemptible VMs are still reclaimed, and persistent disks only preserve data, not running executors. Option D is wrong because restarting the Spark driver addresses driver failure, not executor/task failure caused by preemptible worker loss; the driver is typically on a non-preemptible node.

362
MCQmedium

Your company is building a real-time fraud detection system using Google Cloud. Transactions are streamed into Pub/Sub, and you need to process them with low latency (under 100ms per event) and aggregate data over sliding windows. Which Google Cloud service is best suited for this processing logic?

A.Dataflow
B.BigQuery streaming inserts with scheduled queries
C.Dataproc with Spark Streaming
D.Cloud Functions
AnswerA

Dataflow provides exactly-once, low-latency stream processing with native sliding window support.

Why this answer

Dataflow is the best choice because it provides a unified stream and batch processing model with native support for Pub/Sub, exactly-once semantics, and low-latency sliding window aggregations. Its autoscaling and millisecond-level checkpointing enable sub-100ms per event processing, which is critical for real-time fraud detection.

Exam trap

Google Cloud often tests the misconception that BigQuery streaming inserts can handle real-time per-event processing, but candidates overlook that scheduled queries add latency and BigQuery is not designed for stateful per-event aggregations with sliding windows.

How to eliminate wrong answers

Option B is wrong because BigQuery streaming inserts with scheduled queries cannot achieve sub-100ms latency per event; scheduled queries run on a periodic basis (e.g., every minute), introducing significant delay, and BigQuery is optimized for analytical queries, not per-event low-latency processing. Option C is wrong because Dataproc with Spark Streaming introduces higher startup and shuffle overhead, typically achieving latencies in the seconds range, and requires manual cluster management, making it unsuitable for consistent sub-100ms per event. Option D is wrong because Cloud Functions has a maximum timeout of 9 minutes and is designed for stateless, short-lived tasks; it lacks built-in support for stateful sliding window aggregations and cannot maintain per-key state across events without external services.

363
Multi-Selectmedium

You are designing a Dataflow pipeline for processing real-time clickstream data. The pipeline must group events into 30-second windows and handle late data up to 5 minutes. You want to output partial results every 10 seconds for low-latency monitoring. Which THREE configurations should you use? (Choose three.)

Select 3 answers
A.Use sliding windows of 30 seconds with a 10-second period
B.Use a trigger that fires after the end of the window
C.Use fixed windows of 30 seconds
D.Set allowed lateness to 5 minutes
E.Use a trigger with early firings every 10 seconds
AnswersC, D, E

Fixed 30-second windows partition the clickstream into non-overlapping intervals matching the grouping requirement. Combined with allowed lateness and early triggering, they satisfy the windowing constraint while the other settings handle late data and partial results.

Why this answer

Option C is correct because fixed (tumbling) 30-second windows partition the clickstream into non-overlapping 30-second intervals, which is the required grouping for this scenario. Option D is correct because setting allowed lateness to 5 minutes lets the pipeline retain window state and accept events that arrive up to 5 minutes after the window closes, matching the late-data requirement. Option E is correct because a trigger with early firings every 10 seconds emits speculative partial results before the window closes, providing the requested low-latency monitoring output.

Option A is not appropriate because sliding windows with a 10-second period would create overlapping 30-second windows and duplicate events across windows, which is not the specified grouping. Option B is not appropriate because a trigger that fires only after the end of the window would produce results only at window closure and would not deliver the required 10-second partial outputs.

Exam trap

Candidates might focus only on windowing and lateness, forgetting that early firing triggers are essential for periodic partial results as specified.

364
MCQeasy

You want to monitor the latency of messages in a Pub/Sub subscription. Which Cloud Monitoring metric should you use to see the age of the oldest unacknowledged message?

A.pubsub.googleapis.com/subscription/oldest_unacked_message_age
B.pubsub.googleapis.com/subscription/num_undelivered_messages
C.pubsub.googleapis.com/topic/send_request_count
D.pubsub.googleapis.com/topic/publish_latency
AnswerA

The metric `pubsub.googleapis.com/subscription/oldest_unacked_message_age` directly reports the age of the oldest unacknowledged message in a subscription, measured in seconds. This precisely satisfies the stem's requirement to monitor message latency, since backlog age reflects how long messages wait before acknowledgement.

Why this answer

The Cloud Monitoring metric pubsub.googleapis.com/subscription/oldest_unacked_message_age reports the age (in seconds) of the oldest unacknowledged message in a subscription, which is exactly the latency indicator requested. It is a subscription-level metric that directly reflects how long a message has been waiting without acknowledgment, making it the correct choice for monitoring message age.

Exam trap

PDE often tests whether candidates confuse publisher-side latency metrics (publish_latency, send_request_count) with subscription-side backlog/age metrics (oldest_unacked_message_age, num_undelivered_messages).

How to eliminate wrong answers

Option B is wrong because num_undelivered_messages counts how many messages are pending, not how old the oldest one is, so it measures backlog size rather than latency. Option C is wrong because topic/send_request_count measures publish request volume at the topic level, not message age in a subscription. Option D is wrong because topic/publish_latency measures the time to publish to the topic, which is a publisher-side latency, not the age of unacknowledged messages in a subscription.

365
MCQmedium

You are building a data pipeline that runs daily batch jobs on Dataproc, then loads results into BigQuery. You want to orchestrate the entire workflow, including dependencies between steps, retries, and monitoring. Which Google Cloud service is most appropriate?

A.Cloud Scheduler
B.Cloud Composer
C.Cloud Workflows
D.Dataflow
AnswerB

Cloud Composer, built on Apache Airflow, models the pipeline as a directed acyclic graph, so Dataproc job dependencies, retries and monitoring are orchestrated natively. It satisfies the stem's requirement for cross-service workflow orchestration, which BigQuery scheduled queries alone cannot provide.

Why this answer

Cloud Composer is the most appropriate service because it is a fully managed Apache Airflow service that provides workflow orchestration with DAGs, dependency management, retries, scheduling, and monitoring. It natively integrates with Dataproc and BigQuery operators, allowing you to define the entire pipeline as code. This matches the requirement for orchestrating a multi-step daily batch workflow.

Exam trap

The trap here is confusing orchestration services: candidates often pick Cloud Workflows because it sounds like a workflow tool, but the exam expects you to recognize that Airflow-based Cloud Composer is the standard for data pipeline orchestration with dependencies and retries.

How to eliminate wrong answers

Option A is wrong because Cloud Scheduler is only a cron-based job scheduler that triggers a single target (HTTP, Pub/Sub, App Engine); it cannot manage dependencies, retries, or multi-step workflows. Option C is wrong because Cloud Workflows is a serverless orchestration engine for API calls and service chaining, but it lacks the rich operator ecosystem and DAG-based dependency management of Airflow, making it less suited for data pipeline orchestration. Option D is wrong because Dataflow is a data processing service for stream and batch pipelines, not a workflow orchestrator; it does not manage dependencies between Dataproc and BigQuery steps.

366
Multi-Selecthard

A company is migrating an on-premises Hadoop cluster to Google Cloud. They need to run existing Spark jobs with minimal modification. Which THREE strategies should they consider? (Choose THREE.)

Select 3 answers
A.Migrate to BigQuery for all analytics.
B.Use Cloud Dataproc with Spark and Hive components.
C.Store data in Cloud Storage instead of HDFS.
D.Rewrite Spark jobs as Dataflow pipelines.
E.Use Dataproc Jobs API to submit jobs.
AnswersB, C, E

Cloud Dataproc provides managed Spark and Hive, so existing jobs run with minimal modification — satisfying the stem's constraint. Unlike re-platforming to BigQuery or Dataflow, which require rewriting jobs against different APIs, Dataproc preserves the Hadoop ecosystem's execution model and tooling on Google Cloud.

Why this answer

Option B is correct because Cloud Dataproc is a managed Hadoop/Spark service that natively runs Apache Spark and Hive, so existing Spark jobs can be migrated with minimal code changes. Option C is correct because Dataproc jobs commonly read and write data in Cloud Storage (gs://) instead of HDFS, which is the standard cloud-native replacement for on-premises HDFS storage and requires only path changes. Option E is correct because the Dataproc Jobs API lets you submit Spark jobs programmatically to a cluster, preserving the existing job-submission workflow with minimal modification.

Option A is not appropriate because migrating everything to BigQuery would require rewriting analytics logic and does not run existing Spark jobs as-is. Option D is not appropriate because rewriting Spark jobs as Dataflow pipelines is a significant code and framework change, contradicting the minimal-modification requirement.

Exam trap

The trap here is that candidates may assume BigQuery or Dataflow are the only Google Cloud data processing options, overlooking that Dataproc is specifically designed for minimal-change migrations of existing Spark/Hadoop workloads.

367
MCQhard

A Dataflow streaming pipeline writes to BigQuery and has run in production for months. The team wants to add a transformation and deploy the change with zero data loss and no interruption to the running pipeline. Which deployment approach should they use?

A.Run the updated pipeline in parallel with the old one, then delete the old pipeline after a fixed time.
B.Use a Cloud Scheduler job to redeploy the pipeline template on a cron and let the new job take over.
C.Update the pipeline in place by calling the Dataflow update method with the new job graph and a compatible transform name.
D.Stop the existing pipeline, then start a new pipeline with the updated code and the same job name.
AnswerC

Dataflow supports updating a running streaming job with a new pipeline definition when the transforms are named consistently and the update is compatible, preserving the existing state such as windows and timers. The service swaps in the new graph while maintaining exactly-once semantics for the sinks, so processing continues without draining or losing in-flight data. This is the intended zero-downtime deployment path.

Why this answer

Dataflow's update capability is the supported way to change a running streaming pipeline without stopping it. When the new graph keeps transform names consistent and the changes are compatible, the service swaps the graph while preserving streaming state and maintaining the sink's exactly-once guarantees. Stopping and restarting, running duplicate pipelines, or redeploying templates all either create a processing gap, risk data loss, or cause duplicate output.

Exam trap

The trap here is treating a pipeline update like a code redeploy, where any restart is fine, instead of recognizing that streaming state and exactly-once sinks must be preserved across the change.

368
Multi-Selecthard

A company building a real-time analytics pipeline with Pub/Sub and Dataflow. Which THREE best practices should they follow?

Select 3 answers
A.Use event time processing with watermarks and allowed lateness
B.Design idempotent sinks to handle duplicate outputs
C.Use exactly-once processing for all transforms
D.Use at-least-once delivery with deduplication in the pipeline
E.Use event time processing only for batch pipelines
AnswersA, B, D

Event time processing supports out-of-order data and ensures accurate windowing.

Why this answer

In streaming pipelines, event time processing with watermarks and allowed lateness is essential for handling out-of-order data. Watermarks track the progress of event time, and allowed lateness specifies how long to wait for late-arriving data before considering it as late, ensuring accurate windowed aggregations.

Exam trap

Google Cloud often tests the misconception that exactly-once processing must be applied uniformly across all pipeline transforms, when in practice it is only required at sinks and can be replaced by at-least-once with deduplication for better performance.

369
MCQeasy

A startup is building a real-time dashboard that shows aggregated metrics from social media feeds. They expect up to 10,000 events per second. The data must be near-real-time (< 30 seconds latency) and stored in BigQuery for historical analysis. They have limited experience managing infrastructure. The CTO suggests using Apache Kafka on Compute Engine for ingestion. However, the data engineer recommends a fully managed solution. Which approach should the team adopt?

A.Use Cloud Functions to ingest events directly into BigQuery
B.Use Apache Kafka on Compute Engine for ingestion, then use Dataflow to write to BigQuery
C.Use Cloud Pub/Sub for ingestion and Cloud Dataflow for streaming into BigQuery
D.Use App Engine to receive events and write to BigQuery
AnswerC

Cloud Pub/Sub absorbs the 10,000 events per second without capacity planning, and Dataflow's autoscaling streaming pipelines write into BigQuery within the sub-30-second window. This fully managed pairing removes the infrastructure burden that Kafka on Compute Engine would impose on a team with limited operational experience, satisfying the latency and skills constraints in the stem.

Why this answer

Cloud Pub/Sub provides a fully managed, scalable ingestion service that can handle 10,000+ events per second without infrastructure management, and Cloud Dataflow offers exactly-once, auto-scaling streaming into BigQuery with sub-30-second latency. This combination meets the near-real-time requirement while eliminating operational overhead, aligning with the data engineer's recommendation for a fully managed solution.

Exam trap

The trap here is that candidates may choose Option B (Kafka on Compute Engine) because Kafka is a common streaming tool, but the question emphasizes limited infrastructure experience and a fully managed solution, making the self-managed Kafka approach a distraction that ignores operational overhead.

How to eliminate wrong answers

Option A is wrong because Cloud Functions has a maximum invocation timeout of 9 minutes and is designed for event-driven, short-lived tasks, not sustained high-throughput ingestion of 10,000 events per second; it would also lack buffering and retry mechanisms for streaming into BigQuery. Option B is wrong because managing Apache Kafka on Compute Engine requires significant operational expertise for cluster setup, partitioning, and monitoring, contradicting the team's limited experience and the goal of a fully managed solution. Option D is wrong because App Engine is a web application platform, not a streaming ingestion service; it would introduce HTTP overhead and scaling bottlenecks for high-velocity event streams, and writing directly to BigQuery from App Engine would risk data loss without a buffer.

370
MCQeasy

Given the query plan, what is the most likely reason this query is efficient despite processing 10 billion rows?

A.The query uses a wildcard function.
B.The table is partitioned by sale_date.
C.The table is materialized.
D.The table is clustered by product_id.
AnswerB

Partitioning by sale_date lets the engine prune irrelevant partitions, scanning only the date range the query filters on rather than all 10 billion rows. This partition elimination satisfies the stem's efficiency constraint, since the plan shows reduced data processed despite the table's enormous total size.

Why this answer

Partitioning by sale_date enables partition pruning, which allows the query engine to scan only the relevant partitions instead of the entire 10-billion-row table. This drastically reduces the amount of data read and processed, making the query efficient even with a large total row count.

Exam trap

Google Cloud often tests the distinction between partitioning (which reduces scanned rows via pruning) and clustering (which only improves sorting and compression within partitions), leading candidates to mistakenly choose clustering as the primary efficiency driver.

How to eliminate wrong answers

Option A is wrong because using a wildcard function (e.g., SELECT *) typically increases I/O and processing overhead by reading all columns, which would not improve efficiency. Option C is wrong because a materialized table is a precomputed snapshot that can speed up queries, but it does not inherently reduce the number of rows scanned; the efficiency gain here comes from partition pruning, not materialization. Option D is wrong because clustering by product_id organizes data within partitions for better compression and filter performance, but without partition pruning, the query would still need to scan all 10 billion rows, so clustering alone does not explain the efficiency.

371
MCQmedium

You are using Looker to model data from BigQuery. You have a dimension that should be filtered by a user attribute (e.g., user's region). Which LookML concept allows you to apply dynamic row-level security based on user attributes?

A.Custom field
B.Derived table
C.Access filter
D.Required access grant
AnswerC

An access filter applies a user attribute to a dimension or field, restricting each user's query results to rows matching their attribute value. This delivers dynamic row-level security in Looker without duplicating models per region.

Why this answer

Access filters in LookML allow you to apply row-level security by referencing user attributes, which are values passed from the Looker user's account or via SSO. By using an access filter on a dimension, you can dynamically restrict the data a user sees based on their attribute (e.g., region), ensuring they only view rows matching their assigned region. This is the standard LookML mechanism for dynamic row-level security.

Exam trap

PDE often tests the distinction between modeling constructs (derived tables, custom fields) and security constructs (access filters) — candidates pick derived tables thinking they can embed security logic, but derived tables are for data transformation, not dynamic user-based filtering.

How to eliminate wrong answers

Option A is wrong because a custom field is a user-defined dimension or measure created in the Looker UI or LookML for ad-hoc analysis, not a security mechanism for row-level filtering. Option B is wrong because a derived table is a LookML construct that defines a subquery or SQL-based table for modeling purposes, not a security feature for applying user-attribute-based filters. Option D is wrong because 'required access grant' is not a valid LookML concept — the correct term for row-level security in Looker is access filter (or access_grant in some contexts, but the standard term is access filter).

372
MCQeasy

A media company runs a batch data pipeline on Cloud Dataflow that ingests log files from Cloud Storage, transforms them, and writes results to BigQuery for analytics. The pipeline runs daily and has been stable for months. Recently, the source log format changed: a new optional field was added to some records. The pipeline started failing with ParseErrors for rows that contain the new field. The error logs show that the Dataflow job uses a hardcoded JSON schema that does not include the new field. The Dataflow pipeline logs are written to Stackdriver Logging, but no alerts are configured. The team wants to ensure that future schema changes do not break the pipeline and that failures are detected promptly. The team has limited experience with streaming and wants to keep the batch approach. Which course of action should the team take to improve solution quality?

A.Create a Cloud Monitoring alert on any PipelineError log entries from the Dataflow job, and set up a runbook to manually fix schema mismatches within one hour.
B.Schedule a Cloud Function to run every hour that checks the latest log file headers and compares them to the pipeline schema, sending an alert if differences are found.
C.Use BigQuery dry run queries to validate the schema before loading data, and if a mismatch is detected, block the pipeline run and notify the team via email.
D.Implement schema validation and evolution using a schema registry (e.g., AVRO) in the Dataflow pipeline, and configure Stackdriver alerts on pipeline failure or error logs.
AnswerD

AVRO schema registry enforces compatibility and permits optional-field evolution, so added fields no longer trigger ParseErrors against the hardcoded schema. Stackdriver alerts on failure logs give prompt detection. This preserves the existing batch approach, matching the team's limited streaming experience and stated preference.

Why this answer

Option D is correct because it directly addresses both requirements: schema validation and evolution via a schema registry (e.g., AVRO) allows the Dataflow pipeline to handle added optional fields without hardcoded schema ParseErrors, and configuring Stackdriver (Cloud Monitoring) alerts on pipeline failure or error logs ensures prompt detection of future failures. A schema registry supports backward-compatible schema evolution, which is the standard way to prevent new optional fields from breaking batch ingestion. Option A only adds alerting and a manual runbook, so it detects failures but does not prevent schema changes from breaking the pipeline.

Option B's hourly Cloud Function header check is a custom, brittle workaround that does not integrate with Dataflow's parsing or guarantee schema compatibility. Option C's BigQuery dry runs validate BigQuery load compatibility, not the Dataflow pipeline's JSON parsing schema, so it would not prevent the ParseErrors.

373
MCQeasy

You need to create a Looker model that defines a 'sales' view based on a BigQuery table, with a measure for total revenue. Which LookML object defines the table and dimensions?

A.explore
B.view
C.model
D.dimension
AnswerB

The view object maps a LookML view to an underlying BigQuery table and declares its dimensions and measures, including total revenue. It satisfies the requirement to define the table structure and the revenue measure in one reusable object.

Why this answer

In LookML, a view defines the underlying database table and contains its dimensions (columns) and measures (aggregations like total revenue). The view is the fundamental building block that maps a physical table to a logical set of fields, and it is referenced by explores to make those fields queryable.

Exam trap

PDE often tests the confusion between view and explore — candidates who think the explore defines the table pick A, not realizing the view is where sql_table_name and dimensions live.

How to eliminate wrong answers

Option A is wrong because an explore defines the queryable join graph and the user-facing interface, not the table mapping or field definitions. Option C is wrong because a model is a container that groups explores and defines connection and access settings, not the table or its dimensions. Option D is wrong because a dimension defines a single column or derived field within a view — it is a component of a view, not the object that defines the table itself.

374
Multi-Selectmedium

A company wants to use Eventarc to trigger a Cloud Run service when new objects are created in a GCS bucket. They also need to filter events for a specific bucket and object prefix. Which THREE resources must exist or be created?

Select 3 answers
A.Cloud Storage bucket
B.Pub/Sub topic
C.Cloud Scheduler job
D.Eventarc trigger
E.Cloud Run service
AnswersA, D, E

The Cloud Storage bucket is the event source whose object-creation notifications Eventarc consumes. It must exist before any trigger can target it, satisfying the scenario's requirement to react to new objects, and its name forms part of the trigger's filtering criteria.

Why this answer

Option A (Cloud Storage bucket) is correct because Eventarc's Cloud Storage events originate from an actual GCS bucket, and the trigger's bucket filter must reference that existing bucket where objects are created. Option D (Eventarc trigger) is correct because the trigger is the resource that connects the event source to the destination, defines the event type (e.g., google.cloud.storage.object.v1.finalized), and applies the bucket and object-prefix attribute filters. Option E (Cloud Run service) is correct because it is the event destination that Eventarc invokes when a matching object-creation event fires.

Option B (Pub/Sub topic) is not required because Eventarc manages its own internal Pub/Sub transport for Cloud Storage events; you do not create or specify a topic. Option C (Cloud Scheduler job) is not required because scheduling is unrelated to event-driven object-creation triggers.

Exam trap

PDE often tests the misconception that you must manually create a Pub/Sub topic for Eventarc — Eventarc handles it automatically, so it's not a required resource.

375
MCQeasy

A company needs a messaging service for event-driven applications that require low cost for high-throughput, but can tolerate occasional message loss. Which Pub/Sub product should they choose?

A.Pub/Sub with pull subscriptions
B.Pub/Sub with dead letter topics
C.Pub/Sub with push subscriptions
D.Pub/Sub Lite
AnswerD

Pub/Sub Lite provisions capacity in zonal or regional reservations, giving a much lower throughput price than standard Pub/Sub. Its at-least-once delivery offers no replication guarantee across zones, matching the stated tolerance for occasional message loss at high volume.

Why this answer

Pub/Sub Lite is designed for cost-sensitive workloads with relaxed durability. Standard Pub/Sub offers at-least-once delivery and high durability. Push vs pull is irrelevant to cost.

Page 4

Page 5 of 10

Page 6

All pages