Courseiva

CCNA Maintaining and Automating Data Workloads Questions

8 of 83 questions · Page 2/2 · Maintaining and Automating Data Workloads · Answers revealed

76
MCQeasy

You need to orchestrate a simple, linear workflow that calls several Cloud Functions and API endpoints sequentially with conditional logic. The workflow should be defined as code and have minimal overhead. Which GCP service should you use?

A.Cloud Tasks
B.Workflows
C.Dataflow
D.Cloud Composer
AnswerB

Workflows orchestrates sequential steps with conditional branching, defined declaratively in YAML or JSON, and natively invokes Cloud Functions and HTTP endpoints. This satisfies the stem's need for a linear, code-defined workflow with minimal operational overhead.

Why this answer

Workflows is a serverless orchestration service that uses YAML/JSON to define workflows. It is ideal for simpler, linear or conditional orchestrations without the need for full Airflow infrastructure.

77
MCQmedium

A data engineer uses Cloud Composer to orchestrate a daily batch pipeline. A downstream task should only start after an upstream BigQuery load job finishes successfully and a specific file appears in Cloud Storage. Which combination of operators should the engineer use in the Airflow DAG?

A.BigQueryInsertJobOperator with wait_for_downstream=True
B.BigQueryInsertJobOperator and GCSObjectExistenceSensor with upstream dependency
C.DataflowPythonOperator and GCSObjectExistenceSensor
D.BigQueryOperator and FileSensor with downstream dependency
AnswerB

BigQueryInsertJobOperator runs the load job, while GCSObjectExistenceSensor pokes Cloud Storage until the file appears. Setting the sensor as an upstream dependency forces the downstream task to wait for both the successful load and the file's arrival, satisfying the dual trigger condition.

Why this answer

The engineer needs a BigQuery load job to finish successfully and a specific file to appear in Cloud Storage before a downstream task starts. The correct combination is BigQueryInsertJobOperator to run and wait for the BigQuery job, and GCSObjectExistenceSensor to check for the file, with the sensor set as an upstream dependency of the downstream task. This ensures both conditions are met before proceeding.

Exam trap

The trap is assuming that a single operator can handle both conditions or that wait_for_downstream covers external dependencies; candidates must recognize the need for a separate sensor and correct dependency direction.

How to eliminate wrong answers

Option A is wrong because BigQueryInsertJobOperator with wait_for_downstream=True only ensures downstream tasks wait for the BigQuery job, but it does not check for the file in Cloud Storage. Option C is wrong because DataflowPythonOperator is for Dataflow jobs, not BigQuery load jobs, and it lacks the file sensor. Option D is wrong because BigQueryOperator is deprecated in favor of BigQueryInsertJobOperator, and FileSensor checks local filesystem, not Cloud Storage; also the dependency direction is misstated.

78
MCQeasy

A data engineer must give a Dataproc Serverless for Spark batch workload permission to read objects from a specific Cloud Storage bucket and write to a BigQuery dataset, following least privilege. The workload runs as a custom service account. Which approach should be used?

A.Grant the workload's service account `roles/storage.objectViewer` on the bucket and `roles/bigquery.dataEditor` on the dataset.
B.Grant the workload's service account the basic `roles/editor` role on the project so both Cloud Storage and BigQuery calls succeed.
C.Enable the Cloud Storage and BigQuery APIs and rely on the default Compute Engine service account attached to the Dataproc Serverless workload.
D.Create a VPC Service Controls perimeter around the bucket and dataset and add the workload's service account to the access level.
AnswerA

Attaching IAM roles at the bucket and dataset level scopes permissions to exactly the resources the Spark workload touches. `roles/storage.objectViewer` permits reading objects without delete or create rights, and `roles/bigquery.dataEditor` allows writing data while excluding dataset administration. This satisfies least privilege for a Dataproc Serverless workload running as a custom service account.

Why this answer

Least privilege in Google Cloud means binding predefined or custom roles to the narrowest resource scope that still satisfies the workload. Bucket-level object viewer and dataset-level data editor grant precisely the read and write operations the Spark job requires. Broad project roles, default service accounts, and perimeter controls either overshoot the needed permissions or fail to grant them at all.

Exam trap

The trap here is treating VPC Service Controls or basic project roles as substitutes for scoped IAM bindings, when perimeter policies only constrain access and basic roles grant far more than the workload needs.

79
MCQhard

Your company uses Cloud Composer to orchestrate a complex data pipeline. You need to ensure that the pipeline can recover from failures and that tasks are retried automatically with exponential backoff. You also want to be alerted if a task fails after all retries. Which combination of features should you implement?

A.Use Airflow's SLA feature to trigger retries and send alerts.
B.Set retries and retry_delay on each task, and configure email alerts on task failure.
C.Set retries and retry_exponential_backoff on each task, and use Cloud Monitoring alerts based on Airflow metrics.
D.Configure a custom operator that implements retry logic and sends alerts via Pub/Sub.
AnswerC

Setting retries and retry_exponential_backoff on tasks enables automatic retries with exponential backoff. Cloud Composer exports Airflow metrics to Cloud Monitoring, allowing you to create alerts on task failures or other conditions. This combination provides robust retry logic and centralized alerting, meeting the requirements.

Why this answer

Tasks in Airflow can be configured with retries and retry_exponential_backoff to automatically retry failed tasks with increasing delays. Cloud Composer exports metrics to Cloud Monitoring, where you can set up alerts for task failures. This combination provides both automatic recovery and alerting, fulfilling the requirements.

Exam trap

The trap here is confusing Airflow's SLA feature with retry logic; SLA is for monitoring, not retrying.

80
MCQeasy

You need to schedule a recurring BigQuery query that aggregates data from a partitioned table and writes the results to a new table every day at 03:00 UTC. You want a fully managed solution with minimal operational overhead. What should you use?

A.Cloud Scheduler with a cron job that invokes a Cloud Function to run the query.
B.BigQuery scheduled queries.
C.Cloud Composer with a DAG that runs a BigQueryOperator.
D.A cron job on a Compute Engine instance that runs the bq query command.
AnswerB

BigQuery scheduled queries allow you to schedule SQL queries directly in the BigQuery UI or via API. They are fully managed, require no additional infrastructure, and support cron-like schedules. This is the simplest and most operationally efficient way to run a recurring query and write results to a table.

Why this answer

BigQuery scheduled queries are a fully managed feature that lets you schedule SQL queries to run at specified intervals. They require no additional infrastructure, support cron scheduling, and can write results to destination tables. This makes them the ideal choice for a recurring aggregation query with minimal operational overhead, unlike Cloud Functions, Composer, or VM-based cron jobs.

Exam trap

The trap here is overcomplicating the solution by using orchestration tools like Cloud Composer or Cloud Functions when a native BigQuery feature already provides the required scheduling with zero infrastructure management.

81
MCQhard

A BigQuery table has a REQUIRED column 'user_id' that now needs to accept NULL values due to upstream data changes. You want to alter the schema with minimal downtime and no data loss. What should you do?

A.Run `ALTER TABLE dataset.table ALTER COLUMN user_id DROP NOT NULL;`
B.Use the bq command: `bq update --set_nullable_fields user_id dataset.table`
C.Create a view that casts user_id to NULLABLE and use the view instead.
D.Drop the table and recreate it with the column as NULLABLE.
AnswerA

This BigQuery DDL statement changes the column to nullable without downtime or data loss.

Why this answer

BigQuery allows changing a column from REQUIRED to NULLABLE using the ALTER TABLE ALTER COLUMN SET DATA TYPE statement. This operation is a metadata change and does not require table recreation or data copy. Dropping and recreating the table would cause downtime and data loss.

Using a view is a workaround but doesn't change the underlying schema. Exporting and reloading is disruptive.

82
MCQmedium

Your company stores sensitive customer data in Cloud Storage. You need to inspect the data for personally identifiable information (PII) and de-identify it before sharing with a third party. Which Google Cloud service should you use?

A.Security Command Center
B.Dataplex
C.Cloud Data Loss Prevention (DLP)
D.Cloud KMS
AnswerC

Cloud DLP inspects data using infoType detectors to locate PII, then applies de-identification transformations such as masking, tokenisation or redaction. This satisfies both requirements: identifying sensitive customer data and removing it before sharing with the third party.

Why this answer

Cloud Data Loss Prevention (DLP) is the correct service because it is specifically designed to inspect, classify, and de-identify sensitive data such as PII in Cloud Storage. It provides built-in infoType detectors for over 150 types of PII and supports de-identification techniques like masking, tokenization, and encryption. This directly matches the requirement to inspect and de-identify data before sharing with a third party.

Exam trap

Candidates often confuse Cloud KMS (key management only) with Cloud DLP (inspection and de-identification), leading them to mistakenly choose Cloud KMS because they associate 'de-identify' with encryption, but Cloud KMS only manages keys, not the inspection or transformation of data content.

How to eliminate wrong answers

Option A is wrong because Security Command Center is a security and risk management platform that provides threat detection, vulnerability scanning, and compliance monitoring, but it does not have native capabilities to inspect or de-identify PII in data objects. Option B is wrong because Dataplex is a data governance and management service that helps organize, catalog, and manage data across lakes and warehouses, but it lacks built-in PII inspection and de-identification features. Option D is wrong because Cloud KMS is a key management service for creating, storing, and managing encryption keys, but it does not inspect data for PII or perform de-identification; it only provides encryption/decryption operations.

83
MCQhard

You manage a Cloud Composer 2 environment that runs a DAG with a task using the BigQueryInsertJobOperator. The task occasionally fails with 'rateLimitExceeded' when submitting many jobs in parallel. You want to limit the number of concurrent BigQuery jobs submitted by this DAG without affecting other DAGs in the same environment. What should you do?

A.Use a pool with a limited number of slots and assign the BigQuery tasks to that pool.
B.Increase the number of workers in the Cloud Composer environment to distribute the load.
C.Configure the BigQueryInsertJobOperator with a lower priority for the jobs.
D.Set max_active_tasks on the DAG to a low value.
AnswerA

Airflow pools allow you to limit parallelism for a specific set of tasks. By creating a pool with a small number of slots and assigning the BigQuery tasks to it, you control how many BigQuery jobs run concurrently, directly addressing the rate limit. This does not affect other DAGs or tasks that are not assigned to the pool, providing a targeted solution.

Why this answer

Airflow pools are designed to limit parallelism for a set of tasks. Creating a pool with a limited number of slots and assigning the BigQuery tasks to it ensures that only a controlled number of BigQuery jobs are submitted concurrently, preventing rateLimitExceeded errors. This approach is scoped to the specific tasks and does not affect other DAGs or tasks in the environment.

Exam trap

The trap here is confusing task-level concurrency controls like max_active_tasks with the need to throttle a specific external service, which is best handled by Airflow pools.

← PreviousPage 2 of 2 · 83 questions total

Ready to test yourself?

Try a timed practice session using only Maintaining and Automating Data Workloads questions.