Courseiva

CCNA Monitoring And Alerting Questions

31 questions · Monitoring And Alerting topic · All types, answers revealed

1
MCQmedium

A Data Engineer needs to ensure that a notebook job is not consuming excessive costs. Which monitoring tool provides the best view of DBU consumption per job?

A.Cluster event logs
B.Delta Live Tables event logs
C.Billable Usage System Table
D.Notebook execution history
AnswerC

The Billable Usage System Table is the authoritative source for DBU consumption metrics. It associates usage with specific jobs, clusters, and tags, making it the most effective tool for calculating the cost of specific workflows and enforcing cost management policies.

Why this answer

The Databricks Billable Usage system table tracks DBU consumption at a granular level, including by job and cluster. Analyzing this data is essential for cost governance, allowing engineers to identify inefficient workflows and optimize resource allocation. By monitoring these trends, companies can ensure their data platform remains cost-effective, preventing unexpected budget overruns and providing the data necessary to justify investments in infrastructure or job refactoring.

Exam trap

Candidates often confuse 'Billable Usage' system tables with 'Cluster Logs' or 'Query History', failing to recognize that cost-specific tracking is handled by the dedicated billing system tables.

2
Multi-Selectmedium

A data engineer is responsible for a production Databricks SQL warehouse that serves multiple teams. The engineer needs to set up monitoring to detect when query performance degrades due to resource contention. Which two metrics should the engineer monitor to identify this issue? (Choose two.)

Select 2 answers
A.The number of queries waiting in the queue for the warehouse.
B.The total storage used by the Unity Catalog metastore.
C.The number of active clusters in the workspace.
D.The number of failed login attempts to the workspace.
E.The average time queries spend in the 'RUNNING' state.
AnswersA, E

A growing queue indicates that queries are waiting for compute resources, which is a direct sign of resource contention. Monitoring the queue length helps identify when the warehouse is oversubscribed. This metric is available in the Databricks SQL warehouse monitoring dashboard and can be used to trigger alerts or scaling actions.

Why this answer

Resource contention in a SQL warehouse manifests as queries waiting in the queue and increased query execution times. Monitoring queue length and average running time provides direct insight into whether the warehouse is oversubscribed. These metrics are available in the Databricks SQL warehouse monitoring dashboard and can be used to trigger scaling or alerting.

Exam trap

The trap here is confusing general workspace metrics, such as active clusters or storage usage, with SQL warehouse-specific performance indicators that actually reveal contention.

3
MCQmedium

A Data Engineer needs to monitor the health of a Delta Live Tables (DLT) pipeline. Which metric should they monitor to track the number of data quality violations over time?

A.pipeline_latency_seconds
B.expectations_violated_records
C.cluster_cpu_utilization
D.task_retry_count
AnswerB

This metric exposes the number of records that failed a quality expectation constraint during the execution of a DLT pipeline. Tracking this count allows engineers to identify data drift or schema issues, enabling proactive remediation and ensuring that final tables meet established business quality standards.

Why this answer

The 'expectations_violated_records' metric is specifically designed to track data quality issues in DLT pipelines. By integrating these metrics with Databricks SQL alerts or external tools like Grafana, engineers can gain visibility into data lineage and quality degradation. Monitoring these metrics is essential for maintaining pipeline reliability and ensuring that downstream data consumers receive accurate and validated data, preventing the propagation of corrupted data throughout the analytical ecosystem.

Exam trap

Candidates often confuse general pipeline execution logs with specific data quality tracking metrics like expectations_violated_records when monitoring validation issues.

4
MCQeasy

A data engineer has deployed a Databricks SQL dashboard that queries a gold-layer table. The dashboard is used by executives every morning. The engineer wants to be notified if the dashboard's underlying query fails or returns zero rows, which would indicate a data pipeline issue. Which Databricks feature should they use to set up this notification?

A.Databricks SQL alerts
B.Job alerts on the Databricks job that refreshes the gold table
C.Cluster metrics and logging in the Clusters UI
D.Delta Live Tables expectations
AnswerA

Databricks SQL alerts allow you to define a condition on a query result and trigger notifications when the condition is met. You can set an alert to fire if the query returns zero rows or fails, and configure email or webhook destinations. This directly addresses the need to monitor the dashboard's data freshness and query health.

Why this answer

Databricks SQL alerts are designed to monitor query results and trigger notifications based on conditions such as row count, value thresholds, or query failure. By creating an alert on the dashboard's query, the engineer can receive an email if the query fails or returns zero rows, ensuring timely awareness of data pipeline issues.

Exam trap

The trap here is confusing job-level alerts with query-level alerts; job alerts monitor execution status, while SQL alerts monitor the actual data returned by a query.

5
MCQeasy

Which Databricks feature should be used to gain observability into access patterns and security events across the entire workspace?

A.Delta Live Tables (DLT) logs
B.Job run history
C.System tables (Audit logs)
D.Cluster event logs
AnswerC

System tables store audit records for all user and system activities in the Databricks environment. They are the authoritative source for monitoring access patterns, ensuring compliance with security requirements, and identifying potential anomalies or security incidents across the entire Databricks workspace ecosystem.

Why this answer

System tables (specifically the audit log system table) provide comprehensive logs of all activities within a Databricks account. These tables allow engineers and security teams to monitor access patterns, identify unauthorized actions, and perform compliance reporting. Using system tables is critical for maintaining an audit trail, detecting potential threats, and ensuring that workspace operations adhere to organizational security policies and data governance standards.

Exam trap

Candidates often suggest using 'workspace logs' or manual audit scripts. They miss that 'System Tables' are the official, centralized, and queryable source of truth for all workspace-wide security events.

6
MCQmedium

Which capability is provided by Databricks' integration with cloud-native monitoring tools (e.g., CloudWatch, Azure Monitor)?

A.Automated data quality remediation.
B.Centralized metrics collection and dashboarding.
C.Direct execution of SQL queries on the warehouse.
D.Fine-grained control over Spark session configurations.
AnswerB

Cloud-native monitoring tools like CloudWatch or Azure Monitor ingest platform-wide metrics. They allow for the creation of unified dashboards that display Databricks performance alongside other cloud services, providing a comprehensive operational view for site reliability engineers.

Why this answer

Integrating Databricks with cloud-native monitoring provides a unified view of platform health. These tools aggregate logs and metrics from across the entire cloud environment, allowing for centralized dashboarding and advanced alerting. This integration is vital for enterprises maintaining a 'single pane of glass' strategy, as it ensures that Databricks infrastructure metrics are treated with the same governance and visibility standards as other cloud resources.

Exam trap

Candidates often think cloud-native tools replace Databricks monitoring. They fail to recognize that the integration is for aggregation and centralized dashboarding, not for replacing Databricks' internal observability features.

7
MCQhard

You are managing a large-scale data lakehouse. You notice that your Spark jobs are frequently failing due to disk space issues on worker nodes. Which monitoring feature should you implement to proactively capture this trend?

A.Enable standard Databricks job success/failure notifications.
B.Set up a Databricks SQL alert on the system.query_history table.
C.Monitor Spark shuffle spill metrics using the Ganglia UI or custom Spark listeners.
D.Schedule a daily scan of the underlying S3 or ADLS storage buckets.
AnswerC

Shuffle spill metrics are the leading indicator for disk pressure in Spark. When memory is insufficient, Spark spills shuffle data to the local disk. By monitoring these metrics through Ganglia or custom listeners, you can detect early signs of performance degradation and disk usage growth, allowing you to intervene proactively.

Why this answer

Implementing custom Spark listeners or utilizing the Spark UI metrics for 'Disk Spilling' is the most effective proactive measure. Disk space exhaustion is often a symptom of memory pressure, leading to excessive shuffle operations. By monitoring spill-to-disk metrics, you can identify jobs that require more memory or better partitioning, preventing job failure and optimizing compute efficiency before the disk threshold is reached.

Exam trap

Candidates often select cluster scaling or instance resizing instead of specific metrics like shuffle spill, which directly diagnoses out-of-disk errors caused by memory pressure.

8
MCQmedium

Which Databricks feature provides the most granular view of data quality metrics over time for a Delta Live Tables pipeline?

A.Pipeline Run logs
B.DLT Event Log system table
C.Cluster log files
D.SQL warehouse query history
AnswerB

The DLT event log system table contains detailed records for each pipeline update, including every expectation violation. It is the best source for performing historical analysis and creating trends of data quality health across the entire lifecycle of the data.

Why this answer

The DLT 'event log' (stored as a system table) is the most powerful tool for analyzing data quality. It records every expectation check, allowing for historical trend analysis of data quality violations. This is critical for data governance, as it provides auditability and helps engineers identify when and where data quality has degraded over time, enabling proactive fixes to source upstream data.

Exam trap

Candidates often select 'Delta Log' or 'Table History' instead of the 'DLT Event Log', failing to realize that DLT pipelines have a specific, dedicated log for quality expectations.

9
MCQmedium

A data engineer is configuring monitoring for a production Databricks cluster. Which TWO metrics are best suited to identify potential performance bottlenecks related to worker nodes?

A.Cluster CPU utilization percentage.
B.Number of active jobs currently running in the entire workspace.
C.Cluster memory usage percentage.
D.Total number of users logged into the Databricks workspace.
E.The version of the Databricks Runtime used by the cluster.
AnswerA, C

High CPU utilization across worker nodes is a primary indicator of compute-bound tasks or inefficient code. Monitoring this helps identify when a cluster is saturated, potentially leading to slow query execution or timeouts. Regularly tracking this metric allows for proactive auto-scaling configurations to handle varying workload demands effectively.

Why this answer

Monitoring worker nodes is vital for identifying bottlenecks before they cause job failures. Metrics like CPU utilization and memory usage are key indicators of node pressure. By identifying these patterns, engineers can optimize cluster configurations, adjust instance types, or tune spark code, ensuring that production workloads remain stable, performant, and cost-effective throughout their execution lifecycle.

Exam trap

Candidates often pick storage or network metrics like disk IOPS or network throughput, forgetting that CPU and memory are the primary indicators of worker node compute bottlenecks.

10
MCQeasy

Which of the following is the best practice for managing alerts for a mission-critical production pipeline?

A.Configure alerts to send emails to individual developers.
B.Use a central notification channel connected to an incident management tool.
C.Create a dashboard and check it manually every hour.
D.Disable alerts for development environments.
AnswerB

Integrating alerts with incident management systems (like PagerDuty or Opsgenie) ensures that failures are tracked, acknowledged, and escalated appropriately. This provides a professional, scalable approach to observability that guarantees reliable incident resolution and team accountability.

Why this answer

Centralizing alerts within Databricks and routing them to a dedicated incident management tool ensures that issues are tracked, assigned, and resolved systematically. Avoiding fragmented notification methods is key to operational maturity. This approach prevents alert fatigue, ensures that the correct personnel are notified during off-hours, and keeps a clear historical record of system issues for post-mortem analysis and continuous improvement of the data architecture.

Exam trap

Candidates often choose fragmented methods like individual user emails or custom scripts instead of leveraging a centralized incident management tool to avoid alert fatigue.

11
MCQmedium

A data engineer is reviewing the event log of a Databricks job that has just failed. They need to determine the exact cause of the failure. Which event type in the event log indicates that a task failed due to an exception?

A.RUN_FAILED
B.TASK_FAILED
C.JOB_FAILED
D.EXECUTION_FAILED
AnswerB

TASK_FAILED is the event type logged when a task within a job run fails due to an exception or error. It includes details such as the error message and stack trace, which are essential for diagnosing the root cause. This event is the most direct indicator of a task-level failure in the event log.

Why this answer

The TASK_FAILED event is logged when an individual task within a job run encounters an exception. It captures the error message, stack trace, and other relevant metadata, making it the primary source for diagnosing task-level failures. Other event types like JOB_FAILED indicate a higher-level failure but lack the specific task exception details.

Therefore, TASK_FAILED is the correct choice for determining the exact cause of a task failure.

Exam trap

The trap here is confusing JOB_FAILED with TASK_FAILED, assuming that a job failure event contains the specific task exception details.

12
MCQmedium

A data engineer wants to monitor the health of Delta Live Tables (DLT) pipelines and be alerted if a pipeline fails. Which approach is the most efficient and native way to achieve this?

A.Configure an external cron job to poll the Databricks Jobs API every minute for status updates.
B.Enable the 'Log to Workspace' feature and manually query the system logs using SQL every hour.
C.Define a Notification Destination in the DLT pipeline settings and associate it with the 'On failure' event.
D.Use the Databricks SQL Alerts tool to monitor the underlying Delta tables for new record counts.
AnswerC

Defining a Notification Destination is the recommended, built-in feature for monitoring pipeline lifecycle events. It ensures that failure notifications are sent directly to the specified endpoint, such as email or Webhooks. This approach is highly reliable, scalable, and simplifies management by centralizing monitoring configuration within the DLT pipeline definition itself.

Why this answer

Using Notification Destinations in the DLT pipeline settings is the most native method. It allows engineers to configure email or Slack alerts directly within the UI or JSON configuration. This is crucial for production reliability, ensuring that stakeholders receive immediate notifications regarding job failures or data quality issues without needing external orchestration or custom API scripts, maintaining observability across the entire pipeline lifecycle.

Exam trap

Candidates frequently assume they need external orchestration tools or custom webhook scripts to monitor DLT pipelines, ignoring native pipeline settings.

13
MCQeasy

A data engineer needs to receive an email notification when a Databricks job fails. The job is scheduled to run every hour. The engineer wants to configure this notification with minimal effort and without writing additional code. Which approach should the engineer use?

A.Create a Databricks SQL alert that queries the job's run history and triggers when a failure is detected.
B.Use the Databricks REST API to poll the job status every hour and send an email if the status is failed.
C.Configure the job's email notifications in the job settings to send alerts on failure to the engineer's email address.
D.Set up a webhook in the job configuration to call an external service that sends an email.
AnswerC

Databricks jobs have built-in email notification settings that can be configured directly in the job UI or via the API. Enabling failure notifications sends an email to specified recipients when the job fails. This requires no additional code and is the simplest way to meet the requirement.

Why this answer

Databricks jobs include built-in email notification settings that can be configured to send alerts on failure. This is the most straightforward and code-free method to receive email notifications for job failures, meeting the engineer's requirements for minimal effort.

Exam trap

The trap here is overcomplicating the solution by considering external services or custom polling, when the job's native notification settings already provide the required functionality without extra work.

14
Multi-Selecthard

A data engineer is responsible for monitoring a production Databricks job that runs critical ETL tasks. The job occasionally fails due to transient issues such as cloud storage throttling or network timeouts. The engineer wants to set up automated alerts that notify the team only when the job fails after all retries are exhausted. Which TWO actions should the engineer take to achieve this? (Choose two.)

Select 2 answers
A.Set up a Databricks SQL alert that queries the job run history and triggers when a run fails with a specific error code.
B.Enable notifications on the job to send an email or webhook when the job fails after retries.
C.Configure the job with a retry policy that specifies the number of retries and the interval between them.
D.Use the Databricks REST API to monitor job runs and trigger an alert via an external system when a run fails.
E.Create a scheduled notebook that checks the job status every minute and sends an alert if the status is FAILED.
AnswersB, C

Databricks jobs support built-in notifications that can be configured to fire on failure. When combined with a retry policy, these notifications are sent only after the job has exhausted its retries and ultimately fails. This is the native and most direct way to achieve the desired alerting behavior without custom code.

Why this answer

To alert only after all retries are exhausted, the engineer should configure a retry policy on the job and enable job notifications for failure. The retry policy handles transient issues automatically, and the notification fires only on final failure. Other methods either alert on every failure or require custom code to determine final failure status.

Exam trap

The trap here is assuming that any failure notification will suffice, when the requirement is to alert only after retries are exhausted, which is natively supported by combining retry policies with job notifications.

15
MCQhard

A data engineer is troubleshooting a Delta Live Tables pipeline that intermittently fails with 'StreamingQueryException: Job aborted due to stage failure'. The pipeline processes streaming data from a Kafka source. Which monitoring approach will best help identify the root cause of these intermittent failures?

A.Set up a Databricks SQL alert on the target Delta table's row count to detect missing data.
B.Use the Spark UI to inspect the DAG and stage details for each failed job run.
C.Enable and analyze the Delta Live Tables event log for detailed error messages and stack traces.
D.Monitor the cluster's CPU and memory utilization metrics in the Databricks workspace.
AnswerC

The Delta Live Tables event log captures detailed information about pipeline runs, including error messages, stack traces, and data quality metrics. For intermittent streaming failures, this log provides the granular context needed to pinpoint the exact cause, such as deserialization errors or Kafka connectivity issues, making it the most effective monitoring tool for this scenario.

Why this answer

The Delta Live Tables event log is the centralized source for pipeline run details, including error messages and stack traces. For intermittent streaming failures, it offers the most direct and detailed diagnostic information. Other monitoring tools either focus on resource usage or lack the necessary error context, making them less effective for root cause analysis in this scenario.

Exam trap

The trap here is assuming that general cluster metrics or Spark UI will provide sufficient error details, when the DLT event log is specifically designed to capture pipeline-level exceptions.

16
MCQmedium

A data engineer manages a Databricks SQL warehouse that serves a dashboard used by the finance team. The dashboard queries have become slow during peak hours, and the engineer suspects that some queries are scanning excessive data. Which system table should the engineer query to analyze query performance and identify expensive queries?

A.system.billing.usage
B.system.access.audit
C.system.compute.node_timeline
D.system.query.history
AnswerD

system.query.history contains detailed records of query executions, including query text, duration, rows read, bytes scanned, and user information. By querying this table, the engineer can identify long-running or high-scan queries that impact dashboard performance. This is the correct system table for analyzing query performance in Databricks SQL.

Why this answer

The system.query.history table is designed for query observability in Databricks SQL. It captures execution details, including query duration, rows produced, and bytes read, enabling engineers to pinpoint inefficient queries. Other system tables focus on compute metrics, audit events, or billing, and lack the query-level performance data needed to troubleshoot slow dashboards.

Exam trap

The trap here is confusing system tables that track access or billing with those that track query performance, leading to selection of an audit or billing table instead of query history.

17
Multi-Selecthard

A data engineer is responsible for a Delta Live Tables pipeline that ingests streaming data from multiple sources. The pipeline occasionally experiences delays, and the engineer needs to monitor the pipeline's health. Which two metrics should the engineer monitor to detect ingestion backlog and processing latency? (Choose two.)

Select 2 answers
A.The 'processingTime' metric in the streaming query progress
B.The number of records in the event log with severity 'ERROR'
C.The 'numInputRows' metric in the streaming query progress
D.The number of DBUs consumed by the pipeline cluster
E.The total size of the Delta table in storage
AnswersA, C

processingTime measures how long each micro-batch takes to process. If this value consistently increases, it indicates that the pipeline is unable to keep up with the incoming data rate, leading to increased latency. Monitoring processingTime helps identify performance bottlenecks and potential backlog in Delta Live Tables streaming pipelines.

Why this answer

numInputRows and processingTime are key metrics in the streaming query progress that directly reflect ingestion rate and processing duration. Monitoring both allows the engineer to detect when the pipeline is not keeping up with the source, causing backlog and increased latency. Other metrics like error counts, DBU usage, or storage size do not provide the necessary real-time performance insight.

Exam trap

The trap here is focusing on cost or error metrics rather than the streaming-specific throughput and latency metrics that indicate whether the pipeline is falling behind.

18
MCQeasy

A data engineer wants to monitor the performance of a Databricks cluster by tracking the average CPU utilization over time. Which Databricks feature should they use to visualize this metric?

A.Cluster metrics dashboard in the Databricks UI
B.Spark UI
C.Databricks SQL dashboards
D.Ganglia
AnswerA

The cluster metrics dashboard in the Databricks UI provides real-time and historical charts of CPU utilization, memory usage, and other metrics for a cluster. It is the built-in tool for visualizing cluster performance over time, making it the correct choice for this scenario.

Why this answer

The cluster metrics dashboard in the Databricks UI is designed to display cluster performance metrics such as CPU utilization, memory usage, and network activity. It provides both real-time and historical views, allowing engineers to monitor trends and identify performance issues. Other tools like Spark UI focus on job execution, while Databricks SQL dashboards are for data visualization, not infrastructure monitoring.

Therefore, the cluster metrics dashboard is the correct choice.

Exam trap

The trap here is confusing the Spark UI, which shows job-level details, with the cluster metrics dashboard, which shows infrastructure-level metrics.

19
MCQmedium

A data engineer has set up a Databricks SQL alert on a query that returns the count of failed jobs in the last hour. The alert is configured to trigger when the count exceeds 5. The engineer wants to receive notifications via email and also wants to view the alert history to understand past triggers. Which statement accurately describes the alert notification and history capabilities?

A.Alerts can only send notifications to email addresses and do not retain any history of triggers.
B.Alerts can only be configured to trigger once and then must be manually reset, and history is only available via the API.
C.Alerts can send notifications only to the alert owner's email, and history is stored for 7 days.
D.Alerts can send notifications to email, Slack, and other destinations, and the alert history is available in the Databricks SQL UI.
AnswerD

Databricks SQL alerts support email, Slack, PagerDuty, and webhook notifications. The alert history, including trigger times and query results, is accessible in the Databricks SQL UI under the alert's details. This statement correctly describes both the notification flexibility and the history feature, making it the accurate choice.

Why this answer

Databricks SQL alerts are versatile: they can notify via email, Slack, PagerDuty, or webhooks, and they maintain a history of triggers that can be reviewed in the UI. This allows teams to audit past alert conditions and responses. The other options either limit notification channels or incorrectly describe history retention, making them inaccurate.

Exam trap

The trap here is assuming that alerts only support email and that history is not readily available, when in fact Databricks provides multiple notification integrations and UI-based history.

20
MCQhard

An engineer notices that a SQL warehouse is frequently hitting 'Max Concurrency' limits. Which log should they consult to identify which specific queries are consuming most of the warehouse resources?

A.Cluster event logs
B.Query history system table
C.Job run history logs
D.Delta table access logs
AnswerB

The query history system table provides comprehensive details about every query executed, including the duration, user, warehouse, and CPU/memory footprint. It is the most appropriate source for identifying heavy queries causing concurrency limits to be reached.

Why this answer

The 'query_history' system table is the primary resource for analyzing query performance and resource consumption. By querying this table, engineers can attribute resource usage to specific users, notebooks, or warehouses. This is essential for troubleshooting concurrency bottlenecks, enabling teams to optimize expensive queries or adjust warehouse sizing to meet demand without impacting overall system performance or user productivity.

Exam trap

Candidates often choose 'Cluster Logs' to investigate SQL concurrency, missing that 'Query History' is the purpose-built system table for tracking resource consumption by specific SQL queries.

21
MCQhard

Refer to the exhibit. A Databricks administrator wants to restrict access to a specific SQL Alert. Based on the JSON policy, which statement accurately describes the current permission model for this alert?

A.The user is allowed to modify the alert's query criteria and notification frequency.
B.The user can see the alert, but cannot trigger a manual refresh or modify the alert settings.
C.The user has sufficient permissions to delete the alert from the Databricks SQL workspace.
D.This JSON policy provides the user with global access to all alerts within the SQL warehouse.
AnswerB

The 'CAN_VIEW' permission level strictly limits the user to viewing the alert's current status and metadata. It prevents the user from performing administrative actions like refreshing the alert, editing the logic, or altering notification parameters. This ensures that unauthorized users cannot inadvertently impact the alerting infrastructure or its outputs.

Why this answer

The exhibit displays an Access Control List (ACL) entry granting 'CAN_VIEW' permission to a specific user for a single alert resource. In Databricks, SQL Alerts are secured via the workspace-level permission model. Understanding these JSON-based policy definitions is essential for data engineers tasked with auditing governance and ensuring that sensitive monitoring data is only accessible to authorized personnel during security reviews.

Exam trap

Candidates often confuse view permissions with admin privileges, mistakenly assuming that someone with CAN_VIEW access can also modify alert configurations or manually trigger a refresh.

22
MCQmedium

A data engineer supports a Delta Live Tables pipeline that ingests streaming data from Kafka. The pipeline sometimes experiences latency spikes, and the engineer needs to determine whether the bottleneck is in the ingestion stage or in downstream transformations. They want to use built-in observability without adding external tooling. Which approach provides the most direct insight into per-stage event processing times within the DLT pipeline?

A.Configure Databricks SQL alerts on the pipeline's target tables to detect latency by comparing row counts over time.
B.Monitor cluster CPU utilization in the Clusters UI and correlate spikes with pipeline run times.
C.Use the Spark UI's SQL tab to inspect query plans for each streaming micro-batch.
D.Enable the event log and query the event_log table for flow_progress events to inspect stage-level metrics such as backlog and processing time.
AnswerD

The DLT event log captures flow_progress events that include stage-level metrics, including backlog bytes and records, as well as processing time. Querying the event_log table directly surfaces these metrics without external tools, allowing the engineer to compare ingestion versus transformation stages and pinpoint where latency accumulates.

Why this answer

The event log is the built-in observability mechanism for Delta Live Tables. By querying flow_progress events, the engineer gains stage-level metrics that directly show backlog and processing time for each flow. This allows precise identification of whether ingestion or transformation is the source of latency, without deploying external monitoring tools.

Exam trap

The trap here is assuming that cluster CPU metrics or Spark UI query plans can reveal DLT stage-level timing, when only the event log exposes flow_progress details.

23
MCQhard

When troubleshooting a job that frequently crashes due to 'Out of Memory' (OOM) errors, which TWO metrics or logs should be analyzed?

A.Driver and Executor logs.
B.System table audit logs.
C.Spark UI Executor memory metrics.
D.Workspace usage billing reports.
E.The number of active sessions in the SQL warehouse.
AnswerA, C

Driver and executor logs contain the stack traces and error messages that specifically indicate memory exhaustion. Reviewing these logs allows the engineer to determine whether the error occurred during a shuffle operation, a broadcast join, or while loading a large data partition into memory.

Why this answer

Analyzing Spark UI metrics and cluster logs is critical for resolving OOM errors. Examining the 'JVM heap usage' and 'executor memory usage' helps identify if the data partition size exceeds available memory. Addressing these issues is essential for stabilizing production jobs, as OOM errors are a leading cause of pipeline failure, resulting in data gaps and missed SLAs that impact downstream business decisions.

Exam trap

Candidates often look only at cluster-level infrastructure metrics rather than examining the Spark UI memory metrics and driver/executor logs, missing the specific JVM heap usage details needed to resolve OOM errors.

24
MCQeasy

A data engineer wants to monitor the health of a Delta Live Tables pipeline and receive alerts when the pipeline fails to meet its data quality expectations. The pipeline has several expectations defined. Which Databricks feature should the engineer use to set up these alerts?

A.Job notifications configured on the DLT pipeline's underlying job.
B.Databricks SQL dashboard with a query on the target Delta table.
C.Delta Live Tables event log with Databricks SQL alerts.
D.Cluster metrics in the Databricks workspace.
AnswerC

The Delta Live Tables event log records data quality metrics and expectation outcomes. By querying the event log with Databricks SQL alerts, engineers can set up notifications when expectations fail or metrics drop below thresholds. This is the native and most integrated way to monitor DLT pipeline health and data quality.

Why this answer

The Delta Live Tables event log captures detailed data quality metrics and expectation results. By using Databricks SQL alerts to query this log, engineers can set up precise alerts for data quality violations. Other options either lack the necessary data quality context or only provide high-level failure notifications.

Exam trap

The trap here is assuming that job notifications or cluster metrics can alert on data quality expectations, when only the DLT event log contains that granular information.

25
MCQmedium

Which action allows a Data Engineer to receive a Slack notification when a Delta Live Tables pipeline finishes successfully?

A.Configure a SQL Alert to watch the system metadata tables.
B.Use the 'Notifications' setting in the DLT pipeline configuration.
C.Create a Python script that polls the Jobs API every minute.
D.Add a post-pipeline task that sends an email to the Slack email gateway.
AnswerB

DLT pipelines include a dedicated notifications section in the settings. By configuring a destination—such as a webhook URL—engineers can receive automated notifications for pipeline events, including successful completions or failures, ensuring seamless integration with communication tools like Slack or Microsoft Teams.

Why this answer

Databricks supports native notification channels that can be configured for pipelines. By adding a webhook for a specific Slack channel to the pipeline settings, engineers can receive real-time updates on pipeline lifecycle events, including completion. This integration is vital for workflow orchestration, allowing teams to trigger subsequent processes or notify stakeholders immediately without manual intervention, thereby streamlining the overall data platform efficiency.

Exam trap

Candidates mistakenly look for external workflow schedulers or custom Python scripts to trigger notifications, overlooking built-in notification configurations.

26
MCQmedium

A data engineer wants to monitor the data quality of a Delta table over time. Which tool is most appropriate for this task?

A.Use Databricks SQL Alerts to count total rows in the table every hour.
B.Implement DLT Expectations within your pipeline code to validate data constraints.
C.Write a Spark job that manually scans the table for NULL values and sends an email.
D.Configure a cluster-level alert to trigger when the table metadata is updated.
AnswerB

DLT Expectations are built-in features that allow you to specify data quality rules, such as null checks or value ranges, directly in your pipeline. They provide native metrics and alerting capabilities, making it easy to monitor the health of data as it passes through the system without manual oversight.

Why this answer

Delta Live Tables (DLT) Expectations are the most appropriate and integrated way to enforce and monitor data quality. By defining constraints directly within the DLT pipeline code, data engineers can automatically validate data as it arrives, generate failure logs, and even quarantine invalid records. This observability is critical for maintaining high-quality datasets in a lakehouse architecture without requiring complex, separate validation workflows for every table.

Exam trap

Test-takers frequently choose generic Spark dataframe validations or third-party tools instead of native Delta Live Tables expectations built specifically for streaming pipelines.

27
MCQmedium

A Data Engineer wants to monitor cluster health proactively. Which metric is most effective for identifying that a cluster needs to be scaled up to handle increasing workload demands?

A.Query result set size
B.Cluster CPU utilization
C.Delta table file count
D.User login frequency
AnswerB

High CPU utilization is a direct indicator of compute saturation. When worker nodes sustain high CPU loads, processing throughput drops, leading to job latency. Monitoring this metric allows for the implementation of auto-scaling policies to add nodes, effectively balancing the workload across a larger compute resource pool.

Why this answer

Monitoring 'cluster_memory_utilization' or 'node_cpu_load' is essential for identifying saturation. When these metrics consistently approach high thresholds, it signals that the current compute power is insufficient for the data volume. Proactive scaling based on these metrics prevents job failures and performance degradation, ensuring that data pipelines meet their processing deadlines and maintain efficiency without wasting costs on over-provisioned infrastructure.

Exam trap

Candidates often select 'cluster logs' or 'job duration' as the primary metric. While useful, CPU and memory utilization are the direct indicators of resource saturation requiring vertical or horizontal scaling.

28
MCQhard

Refer to the exhibit. The alert is intended to trigger if the data in 'my_table' is older than one hour. Which query modification correctly implements this check?

A.SELECT COUNT(*) FROM my_table WHERE updated_at < now() - interval 1 hour
B.SELECT CASE WHEN max(updated_at) < now() - interval 1 hour THEN 1 ELSE 0 END
C.SELECT updated_at FROM my_table ORDER BY updated_at DESC LIMIT 1
D.SELECT datediff(now(), max(updated_at)) FROM my_table
AnswerB

This query returns 1 if the latest data is older than one hour, allowing an alert threshold of '> 0' to trigger. This is a standard pattern for binary state alerting where the result set needs to indicate a violation condition clearly.

Why this answer

To detect stale data, the alert must compare the latest record timestamp against the current time. If the latest 'updated_at' value is less than the threshold (one hour ago), it signifies that no data has arrived in the last hour. Monitoring data freshness is critical for time-sensitive analytics, ensuring that business stakeholders are working with current data and allowing teams to troubleshoot ingestion pipeline failures proactively.

Exam trap

Candidates often struggle with SQL syntax for time intervals, accidentally choosing options that use incorrect date functions or fail to account for the 'CASE WHEN' logic required for binary alert triggers.

29
MCQhard

A data engineer is troubleshooting a Databricks SQL query that occasionally fails with 'Query exceeded the maximum allowed execution time' on a shared SQL warehouse. The query is a complex aggregation over a large Delta table. The engineer needs to identify the root cause and ensure the query can complete successfully. Which action should the engineer take first?

A.Increase the SQL warehouse size to provide more compute resources for the query.
B.Set the Spark configuration 'spark.sql.adaptive.enabled' to false to disable adaptive query execution.
C.Examine the query profile in the Databricks SQL query history to identify stages with high data skew or spill.
D.Change the query to use a larger cluster by switching from a SQL warehouse to an all-purpose cluster.
AnswerC

The query profile provides detailed execution metrics, including time spent per stage, data skew, and spill to disk. These insights help pinpoint why the query exceeds the time limit, such as an inefficient join or skewed data distribution. Addressing these issues can allow the query to complete within the limit.

Why this answer

The query profile in Databricks SQL provides detailed execution metrics that can reveal the root cause of long-running queries, such as data skew or spill. By analyzing the profile first, the engineer can make targeted optimizations, such as repartitioning or rewriting the query, which may resolve the timeout without unnecessarily increasing compute resources.

Exam trap

The trap here is jumping to scaling up the warehouse, which may mask the symptom but not fix the underlying inefficiency that the query profile would reveal.

30
MCQhard

Refer to the exhibit. An engineer created this alert for a query. Under what condition will the alert status change to 'Triggered'?

A.When the query returns a row count greater than 3600.
B.When the 'duration' column value in the query output is strictly greater than 3600.
C.When the query execution takes longer than 3600 seconds to run.
D.When the query fails and returns an error code equal to 3600.
AnswerB

The 'op' field is set to '>', and the column is 'duration' with a threshold of '3600'. Therefore, any value exceeding 3600 returned by the query result set will satisfy the condition and transition the alert to the 'Triggered' state.

Why this answer

The alert is configured to monitor the 'duration' column from the specified query result. If the returned value exceeds 3600, the alert enters the 'Triggered' state. Understanding this threshold-based mechanism is vital for maintaining pipeline performance.

If queries consistently exceed this time, the alert notifies the engineer to investigate potential resource contention, suboptimal query plans, or data volume spikes, ensuring the stability and timely execution of data processing tasks.

Exam trap

Candidates often misinterpret the alert condition logic, confusing the 'Triggered' state with the 'OK' state or failing to distinguish between the 'threshold' value and the 'column' being monitored.

31
MCQhard

A data engineer is troubleshooting a production Databricks job that intermittently fails with 'SparkOutOfMemoryError'. The job processes large datasets with skewed partitions. The engineer wants to monitor the job to proactively detect memory pressure before failures occur. Which metric should the engineer monitor on the driver and executor nodes?

A.JVM heap memory usage
B.Disk I/O read/write throughput
C.CPU utilization percentage
D.Network throughput between executors
AnswerA

JVM heap memory usage on driver and executor nodes directly reflects memory pressure. When heap usage approaches the maximum, Spark may spill to disk or throw OutOfMemoryError. Monitoring heap usage allows the engineer to detect when memory is nearly exhausted and take action, such as increasing memory or repartitioning data, before job failure occurs.

Why this answer

JVM heap memory usage is the most direct indicator of memory pressure on Spark driver and executor nodes. By monitoring heap usage, engineers can detect when memory is approaching limits and intervene before OutOfMemoryError occurs. Other metrics like disk I/O, network throughput, and CPU utilization provide complementary information but do not directly measure memory consumption.

Exam trap

The trap here is focusing on symptoms like disk spilling or high CPU rather than the root cause—JVM heap memory usage—which directly signals memory pressure.

Ready to test yourself?

Try a timed practice session using only Monitoring And Alerting questions.