Be able to select the right observability surface—Query History for warehouse queries, audit logs for security events, Lakehouse Monitoring for data quality—and write a correct freshness alert comparing current_timestamp to the table's max timestamp with the proper interval.
Start practicing
Monitoring and Alerting — choose a session length
Free · No account required
Domain overview
This domain covers observability on Databricks: query and warehouse metrics, Delta table quality monitoring, audit logging, and alerting on data freshness. Questions present realistic scenarios—warehouse concurrency pressure, stale-table alerts, workspace-wide security visibility—and ask you to pick the correct log, tool, or query modification. Expect log/feature selection plus interpreting SQL alert logic.
Exam objectives
Query History and warehouse query metrics for identifying resource-heavy queries
Databricks SQL alerts with freshness checks using current_timestamp and table timestamps
Audit log delivery and system tables for workspace access and security events
Lakehouse Monitoring (or DBSQL quality checks) for tracking Delta table data quality over time
Confusing cluster driver/Spark logs with Query History when diagnosing SQL warehouse concurrency and per-query resource use
Writing freshness alerts that compare against a static timestamp or use the wrong time unit instead of current_timestamp minus interval
Assuming audit events are visible by default; they require audit log delivery or system table access to query
Click any question to see the full explanation and answer options, or start a focused practice session above.
A Data Engineer needs to monitor the health of a Delta Live Tables (DLT) pipeline. Which metric should they monitor to track the number of data quality violations over time?
2Refer to the exhibit. An engineer created this alert for a query. Under what condition will the alert status change to 'Triggered'?
3Which Databricks feature should be used to gain observability into access patterns and security events across the entire workspace?
4A Data Engineer wants to monitor cluster health proactively. Which metric is most effective for identifying that a cluster needs to be scaled up to handle increasing workload demands?
5Which action allows a Data Engineer to receive a Slack notification when a Delta Live Tables pipeline finishes successfully?
6When troubleshooting a job that frequently crashes due to 'Out of Memory' (OOM) errors, which TWO metrics or logs should be analyzed?
7A Data Engineer needs to ensure that a notebook job is not consuming excessive costs. Which monitoring tool provides the best view of DBU consumption per job?
8Refer to the exhibit. The alert is intended to trigger if the data in 'my_table' is older than one hour. Which query modification correctly implements this check?
9Which of the following is the best practice for managing alerts for a mission-critical production pipeline?
10An engineer notices that a SQL warehouse is frequently hitting 'Max Concurrency' limits. Which log should they consult to identify which specific queries are consuming most of the warehouse resources?
11Which capability is provided by Databricks' integration with cloud-native monitoring tools (e.g., CloudWatch, Azure Monitor)?
12Which Databricks feature provides the most granular view of data quality metrics over time for a Delta Live Tables pipeline?
13A data engineer wants to monitor the health of Delta Live Tables (DLT) pipelines and be alerted if a pipeline fails. Which approach is the most efficient and native way to achieve this?
14Refer to the exhibit. A Databricks administrator wants to restrict access to a specific SQL Alert. Based on the JSON policy, which statement accurately describes the current permission model for this alert?
15A data engineer is configuring monitoring for a production Databricks cluster. Which TWO metrics are best suited to identify potential performance bottlenecks related to worker nodes?
16You are managing a large-scale data lakehouse. You notice that your Spark jobs are frequently failing due to disk space issues on worker nodes. Which monitoring feature should you implement to proactively capture this trend?
17A data engineer wants to monitor the data quality of a Delta table over time. Which tool is most appropriate for this task?
18A data engineer is reviewing the event log of a Databricks job that has just failed. They need to determine the exact cause of the failure. Which event type in the event log indicates that a task failed due to an exception?
19A data engineer wants to monitor the performance of a Databricks cluster by tracking the average CPU utilization over time. Which Databricks feature should they use to visualize this metric?
20A data engineer manages a Databricks SQL warehouse that serves a dashboard used by the finance team. The dashboard queries have become slow during peak hours, and the engineer suspects that some queries are scanning excessive data. Which system table should the engineer query to analyze query performance and identify expensive queries?
21A data engineer is responsible for a Delta Live Tables pipeline that ingests streaming data from multiple sources. The pipeline occasionally experiences delays, and the engineer needs to monitor the pipeline's health. Which two metrics should the engineer monitor to detect ingestion backlog and processing latency? (Choose two.)
22A data engineer supports a Delta Live Tables pipeline that ingests streaming data from Kafka. The pipeline sometimes experiences latency spikes, and the engineer needs to determine whether the bottleneck is in the ingestion stage or in downstream transformations. They want to use built-in observability without adding external tooling. Which approach provides the most direct insight into per-stage event processing times within the DLT pipeline?
23A data engineer has set up a Databricks SQL alert on a query that returns the count of failed jobs in the last hour. The alert is configured to trigger when the count exceeds 5. The engineer wants to receive notifications via email and also wants to view the alert history to understand past triggers. Which statement accurately describes the alert notification and history capabilities?
24A data engineer has deployed a Databricks SQL dashboard that queries a gold-layer table. The dashboard is used by executives every morning. The engineer wants to be notified if the dashboard's underlying query fails or returns zero rows, which would indicate a data pipeline issue. Which Databricks feature should they use to set up this notification?
25A data engineer is troubleshooting a Delta Live Tables pipeline that intermittently fails with 'StreamingQueryException: Job aborted due to stage failure'. The pipeline processes streaming data from a Kafka source. Which monitoring approach will best help identify the root cause of these intermittent failures?
26A data engineer is responsible for a production Databricks SQL warehouse that serves multiple teams. The engineer needs to set up monitoring to detect when query performance degrades due to resource contention. Which two metrics should the engineer monitor to identify this issue? (Choose two.)
27A data engineer is troubleshooting a production Databricks job that intermittently fails with 'SparkOutOfMemoryError'. The job processes large datasets with skewed partitions. The engineer wants to monitor the job to proactively detect memory pressure before failures occur. Which metric should the engineer monitor on the driver and executor nodes?
28A data engineer needs to receive an email notification when a Databricks job fails. The job is scheduled to run every hour. The engineer wants to configure this notification with minimal effort and without writing additional code. Which approach should the engineer use?
29A data engineer is troubleshooting a Databricks SQL query that occasionally fails with 'Query exceeded the maximum allowed execution time' on a shared SQL warehouse. The query is a complex aggregation over a large Delta table. The engineer needs to identify the root cause and ensure the query can complete successfully. Which action should the engineer take first?
30A data engineer is responsible for monitoring a production Databricks job that runs critical ETL tasks. The job occasionally fails due to transient issues such as cloud storage throttling or network timeouts. The engineer wants to set up automated alerts that notify the team only when the job fails after all retries are exhausted. Which TWO actions should the engineer take to achieve this? (Choose two.)
31A data engineer wants to monitor the health of a Delta Live Tables pipeline and receive alerts when the pipeline fails to meet its data quality expectations. The pipeline has several expectations defined. Which Databricks feature should the engineer use to set up these alerts?
Be able to select the right observability surface—Query History for warehouse queries, audit logs for security events, Lakehouse Monitoring for data quality—and write a correct freshness alert comparing current_timestamp to the table's max timestamp with the proper interval.
The Courseiva Databricks-DE-Pro question bank contains 31 questions in the Monitoring and Alerting domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Monitoring and Alerting domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included