Databricks-DE-Pro · domain
Monitoring and Alerting
This domain covers observability on Databricks: query and warehouse metrics, Delta table quality monitoring, audit logging, and alerting on data freshness. Questions present realistic scenarios—warehouse concurrency pressure, stale-table alerts, workspace-wide security visibility—and ask you to pick the correct log, tool, or query modification. Expect log/feature selection plus interpreting SQL alert logic.
Focused practice
Practice Monitoring and Alerting questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Monitoring and Alerting
Be able to select the right observability surface—Query History for warehouse queries, audit logs for security events, Lakehouse Monitoring for data quality—and write a correct freshness alert comparing current_timestamp to the table's max timestamp with the proper interval.
Query History and warehouse query metrics for identifying resource-heavy queries
Databricks SQL alerts with freshness checks using current_timestamp and table timestamps
Audit log delivery and system tables for workspace access and security events
Lakehouse Monitoring (or DBSQL quality checks) for tracking Delta table data quality over time
Watch out for
Common Monitoring and Alerting exam traps
- ▸Confusing cluster driver/Spark logs with Query History when diagnosing SQL warehouse concurrency and per-query resource use
- ▸Writing freshness alerts that compare against a static timestamp or use the wrong time unit instead of current_timestamp minus interval
- ▸Assuming audit events are visible by default; they require audit log delivery or system table access to query
Question index
All Monitoring and Alerting questions (31)
Click any question to see the full explanation, or start a practice session above.
A Data Engineer needs to ensure that a notebook job is not consuming excessive costs. Which monitoring tool provides the best view of DBU consumption per job?
Medium2A data engineer is responsible for a production Databricks SQL warehouse that serves multiple teams. The engineer needs to set up monitoring to detect when query performance degrades due to resource contention. Which two metrics should the engineer monitor to identify this issue? (Choose two.)
Medium3A Data Engineer needs to monitor the health of a Delta Live Tables (DLT) pipeline. Which metric should they monitor to track the number of data quality violations over time?
Medium4A data engineer has deployed a Databricks SQL dashboard that queries a gold-layer table. The dashboard is used by executives every morning. The engineer wants to be notified if the dashboard's underlying query fails or returns zero rows, which would indicate a data pipeline issue. Which Databricks feature should they use to set up this notification?
Easy5Which Databricks feature should be used to gain observability into access patterns and security events across the entire workspace?
Easy6Which capability is provided by Databricks' integration with cloud-native monitoring tools (e.g., CloudWatch, Azure Monitor)?
Medium7You are managing a large-scale data lakehouse. You notice that your Spark jobs are frequently failing due to disk space issues on worker nodes. Which monitoring feature should you implement to proactively capture this trend?
Hard8Which Databricks feature provides the most granular view of data quality metrics over time for a Delta Live Tables pipeline?
Medium9A data engineer is configuring monitoring for a production Databricks cluster. Which TWO metrics are best suited to identify potential performance bottlenecks related to worker nodes?
Medium10Which of the following is the best practice for managing alerts for a mission-critical production pipeline?
Easy11A data engineer is reviewing the event log of a Databricks job that has just failed. They need to determine the exact cause of the failure. Which event type in the event log indicates that a task failed due to an exception?
Medium12A data engineer wants to monitor the health of Delta Live Tables (DLT) pipelines and be alerted if a pipeline fails. Which approach is the most efficient and native way to achieve this?
Medium13A data engineer needs to receive an email notification when a Databricks job fails. The job is scheduled to run every hour. The engineer wants to configure this notification with minimal effort and without writing additional code. Which approach should the engineer use?
Easy14A data engineer is responsible for monitoring a production Databricks job that runs critical ETL tasks. The job occasionally fails due to transient issues such as cloud storage throttling or network timeouts. The engineer wants to set up automated alerts that notify the team only when the job fails after all retries are exhausted. Which TWO actions should the engineer take to achieve this? (Choose two.)
Hard15A data engineer is troubleshooting a Delta Live Tables pipeline that intermittently fails with 'StreamingQueryException: Job aborted due to stage failure'. The pipeline processes streaming data from a Kafka source. Which monitoring approach will best help identify the root cause of these intermittent failures?
Hard16A data engineer manages a Databricks SQL warehouse that serves a dashboard used by the finance team. The dashboard queries have become slow during peak hours, and the engineer suspects that some queries are scanning excessive data. Which system table should the engineer query to analyze query performance and identify expensive queries?
Medium17A data engineer is responsible for a Delta Live Tables pipeline that ingests streaming data from multiple sources. The pipeline occasionally experiences delays, and the engineer needs to monitor the pipeline's health. Which two metrics should the engineer monitor to detect ingestion backlog and processing latency? (Choose two.)
Hard18A data engineer wants to monitor the performance of a Databricks cluster by tracking the average CPU utilization over time. Which Databricks feature should they use to visualize this metric?
Easy19A data engineer has set up a Databricks SQL alert on a query that returns the count of failed jobs in the last hour. The alert is configured to trigger when the count exceeds 5. The engineer wants to receive notifications via email and also wants to view the alert history to understand past triggers. Which statement accurately describes the alert notification and history capabilities?
Medium20An engineer notices that a SQL warehouse is frequently hitting 'Max Concurrency' limits. Which log should they consult to identify which specific queries are consuming most of the warehouse resources?
Hard21Refer to the exhibit. A Databricks administrator wants to restrict access to a specific SQL Alert. Based on the JSON policy, which statement accurately describes the current permission model for this alert?
Hard22A data engineer supports a Delta Live Tables pipeline that ingests streaming data from Kafka. The pipeline sometimes experiences latency spikes, and the engineer needs to determine whether the bottleneck is in the ingestion stage or in downstream transformations. They want to use built-in observability without adding external tooling. Which approach provides the most direct insight into per-stage event processing times within the DLT pipeline?
Medium23When troubleshooting a job that frequently crashes due to 'Out of Memory' (OOM) errors, which TWO metrics or logs should be analyzed?
Hard24A data engineer wants to monitor the health of a Delta Live Tables pipeline and receive alerts when the pipeline fails to meet its data quality expectations. The pipeline has several expectations defined. Which Databricks feature should the engineer use to set up these alerts?
Easy25Which action allows a Data Engineer to receive a Slack notification when a Delta Live Tables pipeline finishes successfully?
Medium26A data engineer wants to monitor the data quality of a Delta table over time. Which tool is most appropriate for this task?
Medium27A Data Engineer wants to monitor cluster health proactively. Which metric is most effective for identifying that a cluster needs to be scaled up to handle increasing workload demands?
Medium28Refer to the exhibit. The alert is intended to trigger if the data in 'my_table' is older than one hour. Which query modification correctly implements this check?
Hard29A data engineer is troubleshooting a Databricks SQL query that occasionally fails with 'Query exceeded the maximum allowed execution time' on a shared SQL warehouse. The query is a complex aggregation over a large Delta table. The engineer needs to identify the root cause and ensure the query can complete successfully. Which action should the engineer take first?
Hard30Refer to the exhibit. An engineer created this alert for a query. Under what condition will the alert status change to 'Triggered'?
Hard31A data engineer is troubleshooting a production Databricks job that intermittently fails with 'SparkOutOfMemoryError'. The job processes large datasets with skewed partitions. The engineer wants to monitor the job to proactively detect memory pressure before failures occur. Which metric should the engineer monitor on the driver and executor nodes?
HardOther domains
All Databricks-DE-Pro exam domains
Frequently asked questions
- What does the Monitoring and Alerting domain cover on the Databricks-DE-Pro exam?
- Be able to select the right observability surface—Query History for warehouse queries, audit logs for security events, Lakehouse Monitoring for data quality—and write a correct freshness alert comparing current_timestamp to the table's max timestamp with the proper interval.
- How many questions are in this domain?
- This page lists all 31 Monitoring and Alerting questions in the Databricks-DE-Pro question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Monitoring and Alerting questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.