Databricks-DE-Pro Monitoring and Alerting Practice Question
A data engineer is troubleshooting a Delta Live Tables pipeline that intermittently fails with 'StreamingQueryException: Job aborted due to stage failure'. The pipeline processes streaming data from a Kafka source. Which monitoring approach will best help identify the root cause of these intermittent failures?
⚠ Common exam trap
The trap here is assuming that general cluster metrics or Spark UI will provide sufficient error details, when the DLT event log is specifically designed to capture pipeline-level exceptions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable and analyze the Delta Live Tables event log for detailed error messages and stack traces.
The Delta Live Tables event log is the centralized source for pipeline run details, including error messages and stack traces. For intermittent streaming failures, it offers the most direct and detailed diagnostic information. Other monitoring tools either focus on resource usage or lack the necessary error context, making them less effective for root cause analysis in this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set up a Databricks SQL alert on the target Delta table's row count to detect missing data.
Why it's wrong here
A row count alert can indicate that data is not being written, but it does not explain why the pipeline failed. The intermittent nature of the failure requires access to error details, which a row count metric cannot provide. This method is reactive and does not help in diagnosing the underlying cause of the StreamingQueryException.
- ✗
Use the Spark UI to inspect the DAG and stage details for each failed job run.
Why it's wrong here
The Spark UI is valuable for performance tuning and understanding execution plans, but it may not retain logs for intermittent failures that occur after the cluster terminates. Delta Live Tables abstracts much of the Spark execution, and the event log is specifically designed to surface pipeline-level errors. The Spark UI alone would not provide the integrated error context needed.
- ✓
Enable and analyze the Delta Live Tables event log for detailed error messages and stack traces.
Why this is correct
The Delta Live Tables event log captures detailed information about pipeline runs, including error messages, stack traces, and data quality metrics. For intermittent streaming failures, this log provides the granular context needed to pinpoint the exact cause, such as deserialization errors or Kafka connectivity issues, making it the most effective monitoring tool for this scenario.
- ✗
Monitor the cluster's CPU and memory utilization metrics in the Databricks workspace.
Why it's wrong here
While resource metrics can indicate performance bottlenecks, they do not provide the detailed error information needed to diagnose a StreamingQueryException. The exception could be caused by data issues, network problems, or code bugs that are not reflected in CPU or memory usage. This approach lacks the specificity required for root cause analysis.
About these practice questions
This Databricks-DE-Pro question is part of Courseiva's 267-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.