Databricks-DE-Pro Monitoring and Alerting Practice Question
You are managing a large-scale data lakehouse. You notice that your Spark jobs are frequently failing due to disk space issues on worker nodes. Which monitoring feature should you implement to proactively capture this trend?
⚠ Common exam trap
Candidates often select cluster scaling or instance resizing instead of specific metrics like shuffle spill, which directly diagnoses out-of-disk errors caused by memory pressure.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Monitor Spark shuffle spill metrics using the Ganglia UI or custom Spark listeners.
Implementing custom Spark listeners or utilizing the Spark UI metrics for 'Disk Spilling' is the most effective proactive measure. Disk space exhaustion is often a symptom of memory pressure, leading to excessive shuffle operations. By monitoring spill-to-disk metrics, you can identify jobs that require more memory or better partitioning, preventing job failure and optimizing compute efficiency before the disk threshold is reached.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable standard Databricks job success/failure notifications.
Why it's wrong here
Standard success/failure notifications are reactive. They only alert the engineer after the job has already failed due to disk space exhaustion. Proactive monitoring requires identifying the trend of disk usage growth before the failure occurs, which standard job alerts are not designed to capture or report.
- ✗
Set up a Databricks SQL alert on the system.query_history table.
Why it's wrong here
The query_history table tracks SQL statements but does not provide granular metrics regarding worker node disk usage or Spark shuffle spill statistics. Relying on query metadata will not reveal the underlying cause of disk space exhaustion, which is a resource-level issue occurring on the worker nodes during execution.
- ✓
Monitor Spark shuffle spill metrics using the Ganglia UI or custom Spark listeners.
Why this is correct
Shuffle spill metrics are the leading indicator for disk pressure in Spark. When memory is insufficient, Spark spills shuffle data to the local disk. By monitoring these metrics through Ganglia or custom listeners, you can detect early signs of performance degradation and disk usage growth, allowing you to intervene proactively.
- ✗
Schedule a daily scan of the underlying S3 or ADLS storage buckets.
Why it's wrong here
Scanning cloud storage buckets relates to data volume growth, not the temporary disk usage on the Spark worker nodes. Temporary disk space issues are caused by shuffle operations and temporary cache files, which are independent of the total size of the persistent data stored in your data lake.
About these practice questions
Courseiva writes every Databricks-DE-Pro question from scratch — 267 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.