Databricks-DE-Assoc · domain
Troubleshooting, Monitoring, and Optimization
This domain covers diagnosing and fixing Databricks workloads: cluster selection, job failures, Delta table maintenance, and Spark performance tuning. Questions present a failure scenario (OOM, too many files, slow job) and ask which Databricks feature, command, or configuration resolves it, so you must map symptoms to the right remedy.
Focused practice
Practice Troubleshooting, Monitoring, and Optimization questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Troubleshooting, Monitoring, and Optimization
You must diagnose a failing or slow Databricks job and pick the correct fix: cluster type, Delta maintenance command, or Spark join strategy. The most important thing is matching the symptom to the right tool, especially VACUUM versus OPTIMIZE and Job versus All-Purpose clusters.
Choosing Job Clusters versus All-Purpose Clusters based on cost and reuse
Using OPTIMIZE, VACUUM, and partitioning to reduce excessive Delta metadata operations
Mitigating join OOM with broadcast hints, partitioning, or larger cluster memory
Recovering storage via VACUUM with the correct retention interval on Delta tables
Watch out for
Common Troubleshooting, Monitoring, and Optimization exam traps
- ▸Assuming VACUUM deletes files immediately; it respects the retention threshold and fails if below the safe default without disabling the check.
- ▸Confusing OPTIMIZE (compacts small files) with VACUUM (removes unreferenced files); they solve different problems.
- ▸Believing an All-Purpose Cluster is always better for jobs; Job Clusters are cheaper and isolated but cannot be shared interactively.
Question index
All Troubleshooting, Monitoring, and Optimization questions (42)
Click any question to see the full explanation, or start a practice session above.
A data engineer runs a nightly Databricks job that reads from a large Delta table and writes results to another Delta table. The job has been taking progressively longer each night. The engineer examines the Spark UI and sees that the 'SQL' tab shows a single stage with many small tasks, each processing very few records, and the 'Storage' tab shows that the source table has thousands of small files. Which action should the engineer take to improve performance?
Medium2A data engineer notices that a production Delta Lake table is experiencing slow read performance during concurrent write operations. The table contains millions of small files. Which action should the engineer take to resolve this performance degradation?
Medium3A data engineer is optimizing a Databricks job that performs a join between a large fact table and a small dimension table. The job is slow, and the engineer suspects that the join strategy is not optimal. Which TWO actions should the engineer take to improve performance? (Choose two.)
Hard4A data engineer is investigating a slow-running SQL query. Which TWO metrics from the Spark UI should the engineer examine to identify data skew?
Medium5A data engineer notices that a Databricks job processing a Delta table with 10,000 partitions runs slowly. The job filters on a column that is not the partition column, and the query plan shows that all partitions are being scanned. The engineer wants to improve performance without repartitioning the table. Which feature should be used?
Hard6Which Spark configuration property can be used to enable Adaptive Query Execution (AQE) in Databricks?
Medium7A data engineer notices that a scheduled Delta Lake maintenance pipeline is running significantly slower than expected. Upon checking the table history, they see that hundreds of tiny, fragmented data files have accumulated due to frequent streaming micro-batches. Which specific optimization command should the data engineer run first to resolve this file-size bottleneck?
Medium8Refer to the exhibit. A job fails with the provided error message. What is the most likely cause of this failure in a Databricks Delta Lake environment?
Hard9A job is failing with a 'Disk Space' error on the worker nodes. The code performs several large joins. Which configuration should the engineer adjust to mitigate the disk space usage?
Medium10A data engineer runs a nightly Databricks job that ingests data into a Delta table using multiple concurrent write streams. The engineer notices that some write transactions are failing with a ConcurrentAppendException. The job writes to the same partition of the table from different tasks. Which action should the engineer take to resolve this issue?
Medium11A data engineer is troubleshooting a slow-running Spark job on Databricks. Which TWO metrics in the Spark UI are most useful for identifying data skew?
Medium12A data engineer is monitoring a Databricks job and notices that the job's duration has gradually increased over the past week. The job reads a large Delta table, performs aggregations, and writes results to another Delta table. The engineer wants to identify the stage that is taking the most time. Which Spark UI tab should the engineer use to quickly identify the slowest stage?
Easy13A data engineer is investigating why a Databricks job that writes to a Delta table is experiencing performance degradation over time. The job performs frequent small appends. Which TWO actions should the engineer take to improve write performance? (Choose two.)
Hard14An engineer needs to identify the root cause of a job failure. Which THREE of the following are valid locations or methods to investigate the logs?
Medium15A data engineer needs to troubleshoot a job that is failing during the 'shuffle' phase. Which Spark UI tab should the engineer examine to analyze the shuffle partitions and identify potential imbalances?
Medium16An engineer has a large table that is frequently joined with small dimension tables. To optimize this, which optimization technique should be applied to the join operation?
Medium17A data engineer is using Databricks Jobs to run a nightly ETL pipeline. The job occasionally fails due to a transient network error when writing to an external database. The engineer wants to automatically retry the job a few times before marking it as failed. What is the most efficient way to configure this in Databricks?
Medium18A data engineer is monitoring a Databricks job that runs a Structured Streaming query. The engineer notices that the query's input rate is high, but the processing rate is low, and the batch duration is increasing over time. The query uses a Delta table as a source and writes to another Delta table. Which action should the engineer take to improve the streaming query's performance?
Hard19An engineer notices that a specific notebook job is consistently taking longer to start. They observe high 'initialization' times in the job logs. Which action should the engineer take to improve startup time?
Medium20A junior data engineer notices that a scheduled Databricks job running a heavy ETL notebook is failing intermittently due to cluster driver out-of-memory errors. Which TWO configuration changes or architectural adjustments should be implemented to resolve this issue? (Select exactly TWO)
Hard21A data engineer is monitoring a Databricks job and notices that the job's tasks are spending a significant amount of time in garbage collection (GC). The job processes large amounts of data with many small objects. Which action should the engineer take to reduce GC overhead?
Medium22A data engineer runs a Structured Streaming job that writes to a Delta table. The job processes data from a Kafka topic and uses a 10-minute watermark. After a few hours, the engineer notices that the streaming query's input rate is steady, but the processing rate has dropped significantly, and the batch duration has increased from 5 seconds to over 2 minutes. The job is running on a cluster with autoscaling enabled. Which action should the engineer take FIRST to diagnose the performance degradation?
Medium23Which Databricks feature should a data engineer use to view the execution plan, including information about the physical operators and data lineage, to troubleshoot a slow-running SQL query?
Easy24A data engineer is optimizing a Delta table that suffers from slow read performance due to small file sizes. Which command should the engineer execute to consolidate these small files into larger, more efficient files without altering the underlying table data?
Medium25A data engineer needs to monitor the costs associated with specific projects running on a shared Databricks workspace. Which feature should the engineer use to attribute these costs accurately?
Easy26A data engineer is optimizing a Databricks job that reads from a large Delta table and performs a join with a smaller table. The job is experiencing performance issues due to shuffling. The engineer wants to reduce the amount of data shuffled during the join. Which technique should the engineer use?
Hard27A data engineer is analyzing a Spark job that is failing with 'Out of Memory' (OOM) errors. Which configuration parameter should be tuned to increase the amount of memory allocated to the execution of joins and aggregations?
Medium28Which tool in Databricks provides real-time monitoring of cluster resource usage, including CPU, memory, and network throughput, for an active job?
Easy29A data engineer wants to monitor the health and performance of Databricks Jobs over time. Which feature should they use to visualize trends, such as job success rates and average execution times, across multiple runs?
Medium30A data engineer is investigating a Databricks job that failed overnight. The job's status in the Jobs UI shows 'Failed', and the engineer needs to view the error message and stack trace to determine the cause. Where should the engineer look to find the detailed error information for the failed run?
Easy31A data engineer runs a nightly Databricks job that reads a large Delta table and writes aggregated results to another Delta table. The cluster logs show many small files in the source table, and the job runtime has increased steadily over weeks. The engineer wants to reduce the number of files without rewriting the entire table. Which command should be used?
Medium32Which THREE strategies are recommended to improve the performance of reading from a Delta table in Databricks?
Medium33Which Databricks command should you use to recover storage space by removing files that are no longer referenced by a Delta table and are older than the retention period?
Medium34Refer to the exhibit. A data engineer receives this error when collecting data from a large transformation back to the driver node. Which approach should be used to fix this issue?
Hard35Refer to the exhibit. An engineer applies these configurations to a cluster. What is the primary benefit of enabling the Databricks IO Cache for a workload that involves repeatedly reading the same Delta tables?
Hard36A data engineer is optimizing a Databricks job that processes a large dataset. The job performs a join between a large Delta table and a small dimension table, then writes the result to a Delta table. The engineer notices that the join is causing a large shuffle and wants to reduce shuffle overhead. Which two actions should the engineer take to improve performance? (Choose two.)
Medium37A data engineer is troubleshooting a Databricks job that intermittently fails with a `SparkException: Job aborted due to stage failure: Task not serializable`. The job reads from a Parquet file, performs a transformation using a custom function defined in a Python class, and writes to a Delta table. The engineer suspects that the custom function is causing the issue. Which action should the engineer take to resolve the serialization error?
Hard38A data engineer is investigating why a Databricks job that reads from a Delta table is slow. The job performs a simple SELECT with a filter on a partition column. The engineer suspects that the table has many small files. Which Spark UI tab should be examined to confirm the number of files read?
Easy39When configuring a Databricks Job, which TWO factors most directly influence the choice between a 'Job Cluster' and an 'All-Purpose Cluster'?
Medium40Which action is recommended to resolve a scenario where a Databricks Job is failing due to excessive metadata operations on a Delta table with millions of files?
Medium41A Databricks job is failing due to 'Out of Memory' (OOM) errors during a join operation on two large datasets. Which TWO actions could help mitigate this issue?
Medium42A streaming job using Structured Streaming is lagging significantly behind the 'current time'. Which THREE of the following could be the root cause of this processing latency?
HardOther domains
All Databricks-DE-Assoc exam domains
Frequently asked questions
- What does the Troubleshooting, Monitoring, and Optimization domain cover on the Databricks-DE-Assoc exam?
- You must diagnose a failing or slow Databricks job and pick the correct fix: cluster type, Delta maintenance command, or Spark join strategy. The most important thing is matching the symptom to the right tool, especially VACUUM versus OPTIMIZE and Job versus All-Purpose clusters.
- How many questions are in this domain?
- This page lists all 42 Troubleshooting, Monitoring, and Optimization questions in the Databricks-DE-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Troubleshooting, Monitoring, and Optimization questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.