You must diagnose a failing or slow Databricks job and pick the correct fix: cluster type, Delta maintenance command, or Spark join strategy. The most important thing is matching the symptom to the right tool, especially VACUUM versus OPTIMIZE and Job versus All-Purpose clusters.
Start practicing
Troubleshooting, Monitoring, and Optimization — choose a session length
Free · No account required
Domain overview
This domain covers diagnosing and fixing Databricks workloads: cluster selection, job failures, Delta table maintenance, and Spark performance tuning. Questions present a failure scenario (OOM, too many files, slow job) and ask which Databricks feature, command, or configuration resolves it, so you must map symptoms to the right remedy.
Exam objectives
Choosing Job Clusters versus All-Purpose Clusters based on cost and reuse
Using OPTIMIZE, VACUUM, and partitioning to reduce excessive Delta metadata operations
Mitigating join OOM with broadcast hints, partitioning, or larger cluster memory
Recovering storage via VACUUM with the correct retention interval on Delta tables
Assuming VACUUM deletes files immediately; it respects the retention threshold and fails if below the safe default without disabling the check.
Confusing OPTIMIZE (compacts small files) with VACUUM (removes unreferenced files); they solve different problems.
Believing an All-Purpose Cluster is always better for jobs; Job Clusters are cheaper and isolated but cannot be shared interactively.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A Databricks job is failing due to 'Out of Memory' (OOM) errors during a join operation on two large datasets. Which TWO actions could help mitigate this issue?
2Which Databricks feature should a data engineer use to view the execution plan, including information about the physical operators and data lineage, to troubleshoot a slow-running SQL query?
3A streaming job using Structured Streaming is lagging significantly behind the 'current time'. Which THREE of the following could be the root cause of this processing latency?
4An engineer has a large table that is frequently joined with small dimension tables. To optimize this, which optimization technique should be applied to the join operation?
5Which Databricks command should you use to recover storage space by removing files that are no longer referenced by a Delta table and are older than the retention period?
6Which tool in Databricks provides real-time monitoring of cluster resource usage, including CPU, memory, and network throughput, for an active job?
7When configuring a Databricks Job, which TWO factors most directly influence the choice between a 'Job Cluster' and an 'All-Purpose Cluster'?
8Which Spark configuration property can be used to enable Adaptive Query Execution (AQE) in Databricks?
9Which THREE strategies are recommended to improve the performance of reading from a Delta table in Databricks?
10A data engineer wants to monitor the health and performance of Databricks Jobs over time. Which feature should they use to visualize trends, such as job success rates and average execution times, across multiple runs?
11A data engineer is optimizing a Delta table that suffers from slow read performance due to small file sizes. Which command should the engineer execute to consolidate these small files into larger, more efficient files without altering the underlying table data?
12Refer to the exhibit. A job fails with the provided error message. What is the most likely cause of this failure in a Databricks Delta Lake environment?
13A data engineer is troubleshooting a slow-running Spark job on Databricks. Which TWO metrics in the Spark UI are most useful for identifying data skew?
14An engineer notices that a specific notebook job is consistently taking longer to start. They observe high 'initialization' times in the job logs. Which action should the engineer take to improve startup time?
15A data engineer is analyzing a Spark job that is failing with 'Out of Memory' (OOM) errors. Which configuration parameter should be tuned to increase the amount of memory allocated to the execution of joins and aggregations?
16Which action is recommended to resolve a scenario where a Databricks Job is failing due to excessive metadata operations on a Delta table with millions of files?
17Refer to the exhibit. An engineer applies these configurations to a cluster. What is the primary benefit of enabling the Databricks IO Cache for a workload that involves repeatedly reading the same Delta tables?
18A data engineer needs to troubleshoot a job that is failing during the 'shuffle' phase. Which Spark UI tab should the engineer examine to analyze the shuffle partitions and identify potential imbalances?
19A data engineer notices that a production Delta Lake table is experiencing slow read performance during concurrent write operations. The table contains millions of small files. Which action should the engineer take to resolve this performance degradation?
20Refer to the exhibit. A data engineer receives this error when collecting data from a large transformation back to the driver node. Which approach should be used to fix this issue?
21A data engineer is investigating a slow-running SQL query. Which TWO metrics from the Spark UI should the engineer examine to identify data skew?
22A data engineer needs to monitor the costs associated with specific projects running on a shared Databricks workspace. Which feature should the engineer use to attribute these costs accurately?
23A job is failing with a 'Disk Space' error on the worker nodes. The code performs several large joins. Which configuration should the engineer adjust to mitigate the disk space usage?
24An engineer needs to identify the root cause of a job failure. Which THREE of the following are valid locations or methods to investigate the logs?
25A junior data engineer notices that a scheduled Databricks job running a heavy ETL notebook is failing intermittently due to cluster driver out-of-memory errors. Which TWO configuration changes or architectural adjustments should be implemented to resolve this issue? (Select exactly TWO)
26A data engineer notices that a scheduled Delta Lake maintenance pipeline is running significantly slower than expected. Upon checking the table history, they see that hundreds of tiny, fragmented data files have accumulated due to frequent streaming micro-batches. Which specific optimization command should the data engineer run first to resolve this file-size bottleneck?
27A data engineer runs a Structured Streaming job that writes to a Delta table. The job processes data from a Kafka topic and uses a 10-minute watermark. After a few hours, the engineer notices that the streaming query's input rate is steady, but the processing rate has dropped significantly, and the batch duration has increased from 5 seconds to over 2 minutes. The job is running on a cluster with autoscaling enabled. Which action should the engineer take FIRST to diagnose the performance degradation?
28A data engineer runs a nightly Databricks job that reads a large Delta table and writes aggregated results to another Delta table. The cluster logs show many small files in the source table, and the job runtime has increased steadily over weeks. The engineer wants to reduce the number of files without rewriting the entire table. Which command should be used?
29A data engineer notices that a Databricks job processing a Delta table with 10,000 partitions runs slowly. The job filters on a column that is not the partition column, and the query plan shows that all partitions are being scanned. The engineer wants to improve performance without repartitioning the table. Which feature should be used?
30A data engineer is investigating a Databricks job that failed overnight. The job's status in the Jobs UI shows 'Failed', and the engineer needs to view the error message and stack trace to determine the cause. Where should the engineer look to find the detailed error information for the failed run?
31A data engineer runs a nightly Databricks job that reads from a large Delta table and writes results to another Delta table. The job has been taking progressively longer each night. The engineer examines the Spark UI and sees that the 'SQL' tab shows a single stage with many small tasks, each processing very few records, and the 'Storage' tab shows that the source table has thousands of small files. Which action should the engineer take to improve performance?
32A data engineer is troubleshooting a Databricks job that intermittently fails with a `SparkException: Job aborted due to stage failure: Task not serializable`. The job reads from a Parquet file, performs a transformation using a custom function defined in a Python class, and writes to a Delta table. The engineer suspects that the custom function is causing the issue. Which action should the engineer take to resolve the serialization error?
33A data engineer is investigating why a Databricks job that reads from a Delta table is slow. The job performs a simple SELECT with a filter on a partition column. The engineer suspects that the table has many small files. Which Spark UI tab should be examined to confirm the number of files read?
34A data engineer is monitoring a Databricks job and notices that the job's duration has gradually increased over the past week. The job reads a large Delta table, performs aggregations, and writes results to another Delta table. The engineer wants to identify the stage that is taking the most time. Which Spark UI tab should the engineer use to quickly identify the slowest stage?
35A data engineer is optimizing a Databricks job that performs a join between a large fact table and a small dimension table. The job is slow, and the engineer suspects that the join strategy is not optimal. Which TWO actions should the engineer take to improve performance? (Choose two.)
36A data engineer is optimizing a Databricks job that processes a large dataset. The job performs a join between a large Delta table and a small dimension table, then writes the result to a Delta table. The engineer notices that the join is causing a large shuffle and wants to reduce shuffle overhead. Which two actions should the engineer take to improve performance? (Choose two.)
37A data engineer is optimizing a Databricks job that reads from a large Delta table and performs a join with a smaller table. The job is experiencing performance issues due to shuffling. The engineer wants to reduce the amount of data shuffled during the join. Which technique should the engineer use?
38A data engineer is using Databricks Jobs to run a nightly ETL pipeline. The job occasionally fails due to a transient network error when writing to an external database. The engineer wants to automatically retry the job a few times before marking it as failed. What is the most efficient way to configure this in Databricks?
39A data engineer is monitoring a Databricks job and notices that the job's tasks are spending a significant amount of time in garbage collection (GC). The job processes large amounts of data with many small objects. Which action should the engineer take to reduce GC overhead?
40A data engineer is investigating why a Databricks job that writes to a Delta table is experiencing performance degradation over time. The job performs frequent small appends. Which TWO actions should the engineer take to improve write performance? (Choose two.)
41A data engineer is monitoring a Databricks job that runs a Structured Streaming query. The engineer notices that the query's input rate is high, but the processing rate is low, and the batch duration is increasing over time. The query uses a Delta table as a source and writes to another Delta table. Which action should the engineer take to improve the streaming query's performance?
42A data engineer runs a nightly Databricks job that ingests data into a Delta table using multiple concurrent write streams. The engineer notices that some write transactions are failing with a ConcurrentAppendException. The job writes to the same partition of the table from different tasks. Which action should the engineer take to resolve this issue?
You must diagnose a failing or slow Databricks job and pick the correct fix: cluster type, Delta maintenance command, or Spark join strategy. The most important thing is matching the symptom to the right tool, especially VACUUM versus OPTIMIZE and Job versus All-Purpose clusters.
The Courseiva Databricks-DE-Assoc question bank contains 42 questions in the Troubleshooting, Monitoring, and Optimization domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Troubleshooting, Monitoring, and Optimization domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included