Read Spark UI and driver/executor logs to find the real cause, then apply the matching fix: broadcast small dimension tables, salt skewed keys, avoid collect() on large data, and tune streaming triggers and Delta layout. Getting the root cause right matters more than memorizing tuning flags.
Start practicing
Troubleshooting and Tuning DataFrame Apps — choose a session length
Free · No account required
Domain overview
This domain covers diagnosing and fixing performance and failure problems in Spark DataFrame and Structured Streaming jobs on Databricks. Questions present logs, symptoms, or code and ask you to identify root causes like executor loss, driver OutOfMemoryError, growing streaming batch durations, or data skew, then select the correct mitigation using Spark UI, Delta Lake, and tuning techniques.
Exam objectives
Diagnosing ExecutorLostFailure from logs and Spark UI stage/task metrics
Fixing driver OutOfMemoryError from collect() by using write, take, or broadcast
Tuning Structured Streaming batch duration with trigger, watermark, and Delta optimization
Mitigating join data skew via broadcast hints, salting, or AQE skew join
Assuming ExecutorLostFailure is always a code bug, ignoring memory, shuffle spill, or node loss causes shown in logs
Calling collect() on large DataFrames and expecting driver memory to scale, instead of writing results or aggregating first
Treating slow streaming batches as a trigger problem while ignoring state growth, small files, or unoptimized Delta reads
Click any question to see the full explanation and answer options, or start a focused practice session above.
A Spark job is experiencing data skew during a join operation on a key column. Which strategy is most effective for mitigating this issue without changing the business logic?
2Which TWO of the following techniques effectively reduce the shuffle volume in a Databricks Spark job?
3Refer to the exhibit. Which action is the most likely cause of the error shown in the Spark job logs?
4A developer needs to optimize a Spark application that performs repetitive filtering and grouping on the same large DataFrame. Which feature should they implement to improve performance?
5Which THREE factors should a developer consider when choosing a partition count for a shuffle operation?
6Which configuration parameter should be adjusted to change the default number of partitions when reading from a shuffle-heavy operation?
7Refer to the exhibit. What is the most likely cause of the repeated ExecutorLostFailure messages in the logs?
8Which file format is best suited for performance-critical Spark applications that require efficient schema enforcement and column pruning?
9When a Spark job is stuck in a shuffle phase, what is the most effective first step to identify the root cause of the performance bottleneck?
10A job is reading a huge amount of data from a table, but only uses three columns. Which optimization technique will provide the most significant I/O performance benefit?
11Refer to the exhibit. What is the best way to resolve this error?
12When analyzing a Spark job's execution plan, what does a 'BroadcastHashJoin' indicate compared to a 'SortMergeJoin'?
13Your Spark application is experiencing severe data skew while performing a join between a large fact table and a small dimension table. Which technique should you apply to optimize performance?
14A developer notices a Spark job is failing with an OutOfMemoryError during a join operation on two large tables. The join key is highly skewed, causing one task to process significantly more data than others. Which technique should be applied to resolve this skew?
15Which action should be taken to optimize a Spark application that performs multiple operations on the same DataFrame and shows evidence of redundant re-computations in the Spark UI DAG visualization?
16A Spark job is running slower than expected due to excessive shuffling. Which TWO of the following techniques would directly reduce the volume of data transferred over the network?
17Refer to the exhibit. You are reviewing the logs for a Spark application and notice the warning regarding broadcasting a large task binary. What is the most likely cause and mitigation?
18You are debugging a PySpark DataFrame job on Databricks that performs multiple transformations and actions on a large delta table. You notice that the execution plan shows redundant computations where the same upstream DataFrame is evaluated repeatedly. Which transformation should you apply to optimize this workflow and avoid recomputing the upstream lineage?
19A developer runs a PySpark job on Databricks that reads a large Delta table, filters on a timestamp column, and writes results to another Delta table. The job takes 45 minutes, but the Spark UI shows that 90% of task time is spent reading from the source table. The developer wants to reduce the read time. Which action should the developer take?
20A Spark Structured Streaming job on Databricks reads from a Delta table and writes micro-batches to another Delta table with a 30-second trigger. After several hours, the batch duration grows from 4 seconds to over 60 seconds and the job falls behind. The source table is compacted regularly, and the cluster has enough CPU. Which tuning action is most likely to restore the original batch duration?
21A developer is troubleshooting a Spark job that fails with an OutOfMemoryError on the driver. The job collects a large DataFrame to the driver using .collect() and then processes it locally. The developer wants to avoid the driver OOM while still obtaining the results. Which approach is most appropriate?
22A PySpark job on Databricks repeatedly calls `df.count()` and `df.show()` inside a loop across 40 iterations, and the Spark UI shows the identical lineage being recomputed on every iteration even though the source Delta table is unchanged. You want to avoid re-executing the upstream transformations without materializing the data to disk. What should you do?
23A developer notices that a PySpark DataFrame transformation chain runs a full scan of a Delta table each time a new action is invoked, even though the source data has not changed. The developer wants to persist the intermediate DataFrame in memory across actions. Which method should be used?
24A Databricks notebook job calls a PySpark UDF built on a Python function that performs string parsing. The job completes successfully on a small sample, but on the full production dataset it fails with a PythonException and the executor logs show high garbage collection time. Which change is the most appropriate first step to make the job reliable without changing business logic?
25A developer is tuning a Spark job that performs a join between a large fact table and a medium-sized dimension table. The job suffers from data skew, with a few keys having a disproportionately large number of rows. The developer wants to mitigate the skew. Which two actions are most effective? (Choose two.)
26A Spark job performing a join between a 10 GB table and a 5 MB lookup table is running slowly, and the physical plan shows a SortMergeJoin. You want to avoid the shuffle. What should you do?
27A structured streaming DataFrame writes to a Delta table with a foreachBatch function that performs an upsert. After a cluster restart, the stream reprocesses some micro-batches and duplicate rows appear in the target table. The foreachBatch code already uses MERGE keyed on a unique id. Which change best prevents duplicates after restart?
28A Databricks job joins a 500 GB sales table with a 300 GB returns table on a customer_id key. A few customer_id values account for a large fraction of rows on both sides, and the job fails with executor OOM during the join. The developer wants to distribute the hot keys across more partitions without changing the query logic. Which technique should be used?
29A Spark job writes a large DataFrame to a Delta table partitioned by date. The job is taking much longer than expected, and the Spark UI shows that many tasks are writing very small files. You have already set spark.sql.shuffle.partitions to 200. What is the most effective way to reduce the number of small files written?
30A developer runs a PySpark job on Databricks that filters a Delta table and then calls count() and show(). The Spark UI shows two separate scans of the same table. The developer wants to avoid the second scan without changing the filter logic. Which action should be taken?
31A developer runs a PySpark job that joins a 10 GB DataFrame with a 50 MB lookup DataFrame. The job takes far longer than expected, and the Spark UI shows a SortMergeJoin with a large shuffle read and write for both sides. The developer wants the smallest change that most improves performance. Which action should the developer take?
32A PySpark DataFrame job on Databricks runs slowly. Inspection of the Spark UI shows that a shuffle stage writes 200 partitions but downstream stages process only a few, and the physical plan shows an Exchange before a filter. Which two changes are most likely to improve performance? (Choose two.)
33A Databricks job reads a Parquet dataset, applies a chain of `withColumn` transformations, and writes it back with `.write.mode("overwrite").parquet(path)`. The Spark UI shows 6000 small output files totaling 50 GB, and a downstream reader is slow because of per-file overhead. You want to reduce the file count while keeping the write as a Spark-native operation. What should you do?
34A Spark application reads a large CSV file and performs a series of transformations, including a filter and a join with a small lookup table. The job is running slowly, and the Spark UI shows that the CSV parsing stage is taking a long time. Which action would most improve performance?
35A streaming DataFrame job on Databricks writes to a Delta table every 10 seconds. Over several hours, the number of files in the target directory grows into the hundreds of thousands, and downstream reads slow dramatically. The job uses foreachBatch with a write that produces many small files per micro-batch. Which action should be taken to reduce the small-files problem for this streaming write?
36A developer notices that a DataFrame transformation chain is executed twice: once for a count action used for logging and again for a write action. The source is a large Delta table and the repeated scan adds several minutes. Which action avoids the duplicate computation with the least risk?
37A PySpark job reads a Delta table, applies a filter on a timestamp column, and then performs a window function partitioned by customer_id ordered by event_time. The job is slow, and the physical plan shows the filter is applied after the window. The developer wants the filter to reduce data before the window shuffle. Which action should the developer take?
38A job joining a large fact table with a small dimension table runs out of memory on executors during the join. The dimension table is about 40 MB after filtering and the configured spark.sql.autoBroadcastJoinThreshold is 10 MB. The join key is highly skewed in the fact table. Which action is most appropriate?
39A developer notices that a Spark DataFrame job on Databricks is running slowly and the Spark UI shows that many tasks are reading from a Delta table with a large number of small files. The job performs a filter on a date column and then aggregates results. Which optimization technique will most directly improve read performance in this scenario?
40A Databricks job fails with an OutOfMemoryError on the driver. The job collects a large DataFrame to the driver for local processing. Which action should you take to resolve this?
41A Spark job reads a large Parquet dataset, performs a groupBy on a high-cardinality column, and writes the result to a Delta table. The job fails with a FetchFailedException on a particular executor. The Spark UI shows that the executor had sufficient memory but the shuffle fetch failed due to a connection reset. Which configuration change is most likely to resolve this issue?
Read Spark UI and driver/executor logs to find the real cause, then apply the matching fix: broadcast small dimension tables, salt skewed keys, avoid collect() on large data, and tune streaming triggers and Delta layout. Getting the root cause right matters more than memorizing tuning flags.
The Courseiva Databricks-Spark-Assoc question bank contains 41 questions in the Troubleshooting and Tuning DataFrame Apps domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Troubleshooting and Tuning DataFrame Apps domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included