Courseiva

Databricks-Spark-Assoc · domain

Troubleshooting and Tuning DataFrame Apps

This domain covers diagnosing and fixing performance and failure problems in Spark DataFrame and Structured Streaming jobs on Databricks. Questions present logs, symptoms, or code and ask you to identify root causes like executor loss, driver OutOfMemoryError, growing streaming batch durations, or data skew, then select the correct mitigation using Spark UI, Delta Lake, and tuning techniques.

41 questions6 easy19 medium16 hard

Focused practice

Practice Troubleshooting and Tuning DataFrame Apps questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Troubleshooting and Tuning DataFrame Apps

Read Spark UI and driver/executor logs to find the real cause, then apply the matching fix: broadcast small dimension tables, salt skewed keys, avoid collect() on large data, and tune streaming triggers and Delta layout. Getting the root cause right matters more than memorizing tuning flags.

Diagnosing ExecutorLostFailure from logs and Spark UI stage/task metrics

Fixing driver OutOfMemoryError from collect() by using write, take, or broadcast

Tuning Structured Streaming batch duration with trigger, watermark, and Delta optimization

Mitigating join data skew via broadcast hints, salting, or AQE skew join

Why learners struggle

Why Troubleshooting and Tuning DataFrame Apps questions are commonly missed

RAM questions are commonly missed because learners confuse physical form factors (DIMM vs SO-DIMM) and fail to distinguish between memory speed (MHz) and latency (CL).

  • ·DIMM vs SO-DIMM — desktop vs laptop form factor confusion
  • ·DDR3 vs DDR4 vs DDR5 — notch position and voltage differences
  • ·MHz vs CL — speed vs latency trade-offs in performance
  • ·Single-channel vs dual-channel — bandwidth impact misconception
  • ·ECC vs non-ECC — error correction support in servers vs desktops
  • ·32-bit vs 64-bit — maximum addressable RAM limit

Watch out for

Common Troubleshooting and Tuning DataFrame Apps exam traps

  • ▸Assuming ExecutorLostFailure is always a code bug, ignoring memory, shuffle spill, or node loss causes shown in logs
  • ▸Calling collect() on large DataFrames and expecting driver memory to scale, instead of writing results or aggregating first
  • ▸Treating slow streaming batches as a trigger problem while ignoring state growth, small files, or unoptimized Delta reads

Question index

All Troubleshooting and Tuning DataFrame Apps questions (41)

Click any question to see the full explanation, or start a practice session above.

1

A PySpark job reads a Delta table, applies a filter on a timestamp column, and then performs a window function partitioned by customer_id ordered by event_time. The job is slow, and the physical plan shows the filter is applied after the window. The developer wants the filter to reduce data before the window shuffle. Which action should the developer take?

Hard
2

When analyzing a Spark job's execution plan, what does a 'BroadcastHashJoin' indicate compared to a 'SortMergeJoin'?

Medium
3

A Databricks notebook job calls a PySpark UDF built on a Python function that performs string parsing. The job completes successfully on a small sample, but on the full production dataset it fails with a PythonException and the executor logs show high garbage collection time. Which change is the most appropriate first step to make the job reliable without changing business logic?

Medium
4

A Spark job is experiencing data skew during a join operation on a key column. Which strategy is most effective for mitigating this issue without changing the business logic?

Medium
5

Which THREE factors should a developer consider when choosing a partition count for a shuffle operation?

Medium
6

Refer to the exhibit. What is the most likely cause of the repeated ExecutorLostFailure messages in the logs?

Hard
7

A developer runs a PySpark job on Databricks that filters a Delta table and then calls count() and show(). The Spark UI shows two separate scans of the same table. The developer wants to avoid the second scan without changing the filter logic. Which action should be taken?

Medium
8

A Spark job is running slower than expected due to excessive shuffling. Which TWO of the following techniques would directly reduce the volume of data transferred over the network?

Hard
9

A developer is tuning a Spark job that performs a join between a large fact table and a medium-sized dimension table. The job suffers from data skew, with a few keys having a disproportionately large number of rows. The developer wants to mitigate the skew. Which two actions are most effective? (Choose two.)

Hard
10

Which configuration parameter should be adjusted to change the default number of partitions when reading from a shuffle-heavy operation?

Medium
11

A developer runs a PySpark job that joins a 10 GB DataFrame with a 50 MB lookup DataFrame. The job takes far longer than expected, and the Spark UI shows a SortMergeJoin with a large shuffle read and write for both sides. The developer wants the smallest change that most improves performance. Which action should the developer take?

Easy
12

A Spark job reads a large Parquet dataset, performs a groupBy on a high-cardinality column, and writes the result to a Delta table. The job fails with a FetchFailedException on a particular executor. The Spark UI shows that the executor had sufficient memory but the shuffle fetch failed due to a connection reset. Which configuration change is most likely to resolve this issue?

Hard
13

A streaming DataFrame job on Databricks writes to a Delta table every 10 seconds. Over several hours, the number of files in the target directory grows into the hundreds of thousands, and downstream reads slow dramatically. The job uses foreachBatch with a write that produces many small files per micro-batch. Which action should be taken to reduce the small-files problem for this streaming write?

Medium
14

Which file format is best suited for performance-critical Spark applications that require efficient schema enforcement and column pruning?

Medium
15

Refer to the exhibit. What is the best way to resolve this error?

Hard
16

A developer notices that a PySpark DataFrame transformation chain runs a full scan of a Delta table each time a new action is invoked, even though the source data has not changed. The developer wants to persist the intermediate DataFrame in memory across actions. Which method should be used?

Easy
17

Which action should be taken to optimize a Spark application that performs multiple operations on the same DataFrame and shows evidence of redundant re-computations in the Spark UI DAG visualization?

Easy
18

Which TWO of the following techniques effectively reduce the shuffle volume in a Databricks Spark job?

Hard
19

A Databricks job fails with an OutOfMemoryError on the driver. The job collects a large DataFrame to the driver for local processing. Which action should you take to resolve this?

Easy
20

When a Spark job is stuck in a shuffle phase, what is the most effective first step to identify the root cause of the performance bottleneck?

Hard
21

A Spark application reads a large CSV file and performs a series of transformations, including a filter and a join with a small lookup table. The job is running slowly, and the Spark UI shows that the CSV parsing stage is taking a long time. Which action would most improve performance?

Medium
22

A Spark Structured Streaming job on Databricks reads from a Delta table and writes micro-batches to another Delta table with a 30-second trigger. After several hours, the batch duration grows from 4 seconds to over 60 seconds and the job falls behind. The source table is compacted regularly, and the cluster has enough CPU. Which tuning action is most likely to restore the original batch duration?

Hard
23

A structured streaming DataFrame writes to a Delta table with a foreachBatch function that performs an upsert. After a cluster restart, the stream reprocesses some micro-batches and duplicate rows appear in the target table. The foreachBatch code already uses MERGE keyed on a unique id. Which change best prevents duplicates after restart?

Hard
24

A developer notices that a DataFrame transformation chain is executed twice: once for a count action used for logging and again for a write action. The source is a large Delta table and the repeated scan adds several minutes. Which action avoids the duplicate computation with the least risk?

Easy
25

A Databricks job joins a 500 GB sales table with a 300 GB returns table on a customer_id key. A few customer_id values account for a large fraction of rows on both sides, and the job fails with executor OOM during the join. The developer wants to distribute the hot keys across more partitions without changing the query logic. Which technique should be used?

Hard
26

A PySpark DataFrame job on Databricks runs slowly. Inspection of the Spark UI shows that a shuffle stage writes 200 partitions but downstream stages process only a few, and the physical plan shows an Exchange before a filter. Which two changes are most likely to improve performance? (Choose two.)

Medium
27

A Databricks job reads a Parquet dataset, applies a chain of `withColumn` transformations, and writes it back with `.write.mode("overwrite").parquet(path)`. The Spark UI shows 6000 small output files totaling 50 GB, and a downstream reader is slow because of per-file overhead. You want to reduce the file count while keeping the write as a Spark-native operation. What should you do?

Medium
28

Refer to the exhibit. You are reviewing the logs for a Spark application and notice the warning regarding broadcasting a large task binary. What is the most likely cause and mitigation?

Medium
29

Refer to the exhibit. Which action is the most likely cause of the error shown in the Spark job logs?

Medium
30

A job is reading a huge amount of data from a table, but only uses three columns. Which optimization technique will provide the most significant I/O performance benefit?

Medium
31

A developer is troubleshooting a Spark job that fails with an OutOfMemoryError on the driver. The job collects a large DataFrame to the driver using .collect() and then processes it locally. The developer wants to avoid the driver OOM while still obtaining the results. Which approach is most appropriate?

Hard
32

A Spark job performing a join between a 10 GB table and a 5 MB lookup table is running slowly, and the physical plan shows a SortMergeJoin. You want to avoid the shuffle. What should you do?

Hard
33

A developer runs a PySpark job on Databricks that reads a large Delta table, filters on a timestamp column, and writes results to another Delta table. The job takes 45 minutes, but the Spark UI shows that 90% of task time is spent reading from the source table. The developer wants to reduce the read time. Which action should the developer take?

Medium
34

A developer notices a Spark job is failing with an OutOfMemoryError during a join operation on two large tables. The join key is highly skewed, causing one task to process significantly more data than others. Which technique should be applied to resolve this skew?

Medium
35

Your Spark application is experiencing severe data skew while performing a join between a large fact table and a small dimension table. Which technique should you apply to optimize performance?

Medium
36

A job joining a large fact table with a small dimension table runs out of memory on executors during the join. The dimension table is about 40 MB after filtering and the configured spark.sql.autoBroadcastJoinThreshold is 10 MB. The join key is highly skewed in the fact table. Which action is most appropriate?

Hard
37

A developer needs to optimize a Spark application that performs repetitive filtering and grouping on the same large DataFrame. Which feature should they implement to improve performance?

Hard
38

A Spark job writes a large DataFrame to a Delta table partitioned by date. The job is taking much longer than expected, and the Spark UI shows that many tasks are writing very small files. You have already set spark.sql.shuffle.partitions to 200. What is the most effective way to reduce the number of small files written?

Hard
39

A PySpark job on Databricks repeatedly calls `df.count()` and `df.show()` inside a loop across 40 iterations, and the Spark UI shows the identical lineage being recomputed on every iteration even though the source Delta table is unchanged. You want to avoid re-executing the upstream transformations without materializing the data to disk. What should you do?

Medium
40

A developer notices that a Spark DataFrame job on Databricks is running slowly and the Spark UI shows that many tasks are reading from a Delta table with a large number of small files. The job performs a filter on a date column and then aggregates results. Which optimization technique will most directly improve read performance in this scenario?

Easy
41

You are debugging a PySpark DataFrame job on Databricks that performs multiple transformations and actions on a large delta table. You notice that the execution plan shows redundant computations where the same upstream DataFrame is evaluated repeatedly. Which transformation should you apply to optimize this workflow and avoid recomputing the upstream lineage?

Medium

Frequently asked questions

What does the Troubleshooting and Tuning DataFrame Apps domain cover on the Databricks-Spark-Assoc exam?
Read Spark UI and driver/executor logs to find the real cause, then apply the matching fix: broadcast small dimension tables, salt skewed keys, avoid collect() on large data, and tune streaming triggers and Delta layout. Getting the root cause right matters more than memorizing tuning flags.
How many questions are in this domain?
This page lists all 41 Troubleshooting and Tuning DataFrame Apps questions in the Databricks-Spark-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Troubleshooting and Tuning DataFrame Apps questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
databricks-spark-developer-associate DATABRICKS-SPARK-DEVELOPER-ASSOCIATE troubleshooting tuning dataframe apps Practice Questions