Courseiva

Databricks-Spark-Assoc · topic practice

Troubleshooting and Tuning DataFrame Apps practice questions

This domain covers diagnosing and fixing performance and failure problems in Spark DataFrame and Structured Streaming jobs on Databricks. Questions present logs, symptoms, or code and ask you to identify root causes like executor loss, driver OutOfMemoryError, growing streaming batch durations, or data skew, then select the correct mitigation using Spark UI, Delta Lake, and tuning techniques.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Troubleshooting and Tuning DataFrame Apps

What the exam tests

What to know about Troubleshooting and Tuning DataFrame Apps

Read Spark UI and driver/executor logs to find the real cause, then apply the matching fix: broadcast small dimension tables, salt skewed keys, avoid collect() on large data, and tune streaming triggers and Delta layout. Getting the root cause right matters more than memorizing tuning flags.

Diagnosing ExecutorLostFailure from logs and Spark UI stage/task metrics

Fixing driver OutOfMemoryError from collect() by using write, take, or broadcast

Tuning Structured Streaming batch duration with trigger, watermark, and Delta optimization

Mitigating join data skew via broadcast hints, salting, or AQE skew join

Why learners struggle

Why Troubleshooting and Tuning DataFrame Apps questions are commonly missed

RAM questions are commonly missed because learners confuse physical form factors (DIMM vs SO-DIMM) and fail to distinguish between memory speed (MHz) and latency (CL).

  • ·DIMM vs SO-DIMM — desktop vs laptop form factor confusion
  • ·DDR3 vs DDR4 vs DDR5 — notch position and voltage differences
  • ·MHz vs CL — speed vs latency trade-offs in performance
  • ·Single-channel vs dual-channel — bandwidth impact misconception
  • ·ECC vs non-ECC — error correction support in servers vs desktops
  • ·32-bit vs 64-bit — maximum addressable RAM limit

Watch out for

Common Troubleshooting and Tuning DataFrame Apps exam traps

  • ▸Assuming ExecutorLostFailure is always a code bug, ignoring memory, shuffle spill, or node loss causes shown in logs
  • ▸Calling collect() on large DataFrames and expecting driver memory to scale, instead of writing results or aggregating first
  • ▸Treating slow streaming batches as a trigger problem while ignoring state growth, small files, or unoptimized Delta reads

Practice set

Troubleshooting and Tuning DataFrame Apps questions

20 questions · select your answer, then reveal the explanation

A Spark SQL query is running slowly because of a large broadcast join that keeps failing. What is the most likely reason for this failure?

Which TWO actions can help prevent data skew in a join operation?

Which action allows Spark to perform 'predicate pushdown' when reading data from a Parquet file?

Your Spark job is failing with an OutOfMemoryError (OOM) during a shuffle-heavy transformation. Which TWO configurations should you investigate to troubleshoot this memory pressure?

Which action should be prioritized when observing that a DataFrame job is performing frequent 'spilling to disk' during a sort-merge join?

Which THREE factors can lead to 'Data Skew' in a Spark DataFrame application?

What is the primary function of the Spark UI 'SQL' tab when troubleshooting a poorly performing DataFrame application?

You are optimizing a Spark SQL query that performs a sort on a large dataset. The job is currently failing because the sort operation exceeds the memory of the executors. Which configuration should be adjusted to allow the sort to spill to disk?

Which THREE actions help mitigate the 'small files problem' in Spark?

A developer runs a PySpark job on Databricks that joins a 400 GB Delta table with a 2 GB Parquet table. The job is slow and the Spark UI shows that the 2 GB table is being shuffled across the network for every join stage. The developer wants to avoid this shuffle. Which change should be made to the join?

A data engineer notices that a Spark job on Databricks is running slowly and the Spark UI shows that many tasks are spilling to disk. The job performs a groupBy on a high-cardinality key. Which configuration change is most likely to reduce spilling and improve performance?

A PySpark job on Databricks joins a 500 GB Delta table with a 2 GB Delta table using a broadcast hash join. The job runs out of memory on the executors. You have already increased executor memory, but the issue persists. What should you do to resolve the memory pressure while keeping the join efficient?

A Databricks job processing a Delta table with 50,000 small files is slow because each task opens many files. The developer wants to reduce the number of files read per task without rewriting the table. Which two actions should be taken? (Choose two.)

A Spark application on Databricks reads a Parquet dataset with 10,000 small files, each averaging 10 KB. The job performs a simple aggregation and takes hours to complete. The developer wants to improve performance without changing the aggregation logic. Which action is most effective?

A PySpark job on Databricks repeatedly reads the same 200 GB Delta table with different filters across several transformations, and the driver has 64 GB of RAM. You want to avoid recomputing the full table scan for each action while keeping memory pressure manageable. Which approach should you use?

A Spark DataFrame job on Databricks reads a large Parquet dataset and performs a groupBy aggregation. The job is slow, and the Spark UI shows that most tasks are reading the same partitions repeatedly. Which configuration change would most directly improve performance?

A PySpark job on Databricks processes a 200 GB Delta table using df.groupBy('customer_id').agg(sum('amount')). The job runs for 3 hours, but Spark UI shows that 95% of tasks finish in 2 minutes while a few tasks run for over 2 hours. The shuffle read size is evenly distributed. Which Spark configuration change is most likely to resolve the long-running tasks?

A PySpark DataFrame job on Databricks repeatedly runs out of memory during a groupBy aggregation on a 500 GB Parquet dataset. The Spark UI shows that the task with the largest shuffle read processes 20 times more records than the median task, and the job fails with java.lang.OutOfMemoryError on the executor. You want to resolve the failure without changing the aggregation logic. Which action should you take?

A nightly PySpark job joins a 900 GB transactions DataFrame with a 3 GB customers DataFrame. The join key is heavily skewed: one customer ID accounts for roughly 35 percent of all transactions, and a single task runs for hours while the rest of the stage finishes quickly. You cannot change the source data or the join key. Which approach most directly resolves the skew?

A Spark Structured Streaming job on Databricks reads from a Kafka topic and writes to a Delta table using foreachBatch. The job processes 10,000 records per second, but after several hours, the batch duration steadily increases and eventually the job falls behind. The Spark UI shows that the number of active batches is growing and shuffle spill is high. Which action should you take to stabilize the streaming job?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Troubleshooting and Tuning DataFrame Apps sessions

Start a Troubleshooting and Tuning DataFrame Apps only practice session

Every question in these sessions is drawn from the Troubleshooting and Tuning DataFrame Apps domain — nothing else.

Related practice questions

Related Databricks-Spark-Assoc topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-Spark-Assoc exam test about Troubleshooting and Tuning DataFrame Apps?
Read Spark UI and driver/executor logs to find the real cause, then apply the matching fix: broadcast small dimension tables, salt skewed keys, avoid collect() on large data, and tune streaming triggers and Delta layout. Getting the root cause right matters more than memorizing tuning flags.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Troubleshooting and Tuning DataFrame Apps questions in a focused session?
Yes — the session launcher on this page draws every question from the Troubleshooting and Tuning DataFrame Apps domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-Spark-Assoc topics?
Use the topic links above to move to related areas, or go back to the Databricks-Spark-Assoc question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-Spark-Assoc exam covers. They are not copied from any real exam or dump site.