Courseiva

Databricks-DE-Assoc · topic practice

Troubleshooting, Monitoring, and Optimization practice questions

This domain covers diagnosing and fixing Databricks workloads: cluster selection, job failures, Delta table maintenance, and Spark performance tuning. Questions present a failure scenario (OOM, too many files, slow job) and ask which Databricks feature, command, or configuration resolves it, so you must map symptoms to the right remedy.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Troubleshooting, Monitoring, and Optimization

What the exam tests

What to know about Troubleshooting, Monitoring, and Optimization

You must diagnose a failing or slow Databricks job and pick the correct fix: cluster type, Delta maintenance command, or Spark join strategy. The most important thing is matching the symptom to the right tool, especially VACUUM versus OPTIMIZE and Job versus All-Purpose clusters.

Choosing Job Clusters versus All-Purpose Clusters based on cost and reuse

Using OPTIMIZE, VACUUM, and partitioning to reduce excessive Delta metadata operations

Mitigating join OOM with broadcast hints, partitioning, or larger cluster memory

Recovering storage via VACUUM with the correct retention interval on Delta tables

Watch out for

Common Troubleshooting, Monitoring, and Optimization exam traps

  • ▸Assuming VACUUM deletes files immediately; it respects the retention threshold and fails if below the safe default without disabling the check.
  • ▸Confusing OPTIMIZE (compacts small files) with VACUUM (removes unreferenced files); they solve different problems.
  • ▸Believing an All-Purpose Cluster is always better for jobs; Job Clusters are cheaper and isolated but cannot be shared interactively.

Practice set

Troubleshooting, Monitoring, and Optimization questions

20 questions · select your answer, then reveal the explanation

An engineer has discovered that a specific join operation is extremely slow due to severe data skew on the join key. Which strategy should they use to mitigate this?

Refer to the exhibit. An engineer observes that running VACUUM on a Delta table with a retention period of 0 hours results in an error. Why is this configuration likely failing?

Exhibit

{
  "cluster_name": "prod-cluster",
  "spark_conf": {
    "spark.databricks.delta.retentionDurationCheck.enabled": "false"
  }
}

A data engineer is troubleshooting a Databricks job that occasionally fails with a 'java.lang.OutOfMemoryError: GC overhead limit exceeded' on the driver node. The job performs a large collect() operation on a DataFrame and then processes the results locally. Which TWO actions should the engineer take to resolve this issue? (Choose two.)

A data engineer notices that a Databricks job writing to a Delta table is taking much longer than expected. The job uses a MERGE INTO statement that updates a large target table from a small source DataFrame. The engineer runs DESCRIBE HISTORY on the target table and sees that each MERGE operation rewrites a large number of files. Which optimization technique should the engineer apply to improve the performance of the MERGE operation?

A data engineer runs a Structured Streaming job that reads from a Delta table and writes to another Delta table. The job is configured with a 5-minute trigger interval. After several hours, the engineer notices that the write latency has increased significantly and the streaming query is processing each micro-batch slower than before. The source Delta table is frequently updated with small append operations, and the target table is not optimized. Which action should the engineer take to improve the streaming job's performance?

A data engineer is troubleshooting a Databricks job that fails with an OutOfMemoryError on the driver. The job collects a large DataFrame to the driver using collect() before writing results. The engineer wants to resolve the failure with minimal code changes. Which action should be taken?

A data engineer is troubleshooting a Databricks job that fails with an error indicating that the driver node has run out of memory. The job performs a large collect() operation on a DataFrame. Which action should the engineer take to resolve this issue?

A data engineer is troubleshooting a Databricks job that occasionally fails with a 'FileNotFoundException' when reading from a Delta table. The table is updated by a separate streaming job that performs frequent merges and deletes. The engineer suspects that the issue is related to data retention and file cleanup. Which configuration should the engineer adjust to ensure that the reading job can still access the required historical versions of the table?

A data engineer is optimizing a Databricks job that performs a large join between two Delta tables. The engineer notices that the job is spilling data to disk and taking a long time. Which TWO actions should the engineer take to improve performance? (Choose two.)

A data engineer has a Databricks job that writes to a Delta table. The job sometimes fails due to concurrent write conflicts. The engineer wants to ensure that the job can automatically retry on such conflicts. Which feature should the engineer use?

A data engineer is troubleshooting a Databricks job that reads from a Delta table and writes to another Delta table. The job suddenly starts failing with a `FileNotFoundException` during read operations. The engineer verifies that the source table exists and is accessible. Which action should the engineer take to resolve this error?

A Databricks job is failing due to 'Out of Memory' (OOM) errors during a join operation on two large datasets. Which TWO actions could help mitigate this issue?

Which Databricks feature should a data engineer use to view the execution plan, including information about the physical operators and data lineage, to troubleshoot a slow-running SQL query?

A streaming job using Structured Streaming is lagging significantly behind the 'current time'. Which THREE of the following could be the root cause of this processing latency?

An engineer has a large table that is frequently joined with small dimension tables. To optimize this, which optimization technique should be applied to the join operation?

Which Databricks command should you use to recover storage space by removing files that are no longer referenced by a Delta table and are older than the retention period?

Which tool in Databricks provides real-time monitoring of cluster resource usage, including CPU, memory, and network throughput, for an active job?

When configuring a Databricks Job, which TWO factors most directly influence the choice between a 'Job Cluster' and an 'All-Purpose Cluster'?

Which Spark configuration property can be used to enable Adaptive Query Execution (AQE) in Databricks?

Which THREE strategies are recommended to improve the performance of reading from a Delta table in Databricks?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Troubleshooting, Monitoring, and Optimization sessions

Start a Troubleshooting, Monitoring, and Optimization only practice session

Every question in these sessions is drawn from the Troubleshooting, Monitoring, and Optimization domain — nothing else.

Related practice questions

Related Databricks-DE-Assoc topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-DE-Assoc exam test about Troubleshooting, Monitoring, and Optimization?
You must diagnose a failing or slow Databricks job and pick the correct fix: cluster type, Delta maintenance command, or Spark join strategy. The most important thing is matching the symptom to the right tool, especially VACUUM versus OPTIMIZE and Job versus All-Purpose clusters.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Troubleshooting, Monitoring, and Optimization questions in a focused session?
Yes — the session launcher on this page draws every question from the Troubleshooting, Monitoring, and Optimization domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-DE-Assoc topics?
Use the topic links above to move to related areas, or go back to the Databricks-DE-Assoc question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-DE-Assoc exam covers. They are not copied from any real exam or dump site.