Databricks-DE-Pro Debugging and Deploying Practice Question
A data engineer is debugging a Databricks job that fails with a `SparkException: Job aborted due to stage failure` in production. They need to identify the root cause. Which two actions should they take to gather relevant diagnostic information? (Choose two.)
⚠ Common exam trap
The trap here is overlooking the Spark UI in favor of only logs, missing critical performance metrics that explain why the stage failed.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Review the Spark driver logs for the failed job run in the Databricks Jobs UI.
To debug a Spark stage failure, the driver logs provide the exception details, and the Spark UI offers performance metrics that reveal issues like skew or spills. Together, they help identify whether the failure is due to data, code, or resource problems. The other actions are either for different error types or less direct for this specific failure.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Review the Spark driver logs for the failed job run in the Databricks Jobs UI.
Why this is correct
The Spark driver logs contain detailed error messages, including the stage that failed and the exception stack trace. This is the first place to look for root cause. The logs are accessible from the job run's detail page and provide context about the failure, such as data issues or resource problems.
- ✗
Run the job locally with a small sample of data to reproduce the error.
Why it's wrong here
Running locally with a small sample may not reproduce production-scale issues like data skew or memory pressure. The error is likely due to scale or data distribution, which a small sample would mask. While useful for logic errors, it is not the best first step for a stage failure in production.
- ✗
Examine the event log for the job cluster to see if nodes were terminated unexpectedly.
Why it's wrong here
The event log can show cluster events like node termination, but a stage failure is usually caused by an exception in the Spark job, not by node termination. While node loss can cause failures, it would typically result in a different error, such as `SparkException: Lost executor`. This action is secondary.
- ✗
Check the cluster's init script logs to see if a library installation failed.
Why it's wrong here
Init script logs are useful for cluster startup issues, but the error is a Spark stage failure, which occurs during job execution, not initialization. While library issues can cause failures, they typically manifest as different errors, such as `LibraryInstallationFailed`. This action is less relevant for a stage failure.
- ✓
Inspect the Spark UI for the failed stage to identify skewed tasks or spills.
Why this is correct
The Spark UI provides per-stage metrics, including task durations, input sizes, and spill statistics. Skewed tasks or excessive spills often cause stage failures due to out-of-memory errors. Analyzing the Spark UI helps pinpoint performance bottlenecks or data distribution issues that lead to the failure.
About these practice questions
One of 267 original Databricks-DE-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-DE-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-DE-Pro exam.