A production Databricks workflow involves a task that runs a notebook. The notebook takes 15 minutes to finish, but the workflow is set to timeout after 10 minutes. What happens?
The scheduler enforces the timeout by killing the task process. Once terminated, the job workflow marks the task as 'Failed' because it did not complete successfully within the defined limits. This ensures that the system does not waste time and money on jobs that are behaving abnormally.
Why this answer
When a task exceeds its configured 'timeout' value, the Databricks scheduler forcibly terminates the task execution. This is a deliberate safety measure to prevent runaway processes from consuming cluster resources indefinitely. In production, this highlights the necessity of monitoring execution times and setting appropriate timeouts that account for normal data volume fluctuations while catching truly stuck jobs that could impact cost and resource availability.
Exam trap
Candidates often assume that a workflow timeout will gracefully cancel the notebook and mark it as 'Canceled' or 'Timed Out', overlooking that Databricks specifically categorizes exceeded task timeouts as 'Failed'.