Courseiva

Databricks-Spark-Assoc Spark Architecture and Components Practice Question

A developer submits a Spark application to a Databricks cluster using spark-submit with deploy mode set to cluster. During execution, one of the worker nodes hosting a task fails and is lost by the cluster manager. Which Spark component is responsible for rescheduling the failed task on another available executor?

⚠ Common exam trap

The trap here is assuming that because the DAG Scheduler builds the execution graph, it also owns per-task retries, when in fact task-level fault tolerance belongs to the Task Scheduler.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The Task Scheduler

When an executor is lost, the driver's Task Scheduler detects the failed tasks and relaunches them on other executors, honoring spark.task.maxFailures before failing the stage. The DAG Scheduler only regenerates stages when shuffle map outputs are lost. The cluster manager merely allocates resources and reports executor loss, and Catalyst is a compile-time optimizer with no runtime task responsibility.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The Task Scheduler

    Why this is correct

    The Task Scheduler in the Spark driver monitors each task within a stage, receives status updates from executors, and when a task fails or its executor is lost, it marks the task as failed and relaunches it on another available executor, up to spark.task.maxFailures. This is exactly the component that handles task-level fault tolerance in the scenario.

  • ✗

    The Catalyst Optimizer

    Why it's wrong here

    The Catalyst Optimizer is a query-planning framework that transforms logical plans into optimized physical plans at analysis time. It operates before execution begins and has no runtime role in monitoring, retrying, or rescheduling tasks when a worker node fails, so it cannot address the failure described in the scenario.

  • ✗

    The Cluster Manager

    Why it's wrong here

    The cluster manager (such as YARN, Kubernetes, or Databricks' internal manager) allocates and revokes containers or pods for executors, but it has no visibility into individual Spark tasks. It notifies the driver when an executor is lost, yet the decision to rerun the specific failed task belongs to the Task Scheduler, not the cluster manager.

  • ✗

    The DAG Scheduler

    Why it's wrong here

    The DAG Scheduler builds the stage graph from the RDD lineage and submits stages as TaskSets to the Task Scheduler, but it does not directly track individual task failures or retry them. When a task fails, the DAG Scheduler is only involved if an entire stage must be recomputed due to a shuffle map output loss, not for simple task-level retries.

About these practice questions

One of 295 original Databricks-Spark-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-Spark-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-Spark-Assoc exam.