Databricks-ML-Assoc Databricks Machine Learning Practice Question
A machine learning engineer is orchestrating a multi-step training workflow on Databricks using Databricks Jobs. The workflow includes data preprocessing, model training, and evaluation. The engineer needs to ensure that the evaluation step runs only if the training step completes successfully, and that the preprocessing step runs first. Which feature of Databricks Jobs should be used to define these dependencies?
⚠ Common exam trap
Many candidates confuse job orchestration features like dependencies with cluster configuration or failure recovery mechanisms, which do not control task execution order.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Task dependencies using the depends_on field
Task dependencies via the depends_on field are the core mechanism in Databricks Jobs for orchestrating multi-task workflows. They allow you to specify that a task runs only after its upstream tasks complete successfully, enabling conditional execution. This directly satisfies the need to run preprocessing first, then training, and evaluation only if training succeeds. Other options address compute scaling, failure recovery, or version control, not workflow dependencies.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Repair run with task-level retries
Why it's wrong here
Repair run allows re-executing failed tasks with different parameters, and task-level retries automatically retry failed tasks. Neither feature defines the order of tasks or conditional execution based on success. They handle failure recovery, not workflow dependencies. Thus, they cannot enforce that evaluation runs only after training succeeds, making them incorrect for this scenario.
- ✗
Databricks Repos integration with Git
Why it's wrong here
Databricks Repos enables version control and collaboration by syncing notebooks and files with Git repositories. It does not provide task orchestration or dependency management within a job. While useful for source control, it cannot define the execution order or conditional logic required for the multi-step workflow, so it is not the correct choice.
- ✗
Job clusters with autoscaling enabled
Why it's wrong here
Autoscaling job clusters adjust compute resources based on workload but do not control the execution order or conditional execution of tasks. They are unrelated to defining dependencies between tasks. While autoscaling can improve performance, it does not ensure that evaluation runs only after successful training, so it fails to meet the scenario's requirement.
- ✓
Task dependencies using the depends_on field
Why this is correct
Task dependencies in Databricks Jobs allow you to specify that a task runs only after its upstream tasks succeed. By setting the depends_on field, you create a directed acyclic graph (DAG) that enforces the order: preprocessing first, then training, then evaluation. This ensures that the evaluation step only runs if training succeeds, meeting the requirement.
About these practice questions
One of 319 original Databricks-ML-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.