Databricks-ML-Assoc Databricks Machine Learning Practice Question
A machine learning engineer needs to track hyperparameter tuning runs and log artifacts using MLflow inside a Databricks Notebook. Which approach should be used to ensure runs are automatically nested under a parent run?
⚠ Common exam trap
Candidates often run multiple mlflow.start_run() calls without nesting, leading to cluttered dashboards. They fail to use the nested parameter, which is specifically designed for hierarchical hyperparameter tuning organization.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Utilize mlflow.start_run(nested=True) inside the hyperparameter optimization loop while a parent run is active.
Using `mlflow.start_run()` with a `nested=True` parameter ensures that child runs are correctly organized beneath an active parent run within the MLflow tracking server. This hierarchical structuring is essential for organizing complex hyperparameter optimization sweeps, enabling data scientists to easily compare child performance metrics against the overarching parent experiment trial in the Databricks UI.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Call mlflow.start_run() inside a loop without any arguments while an active run exists.
Why it's wrong here
Calling mlflow.start_run() without arguments or the nested parameter while another run is active will terminate the existing parent run and start a new independent top-level run, rather than establishing the required hierarchical parent-child relationship needed for hyperparameter sweeps.
- ✓
Utilize mlflow.start_run(nested=True) inside the hyperparameter optimization loop while a parent run is active.
Why this is correct
The nested=True argument tells MLflow to create a child run under the currently active parent run context. This maintains proper organization in the tracking UI, allowing clean grouping of multiple training iterations associated with a single overarching optimization experiment.
- ✗
Set the MLFLOW_PARENT_RUN_ID environment variable manually before invoking mlflow.start_run().
Why it's wrong here
Manually altering environment variables for run IDs is not the standard programmatic API provided by MLflow. Using environment variables directly can lead to race conditions or incorrect context propagation across distributed worker nodes in a Databricks cluster.
- ✗
Configure the MLflow client to use a hierarchical tracking URI scheme before logging parameters.
Why it's wrong here
MLflow tracking URIs specify the location of the backend store and artifact store, such as databricks or a remote server. They do not control run hierarchy or nesting behavior, which is strictly managed via the active run context stack in code.
About these practice questions
This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.