Databricks-ML-Assoc ML Workflows Practice Question
An ML engineer runs a hyperparameter tuning job on Databricks using Hyperopt with SparkTrials. The objective function trains a model and returns a validation metric. The engineer notices that each trial logs to the same MLflow run, making it impossible to compare trials. What should the engineer do to ensure each trial appears as a separate MLflow run?
⚠ Common exam trap
The trap here is thinking that tagging or renaming metrics separates trials, when only creating a nested run per trial actually produces distinct MLflow runs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Call mlflow.start_run(nested=True) at the beginning of the objective function.
With Hyperopt and SparkTrials on Databricks, the tuning job creates a parent MLflow run. To capture each trial separately, the objective function must open a nested run using mlflow.start_run(nested=True). This produces child runs per trial, each with its own parameters and metrics, which can be compared in the MLflow experiment UI. Tags, metric names, or experiment changes do not create distinct runs.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use mlflow.set_tag to assign a unique trial ID tag before logging metrics.
Why it's wrong here
Tags are metadata attached to a run, not a way to create new runs. Setting a unique trial ID tag on the same run would overwrite the tag on each trial, and all metrics would still accumulate in one run. While tags help filter runs, they cannot substitute for separate run creation, so trial comparison would remain impossible.
- ✗
Pass a unique run_name argument to mlflow.log_metric for each trial.
Why it's wrong here
mlflow.log_metric does not accept a run_name parameter; it logs to the currently active run. Without starting a new run per trial, all metrics still land in the same run. Renaming metrics or runs after the fact does not separate trial data. The engineer must create a distinct run context for each trial, typically via nested runs.
- ✓
Call mlflow.start_run(nested=True) at the beginning of the objective function.
Why this is correct
SparkTrials creates a parent run for the tuning job and, when the objective function opens a nested run with mlflow.start_run(nested=True), each trial is logged as a child run under that parent. This yields one run per trial with its own parameters and metrics, enabling comparison in the MLflow UI. Databricks documents this pattern for distributed hyperparameter tuning.
- ✗
Set the environment variable MLFLOW_EXPERIMENT_ID to a unique value inside the objective function.
Why it's wrong here
Changing the experiment ID inside the objective function would move all trials into different experiments, not separate runs within one experiment. This fragments the experiment and still does not create per-trial runs. The correct mechanism is to start a nested run per trial, which SparkTrials supports automatically when the objective function calls mlflow.start_run with nested=True.
About these practice questions
Courseiva writes every Databricks-ML-Assoc question from scratch — 319 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.