Databricks-ML-Pro Model Development Practice Question
A data scientist is using MLflow on Databricks to tune a scikit-learn GradientBoostingRegressor with Hyperopt. They configure fmin with max_evals=50, but notice that runs appear in the experiment without parameters or metrics logged, and the best model cannot be reproduced. They want to ensure every trial is fully tracked. Which change should they make?
⚠ Common exam trap
The trap here is assuming that Hyperopt automatically creates MLflow runs for each trial, when in fact the objective function must manage the run context.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Wrap the objective function's training and evaluation code in `with mlflow.start_run():` (or use `mlflow.start_run(nested=True)` inside the objective) and log params/metrics explicitly.
Hyperopt's fmin does not manage MLflow runs; each trial must explicitly start a run to log parameters and metrics. Wrapping the objective in mlflow.start_run ensures isolated tracking per trial, which is essential for reproducibility and comparison. Other options confuse search breadth, autologging, or distributed execution with run lifecycle management.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use `SparkTrials` instead of `Trials` so that each trial is logged as a separate MLflow run.
Why it's wrong here
SparkTrials distributes trials across Spark workers but does not automatically create MLflow runs for each trial. It still relies on the objective function to manage MLflow logging. Switching to SparkTrials without adding a run context would still result in missing per-trial tracking. Distribution and tracking are separate concerns.
- ✗
Call `mlflow.autolog()` before Hyperopt and set `nested=True` in fmin.
Why it's wrong here
While mlflow.autolog() can capture scikit-learn training details, it does not automatically create a new run for each Hyperopt trial. Autologging logs to the currently active run. Without an explicit run context per trial, all logs may attach to a single parent run or be lost. fmin does not accept a nested parameter to control MLflow run nesting.
- ✓
Wrap the objective function's training and evaluation code in `with mlflow.start_run():` (or use `mlflow.start_run(nested=True)` inside the objective) and log params/metrics explicitly.
Why this is correct
Hyperopt's fmin does not automatically create an MLflow run for each trial. Without an active run context, calls to mlflow.log_param or mlflow.log_metric attach to the parent run or fail silently, leaving trials untracked. Starting a run inside the objective ensures each trial gets its own run with parameters and metrics, enabling reproducibility and comparison across trials.
- ✗
Increase `max_evals` to at least 200 so that MLflow has enough data to create runs automatically.
Why it's wrong here
The number of evaluations affects search thoroughness, not run creation. MLflow does not create runs based on the number of trials. Increasing max_evals would simply produce more untracked trials, wasting compute. The missing runs are due to lack of an explicit run context in the objective function, not insufficient evaluations.
About these practice questions
Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.