Databricks-ML-Assoc Model Development Practice Question
A data scientist is using MLflow to track a hyperparameter tuning experiment with Spark MLlib's CrossValidator. They notice that each run in the MLflow UI shows only a single set of metrics, but they want to compare the performance of each hyperparameter combination across folds. What is the most effective way to log and visualize the per-combination and per-fold metrics in MLflow?
⚠ Common exam trap
The trap here is assuming that logging all metrics into a single run or using the step parameter can substitute for separate runs, when the MLflow UI's comparison capabilities fundamentally depend on runs as distinct entities.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Create a separate MLflow run for each hyperparameter combination, and within each run, log the average metric across folds as well as individual fold metrics using distinct metric names.
The MLflow UI organizes metrics and parameters by run, making runs the natural unit for comparing hyperparameter combinations. By creating a separate run for each combination and logging both aggregate and per-fold metrics with distinct names, the data scientist can leverage the UI's comparison tools, such as parallel coordinates and scatter plots, to identify the best performing combination and understand fold-level variability.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Create a separate MLflow run for each hyperparameter combination, and within each run, log the average metric across folds as well as individual fold metrics using distinct metric names.
Why this is correct
Creating a separate run per hyperparameter combination allows each combination to be compared as a distinct entity in the MLflow UI. Logging both the average and individual fold metrics, with unique names like 'fold_0_auc', provides detailed insight. This structure supports sorting, filtering, and parallel coordinates plots to identify the best combination, directly addressing the need to compare performance across folds.
- ✗
Use MLflow's log_batch API to log all metrics and parameters for each combination in a single batch, without creating separate runs.
Why it's wrong here
log_batch is an efficient way to log multiple metrics and parameters at once, but it still logs to a single run. Without separate runs per combination, the MLflow UI cannot easily compare combinations. The UI's comparison features rely on runs as the primary unit, so this approach would not provide the desired per-combination visualization.
- ✗
Log all metrics from all combinations into a single run, using metric names that encode the hyperparameters and fold numbers.
Why it's wrong here
While encoding hyperparameters and fold numbers into metric names can store all data in one run, it makes comparison difficult. The MLflow UI is designed to compare across runs, not within a single run's many metrics. This approach leads to a cluttered run and prevents effective use of parallel coordinates or scatter plots, which rely on runs as units of comparison.
- ✗
Use mlflow.log_metric with a step parameter for each fold, and log the hyperparameters as a JSON string in a single parameter.
Why it's wrong here
Using the step parameter with log_metric is intended for time-series metrics, such as loss over epochs. For cross-validation, folds are not sequential steps in the same run. Logging hyperparameters as a JSON string in one parameter makes them hard to filter and compare in the MLflow UI. This approach does not effectively enable per-combination and per-fold visualization.
About these practice questions
Courseiva writes every Databricks-ML-Assoc question from scratch — 319 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.