Databricks-ML-Pro ML Ops Practice Question
An ML engineer is instrumenting a Databricks job that trains and evaluates several models. They want each model's metrics, parameters, and artifacts grouped so that a downstream automated promotion step can compare candidates within the same experiment. Which MLflow practice best supports programmatic comparison across runs in this job?
⚠ Common exam trap
The trap here is assuming that MLflow model versions carry metric data that can be ranked, when metrics live on runs within an experiment.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Log all candidates into a single experiment, tagging each run consistently so the promotion step can filter and rank runs by metric within that experiment.
Grouping candidate runs in one experiment lets MLflow's search API rank and filter them by metric and tag in a single query, which is what an automated promotion step needs. Splitting experiments, overloading registry descriptions, or duplicating metrics elsewhere all break that integrated, queryable comparison.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Log all candidates as versions of a single registered model and compare them by reading each version's description field for metric values.
Why it's wrong here
The registry tracks model versions, not run metrics, and descriptions are free text that cannot be reliably queried or ranked. Storing metric values in descriptions loses structure and prevents programmatic comparison, so this approach does not support an automated promotion step that must evaluate candidates by metric.
- ✗
Write metrics to a Delta table outside MLflow and have the promotion step read that table, using MLflow only to store the model artifacts.
Why it's wrong here
Duplicating metrics outside MLflow splits the source of truth and requires custom code to keep the table and runs in sync. It discards MLflow's built-in search, comparison, and lineage, so the promotion step loses the integrated view of parameters, metrics, and artifacts that makes candidate comparison reliable.
- ✗
Create a separate experiment for each model candidate so that runs are isolated and cannot interfere with one another.
Why it's wrong here
Isolating each candidate in its own experiment fragments the comparison, because MLflow search and metric comparison operate within an experiment. A downstream promotion step would need to query many experiments and reconcile run contexts, which is more complex and error-prone than keeping candidates together where they can be ranked directly.
- ✓
Log all candidates into a single experiment, tagging each run consistently so the promotion step can filter and rank runs by metric within that experiment.
Why this is correct
Keeping candidates in one experiment lets the promotion step search runs by metric and filter with tags in a single query, producing a clean ranking of candidates. Consistent tags add the metadata needed to identify job, dataset, or model family, so automated comparison and selection are straightforward and auditable.
About these practice questions
One of 300 original Databricks-ML-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.