Courseiva
Databricks Machine Learning →mediumMultiple Choice

Databricks-ML-Assoc Databricks Machine Learning Practice Question

A machine learning engineer is using MLflow Tracking on Databricks to compare multiple runs. They want to programmatically retrieve the best run based on a metric called 'rmse' from an experiment. Which MLflow API call should they use?

⚠ Common exam trap

The trap here is assuming that `mlflow.get_run()` or `mlflow.list_run_infos()` can directly sort or filter runs by metrics, when they either require a known run ID or lack metric data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

`mlflow.search_runs()` with `order_by=['metrics.rmse ASC']`

`mlflow.search_runs()` is designed to query runs within an experiment and can order results by metrics. By specifying `order_by=['metrics.rmse ASC']`, the run with the lowest RMSE appears first. This is efficient and programmatic, unlike manually iterating runs or using functions that retrieve single runs or experiment metadata.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    `mlflow.get_experiment_by_name()` and then filter runs by metric

    Why it's wrong here

    `mlflow.get_experiment_by_name()` returns the experiment object, not the runs. It does not provide a way to filter or sort runs by metrics. Additional calls are needed to list runs and their metrics, making this option incomplete and not a direct solution.

  • ✗

    `mlflow.get_run()` with the run ID of the best run

    Why it's wrong here

    `mlflow.get_run()` retrieves details for a single run when you already know its run ID. It does not search or compare runs, so it cannot identify the best run based on a metric. It is useful after you have determined the best run ID through other means.

  • ✗

    `mlflow.list_run_infos()` and manually iterate to find the minimum 'rmse'

    Why it's wrong here

    `mlflow.list_run_infos()` returns metadata about runs but does not include metrics. You would need to call `mlflow.get_run()` for each run to retrieve metrics, which is inefficient and requires manual iteration. This approach is not the recommended programmatic way to find the best run.

  • ✓

    `mlflow.search_runs()` with `order_by=['metrics.rmse ASC']`

    Why this is correct

    `mlflow.search_runs()` returns a pandas DataFrame of runs and supports ordering by metrics. Using `order_by=['metrics.rmse ASC']` sorts runs by the 'rmse' metric in ascending order, so the first row is the run with the lowest RMSE. This is the correct programmatic way to retrieve the best run.

About these practice questions

This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.