Databricks-ML-Assoc ML Workflows Practice Question
A machine learning engineer is building a training pipeline in Databricks. They want each run to record the exact Git commit hash, the versions of scikit-learn and MLflow used, and the input data path so the run can be reproduced later. Which MLflow tracking capability should they use to capture this information with the least custom code?
⚠ Common exam trap
The trap here is assuming that any run metadata must be logged manually as a parameter or artifact, rather than using MLflow's built-in source and environment tracking plus tags.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
mlflow.set_tags() combined with MLflow's automatic logging of the source Git commit and environment.
MLflow Tracking automatically records source information, including the Git commit when the run executes from a repository, and the environment details. Tags provide a searchable place for custom metadata such as the input data path. Together they deliver reproducibility context without custom parsing code, unlike parameters, artifacts, or registry entries that are meant for different purposes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
mlflow.register_model() to create a model version that stores the Git commit hash and data path in its description.
Why it's wrong here
register_model publishes a model to the Model Registry and is intended for lifecycle management, not for capturing per-run environment metadata. The registry version description is not designed to store Git commits or data paths automatically, and it would only cover runs that produce a registered model, leaving training experiments without that provenance information.
- ✗
mlflow.log_artifact() to upload a text file containing the Git commit hash, library versions, and data path.
Why it's wrong here
log_artifact can store a metadata file, but it does not automatically capture the Git commit or environment and it hides the information inside a file that must be downloaded to inspect. This requires custom code to write the file and parse it later, and the values are not searchable in the MLflow run list, so it is a poor fit for quick reproducibility checks.
- ✓
mlflow.set_tags() combined with MLflow's automatic logging of the source Git commit and environment.
Why this is correct
MLflow automatically captures the source Git commit hash and the active conda environment when a run starts, and custom metadata such as a data path can be attached with set_tags. Tags are searchable and displayed separately from parameters and metrics, so the engineer gets reproducibility data with almost no code while keeping the run UI clean.
- ✗
mlflow.log_params() to log the Git commit hash, library versions, and data path as run parameters.
Why it's wrong here
log_params records key-value pairs that describe model hyperparameters or configuration, but Git commit hashes, library versions, and data paths are metadata about the environment and source, not tunable parameters. Stuffing them into parameters pollutes the parameter view, makes comparisons noisy, and does not leverage MLflow's dedicated tags and automatic source tracking for reproducibility.
About these practice questions
Courseiva writes every Databricks-ML-Assoc question from scratch — 319 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.