Databricks-ML-Assoc Model Development Practice Question
A machine learning engineer is using MLflow to track experiments. They call mlflow.start_run() and then log a model with mlflow.sklearn.log_model(). After the run completes, they notice that the model artifact is stored in the run's artifact location, but the run's source version and git commit are not captured. They are running from a Databricks notebook with Git integration enabled. Which action will ensure that the Git commit hash and source version are automatically logged to the MLflow run?
⚠ Common exam trap
The trap here is thinking that manual tagging or environment variables are required, when MLflow's Git integration automatically captures the commit hash when a Git repository is detected.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Ensure that the notebook is attached to a Git repository and that the Git integration is properly configured so MLflow can detect the commit hash.
MLflow automatically logs Git metadata when the code runs from a Git repository. In Databricks, if the notebook is in a Git folder with integration enabled, MLflow detects the repository and records the commit hash and source version in the run. This is the standard behavior; manual environment variables or tags are unnecessary and not the automatic mechanism.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set the environment variable MLFLOW_GIT_COMMIT to the current commit hash before starting the run.
Why it's wrong here
MLflow does not use the MLFLOW_GIT_COMMIT environment variable to capture Git information. It automatically detects Git metadata from the repository if the working directory is a Git repository. Manually setting this environment variable will not populate the source version or git commit fields in the run; those fields are set by MLflow's internal tracking when it detects a Git repo.
- ✓
Ensure that the notebook is attached to a Git repository and that the Git integration is properly configured so MLflow can detect the commit hash.
Why this is correct
MLflow automatically captures the Git commit hash and source version when the code is executed from within a Git repository. In Databricks, if the notebook is part of a Git folder and the Git integration is enabled, MLflow will detect the repository and log the commit hash and source version to the run. This is the intended mechanism.
- ✗
Call mlflow.set_tag('mlflow.source.git.commit', commit_hash) after starting the run.
Why it's wrong here
While you can manually set tags, the mlflow.source.git.commit tag is automatically populated by MLflow if it detects a Git repository. Manually setting it would work, but it is not the automatic mechanism. The question asks for automatic logging, so the correct approach is to rely on MLflow's Git detection, not manual tagging.
- ✗
Use mlflow.log_artifact() to upload the .git directory to the run's artifacts.
Why it's wrong here
Uploading the .git directory does not populate the run's source version or git commit fields. MLflow captures Git metadata by reading the repository state at runtime, not by uploading the .git folder. This action would not achieve the desired automatic logging of the commit hash.
About these practice questions
Courseiva writes every Databricks-ML-Assoc question from scratch — 319 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.