Databricks-ML-Assoc Model Development Practice Question
A data scientist is building a pipeline using MLflow. Which TWO of the following tasks should be performed to ensure experiment reproducibility and auditability?
⚠ Common exam trap
Candidates frequently overlook logging the environment (conda_env) or source code (git hash), thinking that logging metrics alone is sufficient for full reproducibility of a machine learning experiment.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Log the git commit hash using mlflow.set_tag().
Reproducibility in machine learning requires strict tracking of code versions, dependencies, and parameters. By logging the source code version (git hash) and the specific Python environment (conda.yaml or requirements.txt), the scientist ensures that any collaborator can recreate the exact training state later. These practices are fundamental to the Databricks ML lifecycle, ensuring that models can be retrained or audited for compliance during the deployment phase.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Log the git commit hash using mlflow.set_tag().
Why this is correct
Logging the git commit hash directly links the model to the exact state of the source code. This is a best practice for tracking changes over time, as it allows developers to revert to specific training configurations and understand the lineage of the model artifacts within the MLflow Tracking server.
- ✗
Manually copy the training data into the MLflow model folder.
Why it's wrong here
Storing raw training data directly inside the MLflow model folder is inefficient and non-scalable. MLflow is designed to store metadata, code, and environment configurations. Large datasets should be referenced by path or URI, maintaining separation between the model logic and the massive data files it was trained upon.
- ✓
Capture the execution environment using log_model(conda_env=...).
Why this is correct
Logging the environment dependencies ensures that when the model is loaded in a different inference environment, the correct library versions are present. This prevents runtime errors and unexpected behavior due to library version drift, which is critical for consistent model performance in production and staging environments.
- ✗
Disable the automatic logging feature to save memory.
Why it's wrong here
Disabling automatic logging defeats the purpose of using MLflow's robust tracking capabilities. Automatic logging is highly efficient and captures essential metrics and parameters without manual intervention. Disabling it increases the burden on the user to manually track every experiment, which introduces risks of human error and incomplete data.
- ✗
Use a global variable for all experiment parameters.
Why it's wrong here
Global variables are volatile and not persisted across sessions or restarts. They do not provide auditability or reproducibility because they are not recorded in the MLflow tracking store. Every experiment parameter should be explicitly logged via MLflow APIs to ensure persistence and traceability for all team members.
About these practice questions
Courseiva writes every Databricks-ML-Assoc question from scratch — 319 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.