Databricks-ML-Pro ML Ops Practice Question
You are auditing a Databricks environment to ensure compliance. Which TWO actions ensure the highest level of model lineage and reproducibility for models registered in MLflow?
⚠ Common exam trap
Candidates often focus on model accuracy metrics, ignoring that compliance and reproducibility require linking the model to specific code versions (Git) and exact dependency environments (Conda/pip).
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Include the Git hash in the log_model metadata.
Linking models to specific Git commits and tracking the environment dependencies via Conda or pip requirements files are essential for reproducibility. These steps ensure that when a model is redeployed, the exact code version and environment can be recreated, fulfilling the core MLOps requirement of auditability and consistent model behavior across different computing environments.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Include the Git hash in the log_model metadata.
Why this is correct
Recording the Git hash during logging directly links the model artifact to the exact codebase used for training. This enables developers to checkout the precise state of the repository at the time of training, ensuring that lineage is preserved and debugging becomes straightforward.
- ✗
Use random seeds for all ML libraries.
Why it's wrong here
While random seeds are necessary for reproducibility of individual runs, they do not manage lineage or environment configurations. Relying solely on seeds is insufficient for MLOps, as it ignores the dependency management and code-versioning aspects required to track how a specific model was created.
- ✓
Log the environment configuration (pip/conda) with the model.
Why this is correct
Logging the environment file ensures that the exact library versions used during training are captured. This is critical for preventing 'dependency hell' in production, where subtle version differences between training and inference environments can cause unexpected model degradation or failure to load correctly.
- ✗
Store all training data as CSV files in the workspace.
Why it's wrong here
Storing training data as CSVs in the workspace is not scalable or secure. MLOps best practices suggest using Delta tables for versioning data through time travel, as CSVs do not provide the transactional integrity or efficient lineage tracking required for enterprise-scale machine learning operations.
- ✗
Manually copy artifacts to a shared folder.
Why it's wrong here
Manual file management is error-prone and lacks auditability. It bypasses the MLflow Model Registry's built-in tracking and versioning mechanisms, rendering lineage tracing impossible. MLOps requires automated, centralized storage solutions to ensure that artifacts are managed consistently throughout their lifecycle without human intervention or oversight.
About these practice questions
Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.