Courseiva
Model Development →mediumMultiple Select

Databricks-ML-Assoc Model Development Practice Question

Which TWO actions should be taken to ensure reproducibility of a Databricks ML model experiment?

⚠ Common exam trap

Candidates often forget that code and environment versioning are insufficient without capturing the exact historical state of the underlying data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Log the Git commit hash along with experiment metadata in MLflow.

Reproducibility in Databricks is achieved by capturing the state of the code, the environment, and the data. By utilizing Git integration and MLflow tracking, data scientists can point to the exact version of the code and environment used for an experiment. Using versioned Delta tables ensures that the data state can also be reconstructed, preventing 'data drift' from invalidating the results of historical experiments as underlying data tables change.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Hardcode the data file paths to local desktop directories.

    Why it's wrong here

    Hardcoded local paths break reproducibility immediately, as other users or production systems will not have access to those files. Data should always be referenced via cloud storage paths or Delta table versions, which are accessible to the whole cluster and can be versioned appropriately.

  • ✓

    Log the Git commit hash along with experiment metadata in MLflow.

    Why this is correct

    Logging the Git commit hash allows developers to link the model artifact to the exact version of the source code that produced it. This is essential for auditing and reproducing experiments months later, as it provides a clear record of the code changes that led to specific model performance.

  • ✗

    Only log the model artifact and ignore the environment dependencies.

    Why it's wrong here

    Ignoring dependencies makes it impossible to recreate the runtime environment. If the library versions used during training are different from those used during reproduction, the model might produce different results or fail entirely, undermining the reproducibility efforts and making it difficult to maintain the model long-term.

  • ✓

    Utilize Delta Time Travel to query the data as it existed during training.

    Why this is correct

    Delta Lake's Time Travel feature allows users to query data at a specific point in time using versioning or timestamps. This ensures that the exact same dataset used during training can be accessed again, which is critical for reproducing experiments even after the underlying data table has been updated.

  • ✗

    Manually delete all previous MLflow runs to keep the tracking server clean.

    Why it's wrong here

    Deleting history prevents the ability to compare current experiments against past performance. Reproducibility relies on having access to the full history of experiments, including failures. Keeping the tracking server clean by deleting data is counter-productive to the goals of experimentation and professional model lifecycle management.

About these practice questions

One of 319 original Databricks-ML-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.