Courseiva
ML Ops →mediumMultiple Choice

Databricks-ML-Pro ML Ops Practice Question

You maintain a Databricks ML pipeline that trains a model nightly and registers new versions in MLflow Model Registry. A downstream batch scoring job in another workspace loads the model by stage. Auditors require that every production scoring run can be traced back to the exact training data snapshot and code commit. Which approach best satisfies this requirement?

⚠ Common exam trap

The trap here is assuming that any logged metadata such as a description or parameter creates an auditable link, when only the run ID and logged dataset entities provide an immutable, queryable chain.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use mlflow.log_input with a Delta table dataset, log the Git commit as a run tag, and register the model version from that run so the run ID links model, data, and code.

The requirement is an immutable, queryable chain from a production model version to the exact training data snapshot and code revision. Logging the dataset via mlflow.log_input and the Git commit as a run tag, then registering from that run, creates that chain because the model version stores the run ID, and the run stores the dataset digest and code tag. Manual descriptions, parameters, signatures, or notebook revisions are mutable or incomplete and cannot satisfy an audit.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use mlflow.log_input with a Delta table dataset, log the Git commit as a run tag, and register the model version from that run so the run ID links model, data, and code.

    Why this is correct

    mlflow.log_input records a dataset entity with its source and digest on the run, and the run ID becomes the immutable link between the model version, the exact data snapshot, and the logged Git commit tag. Because MLflow automatically stores the run ID on the model version, auditors can traverse from a scored model back to the precise training inputs and code revision without relying on human-maintained metadata.

  • ✗

    Store the training data path in the model's signature and use the workspace notebook revision to identify the code commit.

    Why it's wrong here

    A model signature describes input and output schema, not the data source or its version, so it cannot identify a snapshot. Notebook revisions are workspace-local and are not captured automatically on the MLflow run, and they do not survive export to another workspace. This combination therefore fails to provide a portable, immutable record linking the production model to its training data and code, which is what the auditors require.

  • ✗

    Set the model version's description to the Git commit hash and rely on the registered model name to identify the data snapshot.

    Why it's wrong here

    Descriptions are free-text metadata and can be edited or omitted, so they are not a reliable audit trail. The registered model name identifies a model, not a specific training data snapshot, and it does not record the code commit in a machine-verifiable way. This approach also gives no guarantee that the downstream job actually used the same version, because stage transitions can occur independently of the description.

  • ✗

    Log the training run with mlflow.log_param for the Git commit and log the Delta table version as a tag, then register the model version from that run.

    Why it's wrong here

    Parameters and tags are useful, but they are mutable and not enforced. A tag can be overwritten and nothing prevents a user from registering a version from a run that lacks them. For a strict audit trail you need an immutable, queryable linkage between the model version, the run, and the data version, which this manual logging does not guarantee. It is a partial solution rather than a robust one.

About these practice questions

Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.