Databricks-ML-Pro Model Development Practice Question
A data scientist wants to record the exact library dependencies and a code snapshot alongside a model so that a reviewer can later restore the same environment and reproduce the training run. They are logging with MLflow on Databricks. Which practice best satisfies this requirement?
⚠ Common exam trap
The trap here is believing that autolog or artifact storage location alone captures the dependency environment, when the environment must be explicitly pinned.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Log the model with `mlflow.sklearn.log_model` and pass pinned `pip_requirements` plus attach the source notebook as an artifact.
Reproducibility requires capturing two things: the environment and the code. Pinning `pip_requirements` when logging the model embeds a version-exact dependency manifest that MLflow restoration consumes. Attaching the source notebook records the training logic as it existed at that moment. Storing artifacts elsewhere or relying on autolog alone leaves one of these gaps, so the run cannot be faithfully recreated later.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Call `mlflow.log_artifact` on a generated `requirements.txt` and let autolog capture parameters only.
Why it's wrong here
Logging a requirements file as a generic artifact records the list but does not wire it into the model's environment metadata, so restoration tooling will not consume it automatically. Autolog records parameters and metrics, not a full reproducible environment, leaving the reviewer to reconstruct dependencies manually and risking version mismatches.
- ✓
Log the model with `mlflow.sklearn.log_model` and pass pinned `pip_requirements` plus attach the source notebook as an artifact.
Why this is correct
Pinning `pip_requirements` writes an exact dependency manifest into the model's MLmodel metadata, so restoration tooling can rebuild the environment. Attaching the source notebook preserves the code snapshot. Together these give the reviewer both the environment and the code needed to reproduce the training run faithfully, which is precisely what was requested.
- ✗
Log the model with `mlflow.sklearn.log_model` and rely on the cluster's installed libraries at load time.
Why it's wrong here
Relying on whatever is installed on the loading cluster makes reproduction fragile because package versions can drift between runs. Even though the model artifact is captured, there is no recorded environment specification, so a reviewer cannot guarantee the same behavior. This does not satisfy the requirement for explicit dependency capture.
- ✗
Use `mlflow.autolog()` and set the experiment's artifact location to a Unity Catalog volume.
Why it's wrong here
Autolog captures parameters, metrics, and the model artifact automatically, and the artifact location controls where files are stored, but neither action records the package environment spec. Storage location is about durability and governance, not reproducibility, so the reviewer still cannot rebuild the exact dependency set from the run alone.
About these practice questions
This Databricks-ML-Pro question is part of Courseiva's 300-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.