Databricks-ML-Assoc Model Development Practice Question
When logging a model using MLflow in Databricks, which component is required to capture the environment dependencies so that the model can be accurately reproduced on a different cluster?
⚠ Common exam trap
Test-takers often think logging the model weights alone is enough, forgetting that runtime environments require explicit dependency configuration files.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A conda.yaml or requirements.txt file specifying library versions.
Capturing environment dependencies is essential for reproducibility, as models often rely on specific library versions. MLflow uses the conda.yaml or requirements.txt files to snapshot the Python environment during the log_model process. This ensures that when the model is loaded in a deployment environment, it functions exactly as it did during development, preventing runtime errors caused by version mismatches or missing dependencies in the target production environment.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The raw input dataset in Parquet format.
Why it's wrong here
The raw dataset is used for training but does not encapsulate the software environment. Including the dataset in the model signature or artifacts folder does not resolve dependency conflicts such as library version mismatches, which are the primary reason for model loading failures in production or new environments.
- ✓
A conda.yaml or requirements.txt file specifying library versions.
Why this is correct
The conda.yaml or requirements.txt file provides a precise list of dependencies. MLflow uses these files to recreate the environment during model loading or deployment, ensuring that the necessary Python packages and versions are installed, thus maintaining operational consistency between the training environment and the downstream deployment system.
- ✗
The cluster's Spark configuration JSON string.
Why it's wrong here
Spark configuration settings are cluster-specific and usually irrelevant to the model's internal logic. These settings do not contain information about the installed Python libraries or software packages required to execute the model code correctly in a containerized or remote production environment, making them unsuitable for dependency management.
- ✗
The raw model object without additional metadata.
Why it's wrong here
Logging only the model object is insufficient because the runtime environment remains unknown. Without dependency information, the model might fail to load if the target environment lacks the specific versions of libraries (e.g., scikit-learn or pandas) that were utilized during the model's original training and serialization process.
About these practice questions
This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.