Databricks-ML-Pro Model Development Practice Question
Which practice is most effective for managing dependencies to ensure consistent model training and inference results across different Databricks clusters?
⚠ Common exam trap
Candidates often assume manual library installation via %pip install in the notebook is sufficient, ignoring that these changes do not persist across different clusters or automated production deployment jobs.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Include a requirements.txt file with the notebook or project and specify it in the MLflow model logging process.
Using environment management files (e.g., requirements.txt or conda.yaml) ensures that the exact library versions used during development are replicated in the training and serving environments. This consistency is critical for preventing 'works on my machine' issues, where discrepancies in library versions lead to different numerical outputs or runtime crashes during inference, thus maintaining the reliability and reproducibility of the machine learning pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Manually install libraries on each cluster node before starting the job.
Why it's wrong here
Manual installation is fragile and prone to version drift across different clusters. It is not scalable, complicates disaster recovery, and makes it impossible to track exactly which dependencies were present at the time of a specific model training run, violating the principles of reproducible machine learning and automated MLOps.
- ✓
Include a requirements.txt file with the notebook or project and specify it in the MLflow model logging process.
Why this is correct
Including a requirements file ensures that the model environment is explicitly captured as part of the model artifact. When deployed, the serving platform can recreate the exact environment, ensuring that the inference code runs against the same dependency versions used during training, which minimizes numerical inconsistencies and runtime errors.
- ✗
Rely on the default Databricks Runtime library set for all projects.
Why it's wrong here
The default Runtime libraries may not match the specific version requirements of a custom project, especially if the project relies on specific versions of Scikit-Learn, PyTorch, or TensorFlow. Relying on defaults ignores the need for precise dependency management, which is essential to avoid unpredictable model behavior across different runtime versions.
- ✗
Use broad version constraints like '>=1.0' in the environment configuration.
Why it's wrong here
Broad constraints are risky because they allow the environment to pull the latest version of a library, which might contain breaking changes. Pinning dependencies to specific versions is necessary for reproducibility, ensuring that future runs or deployments use the exact same logic and performance characteristics as the validated training run.
About these practice questions
Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.