Databricks-ML-Pro Model Development Practice Question
An ML engineer is training a model on Databricks using MLflow and wants to ensure that the training process is deterministic across runs. They set the random seed for NumPy, Python, and the machine learning framework. However, they observe that the model's performance varies slightly between runs on the same data and cluster configuration. Which factor is most likely causing the non-determinism?
⚠ Common exam trap
The trap here is assuming that setting random seeds alone guarantees determinism, while GPU-accelerated operations may still introduce non-determinism.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Non-deterministic operations in the machine learning framework, such as GPU-accelerated training with cuDNN, which may not be fully deterministic even with seeds set.
Non-determinism in deep learning on GPUs often arises from cuDNN's non-deterministic algorithms. Even with seeds set, operations like convolutions may produce slightly different results due to atomic operations or algorithm selection. To achieve determinism, you must set framework-specific flags to enforce deterministic algorithms and disable benchmarking. Other factors like autologging or tracking server configuration do not affect model training randomness.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Non-deterministic operations in the machine learning framework, such as GPU-accelerated training with cuDNN, which may not be fully deterministic even with seeds set.
Why this is correct
Many deep learning frameworks use cuDNN for GPU acceleration, which can be non-deterministic by default due to atomic operations and algorithms that are not reproducible. Even with seeds set, cuDNN may choose different algorithms for convolution or other operations, leading to slight variations. To enforce determinism, you must set framework-specific flags (e.g., torch.use_deterministic_algorithms(True)) and possibly disable cuDNN benchmarking.
- ✗
The cluster's autoscaling feature causing different numbers of executors for each run.
Why it's wrong here
Autoscaling can affect distributed training if the number of workers changes, but if the training is single-node or the parallelism is fixed, it may not be the cause. However, the question specifies same cluster configuration, implying fixed resources. The more common cause of non-determinism in deep learning is GPU operations, not cluster autoscaling. Autoscaling would more likely affect performance than model weights.
- ✗
The use of `mlflow.autolog()` which introduces randomness in logging.
Why it's wrong here
mlflow.autolog() does not introduce randomness into the training process. It only logs parameters, metrics, and models. While autologging may add overhead, it does not affect the underlying random number generation or model training. The non-determinism is more likely due to non-seeded operations in the training pipeline, such as data shuffling or parallel operations.
- ✗
The MLflow tracking server not being configured with a persistent backend store.
Why it's wrong here
The MLflow tracking server's backend store affects where runs are logged, not the training process. A non-persistent backend would lose run data but would not cause variations in model performance. The observed non-determinism is in the model training itself, not in logging. The backend store is irrelevant to the randomness of training.
About these practice questions
This Databricks-ML-Pro question is part of Courseiva's 300-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.