You must be able to log experiments with MLflow, run scalable cross-validation on Spark, register and transition model versions, and set up drift monitoring. The single most important thing: know where MLflow artifacts are stored and how the Model Registry tracks lifecycle stages.
Start practicing
Model Development — choose a session length
Free · No account required
Domain overview
This domain covers building and training models on Databricks: experiment tracking with MLflow, distributed training, hyperparameter tuning, and lifecycle management via Model Registry. Questions test whether you can choose the right Databricks tool for cross-validation, model versioning, drift monitoring, and artifact storage, and reason about how these integrate across the workspace.
Exam objectives
Using MLflow Tracking to log parameters, metrics, and artifacts during model training runs.
Performing cross-validation on large data with Spark ML or spark-sklearn wrappers.
Managing model versions and stage transitions with the MLflow Model Registry.
Monitoring production feature drift using Databricks Lakehouse Monitoring or model serving metrics.
Assuming MLflow artifacts are stored only in the workspace filesystem; they actually go to the configured artifact store (DBFS, S3, ADLS).
Confusing Model Registry stage transitions with deployment; registering a model does not automatically serve it.
Using default scikit-learn cross-validation on Spark DataFrames, which collects data to the driver and fails on large datasets.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A machine learning engineer is training a PyTorch model on a Databricks cluster and needs to distribute the training across multiple worker nodes. Which framework should be integrated natively within Databricks to handle this distributed deep learning workflow efficiently?
2An ML engineer is training an XGBoost model on Databricks and wants to leverage hyperparameter tuning using Hyperopt while automatically logging all trial parameters, metrics, and models to MLflow. Which built-in MLflow function should be used to achieve this automatic integration?
3When utilizing Hyperopt with MLflow on Databricks for distributed hyperparameter tuning, which TWO components are strictly required to configure the optimization run properly? (Select TWO)
4Refer to the exhibit. A data scientist is logging their model training process. Which statement accurately describes the storage location of the artifacts referenced in the code snippet?
5When developing a machine learning pipeline on Databricks, which feature provides the most effective way to track the lineage of a model from the raw data used for training to the final deployment?
6A machine learning team is transitioning from local notebooks to Databricks. They want to ensure their code is modular and reusable. Which THREE practices should they implement?
7Which of the following describes the correct usage of the MLflow 'log_param' function in a Databricks environment?
8When evaluating a classification model on Databricks, a team needs to generate a custom performance report that is not natively provided by MLflow. What is the recommended strategy to ensure this report is persisted and associated with the training run?
9Which THREE of the following are considered best practices for handling data preprocessing in a Databricks ML pipeline to prevent data leakage?
10Refer to the exhibit. What happens to these logged metrics in MLflow when the training run completes?
11Which Databricks feature is specifically designed to manage the lifecycle of a machine learning model, including versioning, stage transitions, and deployment tracking?
12A team is building an automated retraining pipeline. They need to ensure that only models exceeding a certain performance threshold are registered. What is the most effective way to implement this logic?
13A machine learning engineer is training a scikit-learn model on Databricks and wants to automatically log hyperparameters, metrics, and the trained artifact without writing extensive boilerplate logging code. Which approach should the engineer use?
14Refer to the exhibit. A data scientist is preparing to log a model. What is the primary benefit of including the explicit 'signature' provided in the exhibit during the mlflow.log_model process?
15A machine learning engineer wants to ensure that model training artifacts are persistent and accessible even if the ephemeral compute cluster is terminated. What is the standard practice in Databricks for achieving this?
16When developing a model, which THREE actions should a data scientist perform to ensure the model is ready for production deployment via Model Serving?
17Refer to the exhibit. A developer wants to ensure the Random Forest model can be used for automated inference at scale. Based on the provided code, what is missing to enable the model to support the 'predict' method within the Databricks Model Serving environment?
18When using the Databricks Model Registry, what does a 'Model Version' represent in the context of the lifecycle?
19Which method is the most appropriate for logging custom pre-processing logic alongside a model so that it is automatically applied during inference in Databricks?
20When using MLflow to manage the machine learning lifecycle, what is the primary purpose of the 'conda.yaml' or 'requirements.txt' file automatically generated during log_model?
21A data scientist is iterating on a model and notices that their training runs are becoming disorganized. What is the standard Databricks mechanism for tracking different 'attempts' at model improvement within a single project?
22When working in Databricks, where should a data scientist primarily look to monitor the resource utilization and execution logs of an active model training job?
23Refer to the exhibit. A user wants to retrieve the 'accuracy' metric from this run programmatically. Which code snippet correctly accesses this value?
24When utilizing the Databricks Feature Store for model development, why should a developer define a primary key in the Feature Table?
25A team is developing a model to forecast demand. They need to ensure that their feature engineering code is reusable for both training and real-time inference. Which architectural pattern should they adopt?
26When hyperparameter tuning using `mlflow.spark.autolog()` or `hyperopt`, what is the primary advantage of logging the parameters to the MLflow tracking server?
27A data scientist is training a model on a large Delta table. They want to ensure that the training data remains consistent even if the underlying table is updated during the training process. What is the most robust way to achieve this?
28Refer to the exhibit. What is the purpose of the 'signature' section in this model configuration?
29Which THREE features are provided by the Databricks Model Registry for model lifecycle management?
30A team notices that their model performance is significantly lower in production than in training. They suspect 'data drift' in the feature inputs. Which Databricks capability should be used to monitor this?
31You are training a model on Databricks using MLflow. You need to log a custom model flavor to ensure it can be loaded in an environment without the original training code. Which approach is best practice?
32You are performing hyperparameter tuning using Hyperopt on Databricks. Which TWO configurations must be defined to ensure optimal performance and result tracking?
33You are developing a machine learning pipeline where you need to perform feature engineering on a large dataset using Spark, then train a model using Scikit-Learn. Which workflow is most efficient?
34You are preparing a model for deployment in a production Databricks environment. Which THREE steps should be included in your model development pipeline to ensure model quality and traceability?
35You need to perform cross-validation on a large dataset while ensuring that your model remains performant. What is the Databricks-recommended approach?
36Refer to the exhibit. You are loading a model from the registry. What does the 'models:/MyModel/1' URI specifically represent?
37When logging a model that requires custom libraries (e.g., a specific version of a non-standard package), how do you ensure the environment is reproducible on the serving endpoint?
38You are developing an MLflow project and want to ensure that your code is reusable. What is the benefit of defining an MLproject file?
39When logging a model to the MLflow Model Registry, what is the primary benefit of using a registered model name rather than just the model URI?
40A machine learning engineer needs to track hyperparameter tuning experiments in Databricks using MLflow. Which approach best ensures that model training runs are associated with the correct code version and environment settings?
41Refer to the exhibit. A data scientist is logging a Scikit-Learn model to the MLflow Model Registry. Which benefit does providing the `signature` and `input_example` offer during the deployment phase?
42A machine learning team is using `mlflow.autolog()` to track experiments. They notice that certain custom metrics are not being captured. What is the most effective way to address this?
43When training a model in Databricks, which storage layer should you prioritize for training data to ensure maximum throughput and compatibility with Feature Store?
44Which practice is most effective for managing dependencies to ensure consistent model training and inference results across different Databricks clusters?
45What is the primary purpose of registering a model in the MLflow Model Registry?
46When designing a model training pipeline, which TWO features of Unity Catalog best support compliance and model governance?
47An ML engineer wants to ensure that their model training pipeline is robust against data quality issues. Which approach, if integrated into the pipeline, most effectively detects skewed or missing values before training begins?
48What is the best way to handle secrets (like API keys for external feature sources) within a Databricks notebook during model development?
49Which of the following is an advantage of using Databricks AutoML compared to building a custom Scikit-Learn training loop?
50A data scientist is training a machine learning model on Databricks using MLflow. They need to track hyperparameter tuning experiments while ensuring that each iteration is uniquely identifiable and reproducible. Which feature should they use to group related runs within a single experiment?
51When developing a machine learning model on Databricks, what is the primary benefit of using Feature Store over standard Delta Lake tables for feature management?
52A data scientist is using Databricks AutoML to solve a classification problem. After the run completes, they want to modify the feature engineering logic for the best-performing model. Which artifact should they retrieve from the AutoML run?
53A data scientist is building a model that requires custom preprocessing logic that is not available in standard libraries. They need to ensure this logic is bundled with the model for inference. What is the recommended approach to encapsulate this custom logic?
54A data scientist is training a deep learning model on Databricks. They observe that the training process is significantly slower than expected. Upon inspection, they find that data loading from DBFS is the bottleneck. What is the most effective way to improve data loading speed for deep learning training on Databricks?
55A data scientist is developing a scikit-learn model on Databricks and wants to track the full lineage of the training data, including the exact Delta table version used. They are using MLflow Tracking with a Unity Catalog-enabled workspace. Which approach best captures this lineage as part of the MLflow run?
56A data scientist is using MLflow to track experiments on Databricks. They want to compare multiple runs and identify the one with the lowest validation RMSE. Which MLflow UI feature allows them to sort and filter runs by a specific metric?
57A machine learning engineer is developing a custom MLflow Python model that requires a pre-processing step using a scikit-learn pipeline. They want to log the model such that it can be served with the pipeline included. Which approach should they take?
58A machine learning engineer is building a feature engineering pipeline in Databricks using Feature Store. They need to ensure that the same feature computation logic is used for both training and batch scoring, and that features are automatically refreshed. Which approach should they take?
59A data scientist is using MLflow on Databricks to tune a scikit-learn GradientBoostingRegressor with Hyperopt. They configure fmin with max_evals=50, but notice that runs appear in the experiment without parameters or metrics logged, and the best model cannot be reproduced. They want to ensure every trial is fully tracked. Which change should they make?
60A data scientist is using MLflow to track experiments on Databricks. They want to record the value of a hyperparameter named 'learning_rate' for a run. Which MLflow function should they use?
61A data scientist is using MLflow on Databricks to train a model with a custom training loop. They want to log the model so that it can be loaded later with `mlflow.pyfunc.load_model()` and used for batch inference. The model artifacts include a Python class and a configuration file. Which approach should they use to log the model?
62A data scientist is using MLflow on Databricks to track a series of experiments. They want to compare the performance of different runs and identify the best model based on a custom metric called "f1_score". Which MLflow feature should they use to efficiently compare and rank these runs?
63An ML engineer is using MLflow to track a deep learning experiment with PyTorch on Databricks. They want to capture the model's architecture, optimizer state, and training metrics, and later reproduce the exact training run. They call `mlflow.pytorch.autolog()` before training. After several epochs, they notice that metrics are logged but the model signature is missing, and the logged model cannot be loaded for inference without specifying the input example. What should they do to ensure the model is properly logged with a signature?
64A machine learning engineer is preparing a model for deployment using Databricks Model Serving. They need to ensure that the model's input schema is enforced and that the model can be served with a specific version. Which TWO actions should they perform? (Choose two.)
65A machine learning engineer is training a model using MLflow on Databricks and wants to ensure that the model's input schema is captured and enforced during inference. They are using the `mlflow.pyfunc` flavor. Which action should they take to enable schema enforcement?
66A machine learning engineer is using Hyperopt with SparkTrials to tune a scikit-learn model on a Databricks cluster. They set max_evals=100 and parallelism=4. After the run, they notice that some trials report a loss of NaN and that the best model selected by Hyperopt has poor performance. What is the most likely reason for the NaN losses?
67A data scientist is using MLflow to track experiments on Databricks. They notice that some runs are missing the model artifact even though they called mlflow.sklearn.log_model(). What is the most likely cause?
68A data scientist is using MLflow to log a custom PyTorch model on Databricks. They want to ensure that the model can be loaded and used for inference without requiring the original training code. Which MLflow feature should they use to package the model with its dependencies?
69A data scientist is using MLflow to track a deep learning experiment on Databricks. They want to log custom metrics that are computed during training but not automatically captured by `mlflow.autolog()`. What is the correct way to log these custom metrics?
70A data scientist is training a scikit-learn model on Databricks and wants to capture the best hyperparameters found during a hyperparameter sweep. They are using MLflow Tracking with nested runs. Which approach correctly records the best parameters and metrics in the parent run?
71A machine learning team is using MLflow on Databricks to manage experiments. They want to ensure that their model training runs are reproducible and that they can compare different runs effectively. Which TWO practices should they follow? (Choose two.)
72A machine learning engineer is using MLflow to track experiments on Databricks. They want to ensure that the model's input schema is enforced during inference to prevent errors from malformed data. Which MLflow feature should they use when logging the model?
73A data scientist is using Hyperopt with SparkTrials on a Databricks cluster to tune an XGBoost classifier. After several trials, they notice that each trial runs on a single executor and the overall tuning job takes much longer than expected. They want to speed up hyperparameter tuning without changing the search space. Which adjustment is most likely to improve performance?
74An ML engineer is training a model on Databricks using MLflow and wants to ensure that the training process is deterministic across runs. They set the random seed for NumPy, Python, and the machine learning framework. However, they observe that the model's performance varies slightly between runs on the same data and cluster configuration. Which factor is most likely causing the non-determinism?
75A data scientist is training a model with scikit-learn on Databricks and wants to track the experiment using MLflow. They call mlflow.start_run() and then train the model. After training, they call mlflow.log_param() and mlflow.log_metric(), but later find that the run is not visible in the MLflow experiment UI. What is the most likely reason?
76A data scientist is developing a scikit-learn model on Databricks and wants to log the model artifact to MLflow so that it can later be deployed for online inference. They call mlflow.sklearn.log_model() without providing a signature. What is the primary consequence of omitting the model signature?
77A machine learning engineer is using MLflow to log a model trained with a custom algorithm. They want to ensure that the model can be served with a specific input schema and that the schema is enforced during inference. Which MLflow feature should they use?
78A data scientist is using Databricks Feature Store to build a training set for a fraud detection model. They define a feature table with a primary key of `transaction_id` and a timestamp key of `event_ts`. When creating the training set with `create_training_set`, they specify `lookup_key=['transaction_id']`. The resulting training set contains features from multiple feature tables. Which statement describes how point-in-time correctness is ensured during this operation?
79A data scientist is developing a model on Databricks and wants to use MLflow to compare multiple runs. They need to quickly identify the run with the lowest validation loss. Which MLflow UI feature allows them to sort and filter runs based on metrics?
80A machine learning engineer is using MLflow to log a custom PyTorch model. They define a custom pyfunc class that inherits from mlflow.pyfunc.PythonModel and implements predict(). After logging the model with mlflow.pyfunc.log_model(), they load it with mlflow.pyfunc.load_model() and call predict() with a pandas DataFrame. The prediction fails with an error about missing context. What is the most likely cause?
81A machine learning engineer is using Hyperopt with SparkTrials on a Databricks cluster to tune a gradient boosting model. They notice that the tuning job is running slowly because each trial trains on the full dataset, and they want to speed up the search without sacrificing final model quality. Which approach is most appropriate?
82A machine learning engineer is using MLflow to track experiments on Databricks. They want to compare multiple runs of a scikit-learn model and automatically log the best model to the Model Registry. They use `mlflow.sklearn.autolog()` and then call `mlflow.sklearn.log_model` with `registered_model_name`. However, they notice that the model version in the registry does not include the signature or input example. Which action should they take to ensure the signature and input example are logged?
83A machine learning engineer is using MLflow to log a model. They want to include custom preprocessing logic that is not part of the model's native library. Which MLflow model flavor should they use to package the model with custom code?
84A machine learning engineer is developing a model on Databricks and wants to ensure that the model's input schema is enforced during inference. They are using MLflow to log the model. What should they do?
85A data scientist is using MLflow to track a training run. They want to log a dictionary of hyperparameters and a list of evaluation metrics that are computed at the end of each epoch. Which MLflow API calls should they use to log these items?
86A data scientist is using MLflow to log a model that includes a custom preprocessing step. They want to ensure that the preprocessing is applied consistently during both training and inference. Which approach should they take?
87A data scientist is using the Databricks Feature Store to build a training set for a fraud-detection model. The feature table is created with a primary key of `customer_id` and a timestamp key of `transaction_ts`. When calling `create_training_set`, the scientist wants to ensure that each label row receives exactly the most recent feature value available at or before the label's timestamp. Which argument must be supplied to `create_training_set` to enforce this point-in-time behavior?
88A machine learning engineer is preparing to deploy a model to production using MLflow Model Registry. They want to ensure that the model can be easily served and that its dependencies are correctly captured. Which TWO actions should they take when logging the model to guarantee that the serving environment can recreate the necessary Python environment? (Choose two.)
89A machine learning engineer is training a model using scikit-learn on Databricks and wants to track the model's hyperparameters, metrics, and artifacts automatically without adding explicit logging calls. Which MLflow feature should they use?
90A data scientist has trained a model and wants to register it in the MLflow Model Registry on Databricks. They want to indicate that the model is ready for testing in a pre-production environment. Which stage should they transition the model version to?
91A machine learning engineer is preparing a scikit-learn model for batch scoring with MLflow on Databricks. The team wants the logged model to carry a reproducible environment and a machine-readable description of the input and output schema so downstream consumers can validate requests. Which TWO actions should the engineer take when logging the model with `mlflow.sklearn.log_model`? (Choose two.)
92A machine learning engineer is training a model using MLflow on Databricks and wants to compare multiple runs to select the best hyperparameters. They need to view metrics across runs in a single interface. Which MLflow feature should they use?
93A data scientist is using Databricks Feature Store to build a training set for a fraud detection model. The feature table contains a column `transaction_time` that is a timestamp. After creating the training set with `create_training_set`, the resulting DataFrame includes `transaction_time` but the model training code fails because the timestamp is not accepted by the XGBoost trainer. What is the most likely cause and correct resolution?
94A machine learning engineer is using MLflow to track experiments and wants to compare multiple runs to identify the best model. They have logged metrics such as accuracy, precision, and recall. Which MLflow feature allows them to programmatically retrieve and compare these metrics across runs for further analysis?
95A team is developing a model on Databricks and wants to run an automated hyperparameter search over a scikit-learn pipeline. They need to try many parameter combinations in parallel across cluster workers while keeping every trial's parameters and metrics in MLflow. Which Databricks capability should they use to orchestrate the search?
96A data scientist is using MLflow to log a model that includes a custom preprocessing step implemented in Python. They want to ensure that the preprocessing logic is packaged with the model so that it can be served consistently. Which MLflow model flavor should they use?
97A machine learning engineer is using MLflow to track experiments on Databricks. They notice that when they run `mlflow.log_artifact` with a local file path inside a notebook, the artifact is stored in the run's artifact location, but when they run the same code in a job cluster, the artifact is missing. The job cluster uses the same MLflow tracking server and experiment. What is the most likely reason for the missing artifact?
98A data scientist wants to record the exact library dependencies and a code snapshot alongside a model so that a reviewer can later restore the same environment and reproduce the training run. They are logging with MLflow on Databricks. Which practice best satisfies this requirement?
99A data scientist is training a model on Databricks and wants to track experiments using MLflow. They need to record the model's hyperparameters, evaluation metrics, and the resulting model artifact. They also want to be able to compare runs and reproduce results later. Which MLflow component should they use to organize these runs?
100A data scientist wants to use MLflow to track a scikit-learn model training run on Databricks. They call `mlflow.sklearn.autolog()` before training. Which of the following will MLflow automatically log for this run?
101An ML engineer is training a model with a custom Python loop and wants MLflow to capture training metrics at regular intervals so that partial progress is visible before the run finishes. They are using `mlflow.start_run` and manual logging. Which approach correctly makes intermediate metrics visible during the run?
102A data scientist is using MLflow to log a model trained with scikit-learn. They want to ensure that the model can be loaded later for batch inference using `mlflow.pyfunc.load_model`. Which condition must be met for the model to be loadable as a PyFunc model?
103A machine learning team is using Databricks Feature Store to manage features for their models. They want to ensure that the features used during training are consistent with those served in production. Which TWO practices should they follow? (Choose two.)
104A machine learning team is using Databricks to develop a model and wants to ensure that the model's input schema is validated at inference time to prevent errors from malformed data. Which TWO approaches allow them to enforce schema validation when serving the model with MLflow Model Serving? (Choose two.)
105A machine learning engineer is using Databricks AutoML to train a classification model. They notice that the best model from AutoML has a high F1 score on the validation set but performs poorly on a holdout test set. They suspect that the data has a temporal component and that the default train/validation split is causing data leakage. What should they do to address this?
106A machine learning engineer is building a model on Databricks and wants to use MLflow to track experiments. They need to log a custom metric that is calculated during training but is not automatically captured by `mlflow.autolog()`. They also want to ensure that the metric is associated with the correct run. Which code snippet should they use inside their training script?
107A data scientist is using MLflow to track experiments in a Databricks notebook. They want to record the source code version (Git commit hash) automatically with each run. Which MLflow feature should they enable to capture this information?
108A machine learning engineer is developing a custom PyFunc model that combines a scikit-learn preprocessing step and a TensorFlow model. They log the model with MLflow and specify a signature. When they attempt to serve the model using Databricks Model Serving, the endpoint returns errors about incompatible input types. The signature was inferred from a pandas DataFrame with integer columns, but the serving request sends JSON with floating-point numbers. Which modification to the model signature will resolve this issue?
109An ML engineer is training a PyTorch model on a Databricks cluster and wants to automatically log training metrics, parameters, and the model artifact to MLflow without writing explicit mlflow.log_* calls in the training script. The engineer also needs the run to be nested under a parent run that tracks the overall experiment. Which approach should the engineer use?
You must be able to log experiments with MLflow, run scalable cross-validation on Spark, register and transition model versions, and set up drift monitoring. The single most important thing: know where MLflow artifacts are stored and how the Model Registry tracks lifecycle stages.
The Courseiva Databricks-ML-Pro question bank contains 109 questions in the Model Development domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Model Development domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included