Courseiva

Databricks-ML-Pro · domain

Model Development

This domain covers building and training models on Databricks: experiment tracking with MLflow, distributed training, hyperparameter tuning, and lifecycle management via Model Registry. Questions test whether you can choose the right Databricks tool for cross-validation, model versioning, drift monitoring, and artifact storage, and reason about how these integrate across the workspace.

109 questions20 easy55 medium34 hard

Focused practice

Practice Model Development questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Model Development

You must be able to log experiments with MLflow, run scalable cross-validation on Spark, register and transition model versions, and set up drift monitoring. The single most important thing: know where MLflow artifacts are stored and how the Model Registry tracks lifecycle stages.

Using MLflow Tracking to log parameters, metrics, and artifacts during model training runs.

Performing cross-validation on large data with Spark ML or spark-sklearn wrappers.

Managing model versions and stage transitions with the MLflow Model Registry.

Monitoring production feature drift using Databricks Lakehouse Monitoring or model serving metrics.

Watch out for

Common Model Development exam traps

  • ▸Assuming MLflow artifacts are stored only in the workspace filesystem; they actually go to the configured artifact store (DBFS, S3, ADLS).
  • ▸Confusing Model Registry stage transitions with deployment; registering a model does not automatically serve it.
  • ▸Using default scikit-learn cross-validation on Spark DataFrames, which collects data to the driver and fails on large datasets.

Question index

All Model Development questions (109)

Click any question to see the full explanation, or start a practice session above.

1

A data scientist is using MLflow on Databricks to track a series of experiments. They want to compare the performance of different runs and identify the best model based on a custom metric called "f1_score". Which MLflow feature should they use to efficiently compare and rank these runs?

Easy
2

A data scientist is using MLflow to log a model that includes a custom preprocessing step. They want to ensure that the preprocessing is applied consistently during both training and inference. Which approach should they take?

Hard
3

When using the Databricks Model Registry, what does a 'Model Version' represent in the context of the lifecycle?

Medium
4

A machine learning team is using Databricks to develop a model and wants to ensure that the model's input schema is validated at inference time to prevent errors from malformed data. Which TWO approaches allow them to enforce schema validation when serving the model with MLflow Model Serving? (Choose two.)

Medium
5

A data scientist is using MLflow to track a deep learning experiment on Databricks. They want to log custom metrics that are computed during training but not automatically captured by `mlflow.autolog()`. What is the correct way to log these custom metrics?

Hard
6

When developing a machine learning pipeline on Databricks, which feature provides the most effective way to track the lineage of a model from the raw data used for training to the final deployment?

Easy
7

A data scientist is using MLflow on Databricks to train a model with a custom training loop. They want to log the model so that it can be loaded later with `mlflow.pyfunc.load_model()` and used for batch inference. The model artifacts include a Python class and a configuration file. Which approach should they use to log the model?

Hard
8

When developing a machine learning model on Databricks, what is the primary benefit of using Feature Store over standard Delta Lake tables for feature management?

Easy
9

Which of the following describes the correct usage of the MLflow 'log_param' function in a Databricks environment?

Medium
10

A machine learning engineer is using Hyperopt with SparkTrials to tune a scikit-learn model on a Databricks cluster. They set max_evals=100 and parallelism=4. After the run, they notice that some trials report a loss of NaN and that the best model selected by Hyperopt has poor performance. What is the most likely reason for the NaN losses?

Hard
11

A data scientist is training a model on Databricks and wants to track experiments using MLflow. They need to record the model's hyperparameters, evaluation metrics, and the resulting model artifact. They also want to be able to compare runs and reproduce results later. Which MLflow component should they use to organize these runs?

Easy
12

When utilizing Hyperopt with MLflow on Databricks for distributed hyperparameter tuning, which TWO components are strictly required to configure the optimization run properly? (Select TWO)

Hard
13

A data scientist is using MLflow to track experiments on Databricks. They want to record the value of a hyperparameter named 'learning_rate' for a run. Which MLflow function should they use?

Easy
14

A machine learning engineer is using MLflow to track experiments on Databricks. They want to ensure that the model's input schema is enforced during inference to prevent errors from malformed data. Which MLflow feature should they use when logging the model?

Hard
15

You need to perform cross-validation on a large dataset while ensuring that your model remains performant. What is the Databricks-recommended approach?

Medium
16

A data scientist is developing a scikit-learn model on Databricks and wants to track the full lineage of the training data, including the exact Delta table version used. They are using MLflow Tracking with a Unity Catalog-enabled workspace. Which approach best captures this lineage as part of the MLflow run?

Medium
17

A machine learning engineer wants to ensure that model training artifacts are persistent and accessible even if the ephemeral compute cluster is terminated. What is the standard practice in Databricks for achieving this?

Medium
18

When utilizing the Databricks Feature Store for model development, why should a developer define a primary key in the Feature Table?

Medium
19

A machine learning engineer is using MLflow to track experiments on Databricks. They notice that when they run `mlflow.log_artifact` with a local file path inside a notebook, the artifact is stored in the run's artifact location, but when they run the same code in a job cluster, the artifact is missing. The job cluster uses the same MLflow tracking server and experiment. What is the most likely reason for the missing artifact?

Hard
20

Refer to the exhibit. A user wants to retrieve the 'accuracy' metric from this run programmatically. Which code snippet correctly accesses this value?

Medium
21

Refer to the exhibit. What is the purpose of the 'signature' section in this model configuration?

Hard
22

An ML engineer is training a model with a custom Python loop and wants MLflow to capture training metrics at regular intervals so that partial progress is visible before the run finishes. They are using `mlflow.start_run` and manual logging. Which approach correctly makes intermediate metrics visible during the run?

Hard
23

A machine learning engineer is training a model using MLflow on Databricks and wants to ensure that the model's input schema is captured and enforced during inference. They are using the `mlflow.pyfunc` flavor. Which action should they take to enable schema enforcement?

Medium
24

You are training a model on Databricks using MLflow. You need to log a custom model flavor to ensure it can be loaded in an environment without the original training code. Which approach is best practice?

Medium
25

When training a model in Databricks, which storage layer should you prioritize for training data to ensure maximum throughput and compatibility with Feature Store?

Easy
26

A machine learning engineer is building a model on Databricks and wants to use MLflow to track experiments. They need to log a custom metric that is calculated during training but is not automatically captured by `mlflow.autolog()`. They also want to ensure that the metric is associated with the correct run. Which code snippet should they use inside their training script?

Hard
27

A machine learning engineer is training a scikit-learn model on Databricks and wants to automatically log hyperparameters, metrics, and the trained artifact without writing extensive boilerplate logging code. Which approach should the engineer use?

Medium
28

A data scientist is using MLflow to track experiments in a Databricks notebook. They want to record the source code version (Git commit hash) automatically with each run. Which MLflow feature should they enable to capture this information?

Easy
29

A machine learning engineer is using Hyperopt with SparkTrials on a Databricks cluster to tune a gradient boosting model. They notice that the tuning job is running slowly because each trial trains on the full dataset, and they want to speed up the search without sacrificing final model quality. Which approach is most appropriate?

Hard
30

A machine learning engineer is developing a custom MLflow Python model that requires a pre-processing step using a scikit-learn pipeline. They want to log the model such that it can be served with the pipeline included. Which approach should they take?

Hard
31

A data scientist is training a model on a large Delta table. They want to ensure that the training data remains consistent even if the underlying table is updated during the training process. What is the most robust way to achieve this?

Medium
32

What is the best way to handle secrets (like API keys for external feature sources) within a Databricks notebook during model development?

Medium
33

A team is developing a model to forecast demand. They need to ensure that their feature engineering code is reusable for both training and real-time inference. Which architectural pattern should they adopt?

Medium
34

A team notices that their model performance is significantly lower in production than in training. They suspect 'data drift' in the feature inputs. Which Databricks capability should be used to monitor this?

Hard
35

A data scientist is using MLflow to track experiments on Databricks. They want to compare multiple runs and identify the one with the lowest validation RMSE. Which MLflow UI feature allows them to sort and filter runs by a specific metric?

Easy
36

A data scientist is using MLflow to log a model that includes a custom preprocessing step implemented in Python. They want to ensure that the preprocessing logic is packaged with the model so that it can be served consistently. Which MLflow model flavor should they use?

Hard
37

A data scientist is using Databricks AutoML to solve a classification problem. After the run completes, they want to modify the feature engineering logic for the best-performing model. Which artifact should they retrieve from the AutoML run?

Medium
38

A data scientist is developing a scikit-learn model on Databricks and wants to log the model artifact to MLflow so that it can later be deployed for online inference. They call mlflow.sklearn.log_model() without providing a signature. What is the primary consequence of omitting the model signature?

Medium
39

When evaluating a classification model on Databricks, a team needs to generate a custom performance report that is not natively provided by MLflow. What is the recommended strategy to ensure this report is persisted and associated with the training run?

Medium
40

A data scientist is using MLflow to log a model trained with scikit-learn. They want to ensure that the model can be loaded later for batch inference using `mlflow.pyfunc.load_model`. Which condition must be met for the model to be loadable as a PyFunc model?

Easy
41

Refer to the exhibit. A data scientist is logging their model training process. Which statement accurately describes the storage location of the artifacts referenced in the code snippet?

Medium
42

A machine learning engineer is preparing a scikit-learn model for batch scoring with MLflow on Databricks. The team wants the logged model to carry a reproducible environment and a machine-readable description of the input and output schema so downstream consumers can validate requests. Which TWO actions should the engineer take when logging the model with `mlflow.sklearn.log_model`? (Choose two.)

Hard
43

A machine learning engineer is preparing to deploy a model to production using MLflow Model Registry. They want to ensure that the model can be easily served and that its dependencies are correctly captured. Which TWO actions should they take when logging the model to guarantee that the serving environment can recreate the necessary Python environment? (Choose two.)

Hard
44

A machine learning engineer is building a feature engineering pipeline in Databricks using Feature Store. They need to ensure that the same feature computation logic is used for both training and batch scoring, and that features are automatically refreshed. Which approach should they take?

Hard
45

Which THREE of the following are considered best practices for handling data preprocessing in a Databricks ML pipeline to prevent data leakage?

Medium
46

Refer to the exhibit. A developer wants to ensure the Random Forest model can be used for automated inference at scale. Based on the provided code, what is missing to enable the model to support the 'predict' method within the Databricks Model Serving environment?

Hard
47

An ML engineer wants to ensure that their model training pipeline is robust against data quality issues. Which approach, if integrated into the pipeline, most effectively detects skewed or missing values before training begins?

Medium
48

Refer to the exhibit. A data scientist is logging a Scikit-Learn model to the MLflow Model Registry. Which benefit does providing the `signature` and `input_example` offer during the deployment phase?

Medium
49

A machine learning engineer is using MLflow to track experiments and wants to compare multiple runs to identify the best model. They have logged metrics such as accuracy, precision, and recall. Which MLflow feature allows them to programmatically retrieve and compare these metrics across runs for further analysis?

Hard
50

A machine learning engineer is developing a model on Databricks and wants to ensure that the model's input schema is enforced during inference. They are using MLflow to log the model. What should they do?

Medium
51

You are preparing a model for deployment in a production Databricks environment. Which THREE steps should be included in your model development pipeline to ensure model quality and traceability?

Hard
52

A data scientist is using MLflow on Databricks to tune a scikit-learn GradientBoostingRegressor with Hyperopt. They configure fmin with max_evals=50, but notice that runs appear in the experiment without parameters or metrics logged, and the best model cannot be reproduced. They want to ensure every trial is fully tracked. Which change should they make?

Medium
53

A data scientist wants to use MLflow to track a scikit-learn model training run on Databricks. They call `mlflow.sklearn.autolog()` before training. Which of the following will MLflow automatically log for this run?

Easy
54

A data scientist is training a model with scikit-learn on Databricks and wants to track the experiment using MLflow. They call mlflow.start_run() and then train the model. After training, they call mlflow.log_param() and mlflow.log_metric(), but later find that the run is not visible in the MLflow experiment UI. What is the most likely reason?

Easy
55

When working in Databricks, where should a data scientist primarily look to monitor the resource utilization and execution logs of an active model training job?

Easy
56

A machine learning engineer is using MLflow to log a model trained with a custom algorithm. They want to ensure that the model can be served with a specific input schema and that the schema is enforced during inference. Which MLflow feature should they use?

Hard
57

A data scientist is using Hyperopt with SparkTrials on a Databricks cluster to tune an XGBoost classifier. After several trials, they notice that each trial runs on a single executor and the overall tuning job takes much longer than expected. They want to speed up hyperparameter tuning without changing the search space. Which adjustment is most likely to improve performance?

Medium
58

An ML engineer is training a PyTorch model on a Databricks cluster and wants to automatically log training metrics, parameters, and the model artifact to MLflow without writing explicit mlflow.log_* calls in the training script. The engineer also needs the run to be nested under a parent run that tracks the overall experiment. Which approach should the engineer use?

Hard
59

A data scientist is using MLflow to track a training run. They want to log a dictionary of hyperparameters and a list of evaluation metrics that are computed at the end of each epoch. Which MLflow API calls should they use to log these items?

Easy
60

When logging a model that requires custom libraries (e.g., a specific version of a non-standard package), how do you ensure the environment is reproducible on the serving endpoint?

Medium
61

When developing a model, which THREE actions should a data scientist perform to ensure the model is ready for production deployment via Model Serving?

Medium
62

Which THREE features are provided by the Databricks Model Registry for model lifecycle management?

Medium
63

When hyperparameter tuning using `mlflow.spark.autolog()` or `hyperopt`, what is the primary advantage of logging the parameters to the MLflow tracking server?

Medium
64

A machine learning team is using MLflow on Databricks to manage experiments. They want to ensure that their model training runs are reproducible and that they can compare different runs effectively. Which TWO practices should they follow? (Choose two.)

Medium
65

A machine learning engineer is using MLflow to log a model. They want to include custom preprocessing logic that is not part of the model's native library. Which MLflow model flavor should they use to package the model with custom code?

Easy
66

A machine learning engineer is using MLflow to track experiments on Databricks. They want to compare multiple runs of a scikit-learn model and automatically log the best model to the Model Registry. They use `mlflow.sklearn.autolog()` and then call `mlflow.sklearn.log_model` with `registered_model_name`. However, they notice that the model version in the registry does not include the signature or input example. Which action should they take to ensure the signature and input example are logged?

Hard
67

A data scientist is building a model that requires custom preprocessing logic that is not available in standard libraries. They need to ensure this logic is bundled with the model for inference. What is the recommended approach to encapsulate this custom logic?

Medium
68

A data scientist has trained a model and wants to register it in the MLflow Model Registry on Databricks. They want to indicate that the model is ready for testing in a pre-production environment. Which stage should they transition the model version to?

Easy
69

A machine learning engineer is training a model using MLflow on Databricks and wants to compare multiple runs to select the best hyperparameters. They need to view metrics across runs in a single interface. Which MLflow feature should they use?

Easy
70

An ML engineer is training a model on Databricks using MLflow and wants to ensure that the training process is deterministic across runs. They set the random seed for NumPy, Python, and the machine learning framework. However, they observe that the model's performance varies slightly between runs on the same data and cluster configuration. Which factor is most likely causing the non-determinism?

Hard
71

You are developing an MLflow project and want to ensure that your code is reusable. What is the benefit of defining an MLproject file?

Medium
72

Which method is the most appropriate for logging custom pre-processing logic alongside a model so that it is automatically applied during inference in Databricks?

Medium
73

Which of the following is an advantage of using Databricks AutoML compared to building a custom Scikit-Learn training loop?

Medium
74

An ML engineer is using MLflow to track a deep learning experiment with PyTorch on Databricks. They want to capture the model's architecture, optimizer state, and training metrics, and later reproduce the exact training run. They call `mlflow.pytorch.autolog()` before training. After several epochs, they notice that metrics are logged but the model signature is missing, and the logged model cannot be loaded for inference without specifying the input example. What should they do to ensure the model is properly logged with a signature?

Hard
75

A data scientist is iterating on a model and notices that their training runs are becoming disorganized. What is the standard Databricks mechanism for tracking different 'attempts' at model improvement within a single project?

Medium
76

A machine learning engineer is using Databricks AutoML to train a classification model. They notice that the best model from AutoML has a high F1 score on the validation set but performs poorly on a holdout test set. They suspect that the data has a temporal component and that the default train/validation split is causing data leakage. What should they do to address this?

Hard
77

A machine learning engineer is using MLflow to log a custom PyTorch model. They define a custom pyfunc class that inherits from mlflow.pyfunc.PythonModel and implements predict(). After logging the model with mlflow.pyfunc.log_model(), they load it with mlflow.pyfunc.load_model() and call predict() with a pandas DataFrame. The prediction fails with an error about missing context. What is the most likely cause?

Hard
78

A data scientist is using MLflow to log a custom PyTorch model on Databricks. They want to ensure that the model can be loaded and used for inference without requiring the original training code. Which MLflow feature should they use to package the model with its dependencies?

Hard
79

Which Databricks feature is specifically designed to manage the lifecycle of a machine learning model, including versioning, stage transitions, and deployment tracking?

Medium
80

When using MLflow to manage the machine learning lifecycle, what is the primary purpose of the 'conda.yaml' or 'requirements.txt' file automatically generated during log_model?

Hard
81

A team is building an automated retraining pipeline. They need to ensure that only models exceeding a certain performance threshold are registered. What is the most effective way to implement this logic?

Hard
82

A data scientist is using Databricks Feature Store to build a training set for a fraud detection model. The feature table contains a column `transaction_time` that is a timestamp. After creating the training set with `create_training_set`, the resulting DataFrame includes `transaction_time` but the model training code fails because the timestamp is not accepted by the XGBoost trainer. What is the most likely cause and correct resolution?

Medium
83

A data scientist is using the Databricks Feature Store to build a training set for a fraud-detection model. The feature table is created with a primary key of `customer_id` and a timestamp key of `transaction_ts`. When calling `create_training_set`, the scientist wants to ensure that each label row receives exactly the most recent feature value available at or before the label's timestamp. Which argument must be supplied to `create_training_set` to enforce this point-in-time behavior?

Medium
84

You are performing hyperparameter tuning using Hyperopt on Databricks. Which TWO configurations must be defined to ensure optimal performance and result tracking?

Medium
85

Refer to the exhibit. A data scientist is preparing to log a model. What is the primary benefit of including the explicit 'signature' provided in the exhibit during the mlflow.log_model process?

Medium
86

You are developing a machine learning pipeline where you need to perform feature engineering on a large dataset using Spark, then train a model using Scikit-Learn. Which workflow is most efficient?

Medium
87

A machine learning team is using `mlflow.autolog()` to track experiments. They notice that certain custom metrics are not being captured. What is the most effective way to address this?

Medium
88

Which practice is most effective for managing dependencies to ensure consistent model training and inference results across different Databricks clusters?

Medium
89

A data scientist is using Databricks Feature Store to build a training set for a fraud detection model. They define a feature table with a primary key of `transaction_id` and a timestamp key of `event_ts`. When creating the training set with `create_training_set`, they specify `lookup_key=['transaction_id']`. The resulting training set contains features from multiple feature tables. Which statement describes how point-in-time correctness is ensured during this operation?

Medium
90

What is the primary purpose of registering a model in the MLflow Model Registry?

Easy
91

A machine learning team is transitioning from local notebooks to Databricks. They want to ensure their code is modular and reusable. Which THREE practices should they implement?

Medium
92

A data scientist is training a scikit-learn model on Databricks and wants to capture the best hyperparameters found during a hyperparameter sweep. They are using MLflow Tracking with nested runs. Which approach correctly records the best parameters and metrics in the parent run?

Medium
93

A machine learning engineer is developing a custom PyFunc model that combines a scikit-learn preprocessing step and a TensorFlow model. They log the model with MLflow and specify a signature. When they attempt to serve the model using Databricks Model Serving, the endpoint returns errors about incompatible input types. The signature was inferred from a pandas DataFrame with integer columns, but the serving request sends JSON with floating-point numbers. Which modification to the model signature will resolve this issue?

Hard
94

A machine learning engineer is training a model using scikit-learn on Databricks and wants to track the model's hyperparameters, metrics, and artifacts automatically without adding explicit logging calls. Which MLflow feature should they use?

Easy
95

A data scientist is training a machine learning model on Databricks using MLflow. They need to track hyperparameter tuning experiments while ensuring that each iteration is uniquely identifiable and reproducible. Which feature should they use to group related runs within a single experiment?

Medium
96

A machine learning team is using Databricks Feature Store to manage features for their models. They want to ensure that the features used during training are consistent with those served in production. Which TWO practices should they follow? (Choose two.)

Medium
97

When logging a model to the MLflow Model Registry, what is the primary benefit of using a registered model name rather than just the model URI?

Easy
98

A machine learning engineer is preparing a model for deployment using Databricks Model Serving. They need to ensure that the model's input schema is enforced and that the model can be served with a specific version. Which TWO actions should they perform? (Choose two.)

Hard
99

A machine learning engineer is training a PyTorch model on a Databricks cluster and needs to distribute the training across multiple worker nodes. Which framework should be integrated natively within Databricks to handle this distributed deep learning workflow efficiently?

Medium
100

Refer to the exhibit. You are loading a model from the registry. What does the 'models:/MyModel/1' URI specifically represent?

Hard
101

A data scientist wants to record the exact library dependencies and a code snapshot alongside a model so that a reviewer can later restore the same environment and reproduce the training run. They are logging with MLflow on Databricks. Which practice best satisfies this requirement?

Easy
102

A data scientist is developing a model on Databricks and wants to use MLflow to compare multiple runs. They need to quickly identify the run with the lowest validation loss. Which MLflow UI feature allows them to sort and filter runs based on metrics?

Medium
103

An ML engineer is training an XGBoost model on Databricks and wants to leverage hyperparameter tuning using Hyperopt while automatically logging all trial parameters, metrics, and models to MLflow. Which built-in MLflow function should be used to achieve this automatic integration?

Medium
104

A team is developing a model on Databricks and wants to run an automated hyperparameter search over a scikit-learn pipeline. They need to try many parameter combinations in parallel across cluster workers while keeping every trial's parameters and metrics in MLflow. Which Databricks capability should they use to orchestrate the search?

Medium
105

A data scientist is using MLflow to track experiments on Databricks. They notice that some runs are missing the model artifact even though they called mlflow.sklearn.log_model(). What is the most likely cause?

Medium
106

When designing a model training pipeline, which TWO features of Unity Catalog best support compliance and model governance?

Hard
107

A data scientist is training a deep learning model on Databricks. They observe that the training process is significantly slower than expected. Upon inspection, they find that data loading from DBFS is the bottleneck. What is the most effective way to improve data loading speed for deep learning training on Databricks?

Medium
108

Refer to the exhibit. What happens to these logged metrics in MLflow when the training run completes?

Medium
109

A machine learning engineer needs to track hyperparameter tuning experiments in Databricks using MLflow. Which approach best ensures that model training runs are associated with the correct code version and environment settings?

Medium

Frequently asked questions

What does the Model Development domain cover on the Databricks-ML-Pro exam?
You must be able to log experiments with MLflow, run scalable cross-validation on Spark, register and transition model versions, and set up drift monitoring. The single most important thing: know where MLflow artifacts are stored and how the Model Registry tracks lifecycle stages.
How many questions are in this domain?
This page lists all 109 Model Development questions in the Databricks-ML-Pro question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Model Development questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
databricks-ml-professional DATABRICKS-ML-PROFESSIONAL ml pro model development Practice Questions