Databricks-ML-Pro · domain
Model Development
This domain covers building and training models on Databricks: experiment tracking with MLflow, distributed training, hyperparameter tuning, and lifecycle management via Model Registry. Questions test whether you can choose the right Databricks tool for cross-validation, model versioning, drift monitoring, and artifact storage, and reason about how these integrate across the workspace.
Focused practice
Practice Model Development questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Model Development
You must be able to log experiments with MLflow, run scalable cross-validation on Spark, register and transition model versions, and set up drift monitoring. The single most important thing: know where MLflow artifacts are stored and how the Model Registry tracks lifecycle stages.
Using MLflow Tracking to log parameters, metrics, and artifacts during model training runs.
Performing cross-validation on large data with Spark ML or spark-sklearn wrappers.
Managing model versions and stage transitions with the MLflow Model Registry.
Monitoring production feature drift using Databricks Lakehouse Monitoring or model serving metrics.
Watch out for
Common Model Development exam traps
- ▸Assuming MLflow artifacts are stored only in the workspace filesystem; they actually go to the configured artifact store (DBFS, S3, ADLS).
- ▸Confusing Model Registry stage transitions with deployment; registering a model does not automatically serve it.
- ▸Using default scikit-learn cross-validation on Spark DataFrames, which collects data to the driver and fails on large datasets.
Question index
All Model Development questions (109)
Click any question to see the full explanation, or start a practice session above.
A data scientist is using MLflow on Databricks to track a series of experiments. They want to compare the performance of different runs and identify the best model based on a custom metric called "f1_score". Which MLflow feature should they use to efficiently compare and rank these runs?
Easy2A data scientist is using MLflow to log a model that includes a custom preprocessing step. They want to ensure that the preprocessing is applied consistently during both training and inference. Which approach should they take?
Hard3When using the Databricks Model Registry, what does a 'Model Version' represent in the context of the lifecycle?
Medium4A machine learning team is using Databricks to develop a model and wants to ensure that the model's input schema is validated at inference time to prevent errors from malformed data. Which TWO approaches allow them to enforce schema validation when serving the model with MLflow Model Serving? (Choose two.)
Medium5A data scientist is using MLflow to track a deep learning experiment on Databricks. They want to log custom metrics that are computed during training but not automatically captured by `mlflow.autolog()`. What is the correct way to log these custom metrics?
Hard6When developing a machine learning pipeline on Databricks, which feature provides the most effective way to track the lineage of a model from the raw data used for training to the final deployment?
Easy7A data scientist is using MLflow on Databricks to train a model with a custom training loop. They want to log the model so that it can be loaded later with `mlflow.pyfunc.load_model()` and used for batch inference. The model artifacts include a Python class and a configuration file. Which approach should they use to log the model?
Hard8When developing a machine learning model on Databricks, what is the primary benefit of using Feature Store over standard Delta Lake tables for feature management?
Easy9Which of the following describes the correct usage of the MLflow 'log_param' function in a Databricks environment?
Medium10A machine learning engineer is using Hyperopt with SparkTrials to tune a scikit-learn model on a Databricks cluster. They set max_evals=100 and parallelism=4. After the run, they notice that some trials report a loss of NaN and that the best model selected by Hyperopt has poor performance. What is the most likely reason for the NaN losses?
Hard11A data scientist is training a model on Databricks and wants to track experiments using MLflow. They need to record the model's hyperparameters, evaluation metrics, and the resulting model artifact. They also want to be able to compare runs and reproduce results later. Which MLflow component should they use to organize these runs?
Easy12When utilizing Hyperopt with MLflow on Databricks for distributed hyperparameter tuning, which TWO components are strictly required to configure the optimization run properly? (Select TWO)
Hard13A data scientist is using MLflow to track experiments on Databricks. They want to record the value of a hyperparameter named 'learning_rate' for a run. Which MLflow function should they use?
Easy14A machine learning engineer is using MLflow to track experiments on Databricks. They want to ensure that the model's input schema is enforced during inference to prevent errors from malformed data. Which MLflow feature should they use when logging the model?
Hard15You need to perform cross-validation on a large dataset while ensuring that your model remains performant. What is the Databricks-recommended approach?
Medium16A data scientist is developing a scikit-learn model on Databricks and wants to track the full lineage of the training data, including the exact Delta table version used. They are using MLflow Tracking with a Unity Catalog-enabled workspace. Which approach best captures this lineage as part of the MLflow run?
Medium17A machine learning engineer wants to ensure that model training artifacts are persistent and accessible even if the ephemeral compute cluster is terminated. What is the standard practice in Databricks for achieving this?
Medium18When utilizing the Databricks Feature Store for model development, why should a developer define a primary key in the Feature Table?
Medium19A machine learning engineer is using MLflow to track experiments on Databricks. They notice that when they run `mlflow.log_artifact` with a local file path inside a notebook, the artifact is stored in the run's artifact location, but when they run the same code in a job cluster, the artifact is missing. The job cluster uses the same MLflow tracking server and experiment. What is the most likely reason for the missing artifact?
Hard20Refer to the exhibit. A user wants to retrieve the 'accuracy' metric from this run programmatically. Which code snippet correctly accesses this value?
Medium21Refer to the exhibit. What is the purpose of the 'signature' section in this model configuration?
Hard22An ML engineer is training a model with a custom Python loop and wants MLflow to capture training metrics at regular intervals so that partial progress is visible before the run finishes. They are using `mlflow.start_run` and manual logging. Which approach correctly makes intermediate metrics visible during the run?
Hard23A machine learning engineer is training a model using MLflow on Databricks and wants to ensure that the model's input schema is captured and enforced during inference. They are using the `mlflow.pyfunc` flavor. Which action should they take to enable schema enforcement?
Medium24You are training a model on Databricks using MLflow. You need to log a custom model flavor to ensure it can be loaded in an environment without the original training code. Which approach is best practice?
Medium25When training a model in Databricks, which storage layer should you prioritize for training data to ensure maximum throughput and compatibility with Feature Store?
Easy26A machine learning engineer is building a model on Databricks and wants to use MLflow to track experiments. They need to log a custom metric that is calculated during training but is not automatically captured by `mlflow.autolog()`. They also want to ensure that the metric is associated with the correct run. Which code snippet should they use inside their training script?
Hard27A machine learning engineer is training a scikit-learn model on Databricks and wants to automatically log hyperparameters, metrics, and the trained artifact without writing extensive boilerplate logging code. Which approach should the engineer use?
Medium28A data scientist is using MLflow to track experiments in a Databricks notebook. They want to record the source code version (Git commit hash) automatically with each run. Which MLflow feature should they enable to capture this information?
Easy29A machine learning engineer is using Hyperopt with SparkTrials on a Databricks cluster to tune a gradient boosting model. They notice that the tuning job is running slowly because each trial trains on the full dataset, and they want to speed up the search without sacrificing final model quality. Which approach is most appropriate?
Hard30A machine learning engineer is developing a custom MLflow Python model that requires a pre-processing step using a scikit-learn pipeline. They want to log the model such that it can be served with the pipeline included. Which approach should they take?
Hard31A data scientist is training a model on a large Delta table. They want to ensure that the training data remains consistent even if the underlying table is updated during the training process. What is the most robust way to achieve this?
Medium32What is the best way to handle secrets (like API keys for external feature sources) within a Databricks notebook during model development?
Medium33A team is developing a model to forecast demand. They need to ensure that their feature engineering code is reusable for both training and real-time inference. Which architectural pattern should they adopt?
Medium34A team notices that their model performance is significantly lower in production than in training. They suspect 'data drift' in the feature inputs. Which Databricks capability should be used to monitor this?
Hard35A data scientist is using MLflow to track experiments on Databricks. They want to compare multiple runs and identify the one with the lowest validation RMSE. Which MLflow UI feature allows them to sort and filter runs by a specific metric?
Easy36A data scientist is using MLflow to log a model that includes a custom preprocessing step implemented in Python. They want to ensure that the preprocessing logic is packaged with the model so that it can be served consistently. Which MLflow model flavor should they use?
Hard37A data scientist is using Databricks AutoML to solve a classification problem. After the run completes, they want to modify the feature engineering logic for the best-performing model. Which artifact should they retrieve from the AutoML run?
Medium38A data scientist is developing a scikit-learn model on Databricks and wants to log the model artifact to MLflow so that it can later be deployed for online inference. They call mlflow.sklearn.log_model() without providing a signature. What is the primary consequence of omitting the model signature?
Medium39When evaluating a classification model on Databricks, a team needs to generate a custom performance report that is not natively provided by MLflow. What is the recommended strategy to ensure this report is persisted and associated with the training run?
Medium40A data scientist is using MLflow to log a model trained with scikit-learn. They want to ensure that the model can be loaded later for batch inference using `mlflow.pyfunc.load_model`. Which condition must be met for the model to be loadable as a PyFunc model?
Easy41Refer to the exhibit. A data scientist is logging their model training process. Which statement accurately describes the storage location of the artifacts referenced in the code snippet?
Medium42A machine learning engineer is preparing a scikit-learn model for batch scoring with MLflow on Databricks. The team wants the logged model to carry a reproducible environment and a machine-readable description of the input and output schema so downstream consumers can validate requests. Which TWO actions should the engineer take when logging the model with `mlflow.sklearn.log_model`? (Choose two.)
Hard43A machine learning engineer is preparing to deploy a model to production using MLflow Model Registry. They want to ensure that the model can be easily served and that its dependencies are correctly captured. Which TWO actions should they take when logging the model to guarantee that the serving environment can recreate the necessary Python environment? (Choose two.)
Hard44A machine learning engineer is building a feature engineering pipeline in Databricks using Feature Store. They need to ensure that the same feature computation logic is used for both training and batch scoring, and that features are automatically refreshed. Which approach should they take?
Hard45Which THREE of the following are considered best practices for handling data preprocessing in a Databricks ML pipeline to prevent data leakage?
Medium46Refer to the exhibit. A developer wants to ensure the Random Forest model can be used for automated inference at scale. Based on the provided code, what is missing to enable the model to support the 'predict' method within the Databricks Model Serving environment?
Hard47An ML engineer wants to ensure that their model training pipeline is robust against data quality issues. Which approach, if integrated into the pipeline, most effectively detects skewed or missing values before training begins?
Medium48Refer to the exhibit. A data scientist is logging a Scikit-Learn model to the MLflow Model Registry. Which benefit does providing the `signature` and `input_example` offer during the deployment phase?
Medium49A machine learning engineer is using MLflow to track experiments and wants to compare multiple runs to identify the best model. They have logged metrics such as accuracy, precision, and recall. Which MLflow feature allows them to programmatically retrieve and compare these metrics across runs for further analysis?
Hard50A machine learning engineer is developing a model on Databricks and wants to ensure that the model's input schema is enforced during inference. They are using MLflow to log the model. What should they do?
Medium51You are preparing a model for deployment in a production Databricks environment. Which THREE steps should be included in your model development pipeline to ensure model quality and traceability?
Hard52A data scientist is using MLflow on Databricks to tune a scikit-learn GradientBoostingRegressor with Hyperopt. They configure fmin with max_evals=50, but notice that runs appear in the experiment without parameters or metrics logged, and the best model cannot be reproduced. They want to ensure every trial is fully tracked. Which change should they make?
Medium53A data scientist wants to use MLflow to track a scikit-learn model training run on Databricks. They call `mlflow.sklearn.autolog()` before training. Which of the following will MLflow automatically log for this run?
Easy54A data scientist is training a model with scikit-learn on Databricks and wants to track the experiment using MLflow. They call mlflow.start_run() and then train the model. After training, they call mlflow.log_param() and mlflow.log_metric(), but later find that the run is not visible in the MLflow experiment UI. What is the most likely reason?
Easy55When working in Databricks, where should a data scientist primarily look to monitor the resource utilization and execution logs of an active model training job?
Easy56A machine learning engineer is using MLflow to log a model trained with a custom algorithm. They want to ensure that the model can be served with a specific input schema and that the schema is enforced during inference. Which MLflow feature should they use?
Hard57A data scientist is using Hyperopt with SparkTrials on a Databricks cluster to tune an XGBoost classifier. After several trials, they notice that each trial runs on a single executor and the overall tuning job takes much longer than expected. They want to speed up hyperparameter tuning without changing the search space. Which adjustment is most likely to improve performance?
Medium58An ML engineer is training a PyTorch model on a Databricks cluster and wants to automatically log training metrics, parameters, and the model artifact to MLflow without writing explicit mlflow.log_* calls in the training script. The engineer also needs the run to be nested under a parent run that tracks the overall experiment. Which approach should the engineer use?
Hard59A data scientist is using MLflow to track a training run. They want to log a dictionary of hyperparameters and a list of evaluation metrics that are computed at the end of each epoch. Which MLflow API calls should they use to log these items?
Easy60When logging a model that requires custom libraries (e.g., a specific version of a non-standard package), how do you ensure the environment is reproducible on the serving endpoint?
Medium61When developing a model, which THREE actions should a data scientist perform to ensure the model is ready for production deployment via Model Serving?
Medium62Which THREE features are provided by the Databricks Model Registry for model lifecycle management?
Medium63When hyperparameter tuning using `mlflow.spark.autolog()` or `hyperopt`, what is the primary advantage of logging the parameters to the MLflow tracking server?
Medium64A machine learning team is using MLflow on Databricks to manage experiments. They want to ensure that their model training runs are reproducible and that they can compare different runs effectively. Which TWO practices should they follow? (Choose two.)
Medium65A machine learning engineer is using MLflow to log a model. They want to include custom preprocessing logic that is not part of the model's native library. Which MLflow model flavor should they use to package the model with custom code?
Easy66A machine learning engineer is using MLflow to track experiments on Databricks. They want to compare multiple runs of a scikit-learn model and automatically log the best model to the Model Registry. They use `mlflow.sklearn.autolog()` and then call `mlflow.sklearn.log_model` with `registered_model_name`. However, they notice that the model version in the registry does not include the signature or input example. Which action should they take to ensure the signature and input example are logged?
Hard67A data scientist is building a model that requires custom preprocessing logic that is not available in standard libraries. They need to ensure this logic is bundled with the model for inference. What is the recommended approach to encapsulate this custom logic?
Medium68A data scientist has trained a model and wants to register it in the MLflow Model Registry on Databricks. They want to indicate that the model is ready for testing in a pre-production environment. Which stage should they transition the model version to?
Easy69A machine learning engineer is training a model using MLflow on Databricks and wants to compare multiple runs to select the best hyperparameters. They need to view metrics across runs in a single interface. Which MLflow feature should they use?
Easy70An ML engineer is training a model on Databricks using MLflow and wants to ensure that the training process is deterministic across runs. They set the random seed for NumPy, Python, and the machine learning framework. However, they observe that the model's performance varies slightly between runs on the same data and cluster configuration. Which factor is most likely causing the non-determinism?
Hard71You are developing an MLflow project and want to ensure that your code is reusable. What is the benefit of defining an MLproject file?
Medium72Which method is the most appropriate for logging custom pre-processing logic alongside a model so that it is automatically applied during inference in Databricks?
Medium73Which of the following is an advantage of using Databricks AutoML compared to building a custom Scikit-Learn training loop?
Medium74An ML engineer is using MLflow to track a deep learning experiment with PyTorch on Databricks. They want to capture the model's architecture, optimizer state, and training metrics, and later reproduce the exact training run. They call `mlflow.pytorch.autolog()` before training. After several epochs, they notice that metrics are logged but the model signature is missing, and the logged model cannot be loaded for inference without specifying the input example. What should they do to ensure the model is properly logged with a signature?
Hard75A data scientist is iterating on a model and notices that their training runs are becoming disorganized. What is the standard Databricks mechanism for tracking different 'attempts' at model improvement within a single project?
Medium76A machine learning engineer is using Databricks AutoML to train a classification model. They notice that the best model from AutoML has a high F1 score on the validation set but performs poorly on a holdout test set. They suspect that the data has a temporal component and that the default train/validation split is causing data leakage. What should they do to address this?
Hard77A machine learning engineer is using MLflow to log a custom PyTorch model. They define a custom pyfunc class that inherits from mlflow.pyfunc.PythonModel and implements predict(). After logging the model with mlflow.pyfunc.log_model(), they load it with mlflow.pyfunc.load_model() and call predict() with a pandas DataFrame. The prediction fails with an error about missing context. What is the most likely cause?
Hard78A data scientist is using MLflow to log a custom PyTorch model on Databricks. They want to ensure that the model can be loaded and used for inference without requiring the original training code. Which MLflow feature should they use to package the model with its dependencies?
Hard79Which Databricks feature is specifically designed to manage the lifecycle of a machine learning model, including versioning, stage transitions, and deployment tracking?
Medium80When using MLflow to manage the machine learning lifecycle, what is the primary purpose of the 'conda.yaml' or 'requirements.txt' file automatically generated during log_model?
Hard81A team is building an automated retraining pipeline. They need to ensure that only models exceeding a certain performance threshold are registered. What is the most effective way to implement this logic?
Hard82A data scientist is using Databricks Feature Store to build a training set for a fraud detection model. The feature table contains a column `transaction_time` that is a timestamp. After creating the training set with `create_training_set`, the resulting DataFrame includes `transaction_time` but the model training code fails because the timestamp is not accepted by the XGBoost trainer. What is the most likely cause and correct resolution?
Medium83A data scientist is using the Databricks Feature Store to build a training set for a fraud-detection model. The feature table is created with a primary key of `customer_id` and a timestamp key of `transaction_ts`. When calling `create_training_set`, the scientist wants to ensure that each label row receives exactly the most recent feature value available at or before the label's timestamp. Which argument must be supplied to `create_training_set` to enforce this point-in-time behavior?
Medium84You are performing hyperparameter tuning using Hyperopt on Databricks. Which TWO configurations must be defined to ensure optimal performance and result tracking?
Medium85Refer to the exhibit. A data scientist is preparing to log a model. What is the primary benefit of including the explicit 'signature' provided in the exhibit during the mlflow.log_model process?
Medium86You are developing a machine learning pipeline where you need to perform feature engineering on a large dataset using Spark, then train a model using Scikit-Learn. Which workflow is most efficient?
Medium87A machine learning team is using `mlflow.autolog()` to track experiments. They notice that certain custom metrics are not being captured. What is the most effective way to address this?
Medium88Which practice is most effective for managing dependencies to ensure consistent model training and inference results across different Databricks clusters?
Medium89A data scientist is using Databricks Feature Store to build a training set for a fraud detection model. They define a feature table with a primary key of `transaction_id` and a timestamp key of `event_ts`. When creating the training set with `create_training_set`, they specify `lookup_key=['transaction_id']`. The resulting training set contains features from multiple feature tables. Which statement describes how point-in-time correctness is ensured during this operation?
Medium90What is the primary purpose of registering a model in the MLflow Model Registry?
Easy91A machine learning team is transitioning from local notebooks to Databricks. They want to ensure their code is modular and reusable. Which THREE practices should they implement?
Medium92A data scientist is training a scikit-learn model on Databricks and wants to capture the best hyperparameters found during a hyperparameter sweep. They are using MLflow Tracking with nested runs. Which approach correctly records the best parameters and metrics in the parent run?
Medium93A machine learning engineer is developing a custom PyFunc model that combines a scikit-learn preprocessing step and a TensorFlow model. They log the model with MLflow and specify a signature. When they attempt to serve the model using Databricks Model Serving, the endpoint returns errors about incompatible input types. The signature was inferred from a pandas DataFrame with integer columns, but the serving request sends JSON with floating-point numbers. Which modification to the model signature will resolve this issue?
Hard94A machine learning engineer is training a model using scikit-learn on Databricks and wants to track the model's hyperparameters, metrics, and artifacts automatically without adding explicit logging calls. Which MLflow feature should they use?
Easy95A data scientist is training a machine learning model on Databricks using MLflow. They need to track hyperparameter tuning experiments while ensuring that each iteration is uniquely identifiable and reproducible. Which feature should they use to group related runs within a single experiment?
Medium96A machine learning team is using Databricks Feature Store to manage features for their models. They want to ensure that the features used during training are consistent with those served in production. Which TWO practices should they follow? (Choose two.)
Medium97When logging a model to the MLflow Model Registry, what is the primary benefit of using a registered model name rather than just the model URI?
Easy98A machine learning engineer is preparing a model for deployment using Databricks Model Serving. They need to ensure that the model's input schema is enforced and that the model can be served with a specific version. Which TWO actions should they perform? (Choose two.)
Hard99A machine learning engineer is training a PyTorch model on a Databricks cluster and needs to distribute the training across multiple worker nodes. Which framework should be integrated natively within Databricks to handle this distributed deep learning workflow efficiently?
Medium100Refer to the exhibit. You are loading a model from the registry. What does the 'models:/MyModel/1' URI specifically represent?
Hard101A data scientist wants to record the exact library dependencies and a code snapshot alongside a model so that a reviewer can later restore the same environment and reproduce the training run. They are logging with MLflow on Databricks. Which practice best satisfies this requirement?
Easy102A data scientist is developing a model on Databricks and wants to use MLflow to compare multiple runs. They need to quickly identify the run with the lowest validation loss. Which MLflow UI feature allows them to sort and filter runs based on metrics?
Medium103An ML engineer is training an XGBoost model on Databricks and wants to leverage hyperparameter tuning using Hyperopt while automatically logging all trial parameters, metrics, and models to MLflow. Which built-in MLflow function should be used to achieve this automatic integration?
Medium104A team is developing a model on Databricks and wants to run an automated hyperparameter search over a scikit-learn pipeline. They need to try many parameter combinations in parallel across cluster workers while keeping every trial's parameters and metrics in MLflow. Which Databricks capability should they use to orchestrate the search?
Medium105A data scientist is using MLflow to track experiments on Databricks. They notice that some runs are missing the model artifact even though they called mlflow.sklearn.log_model(). What is the most likely cause?
Medium106When designing a model training pipeline, which TWO features of Unity Catalog best support compliance and model governance?
Hard107A data scientist is training a deep learning model on Databricks. They observe that the training process is significantly slower than expected. Upon inspection, they find that data loading from DBFS is the bottleneck. What is the most effective way to improve data loading speed for deep learning training on Databricks?
Medium108Refer to the exhibit. What happens to these logged metrics in MLflow when the training run completes?
Medium109A machine learning engineer needs to track hyperparameter tuning experiments in Databricks using MLflow. Which approach best ensures that model training runs are associated with the correct code version and environment settings?
MediumOther domains
All Databricks-ML-Pro exam domains
Frequently asked questions
- What does the Model Development domain cover on the Databricks-ML-Pro exam?
- You must be able to log experiments with MLflow, run scalable cross-validation on Spark, register and transition model versions, and set up drift monitoring. The single most important thing: know where MLflow artifacts are stored and how the Model Registry tracks lifecycle stages.
- How many questions are in this domain?
- This page lists all 109 Model Development questions in the Databricks-ML-Pro question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Model Development questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.