Databricks-ML-Assoc · domain
Model Development
Model Development on Databricks-ML-Assoc covers building and tracking models with MLflow and cluster tooling: experiment runs, parameters and metrics, model signatures, code_path packaging, custom containers, and real-time monitoring. Questions are scenario-based, asking you to pick the correct MLflow API, cluster configuration, or monitoring tool for a stated engineering goal.
Focused practice
Practice Model Development questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Model Development
Be able to log a model with the right MLflow parameters, configure a consistent cluster environment, and compare runs by a custom metric. The single most important thing: know exactly what each mlflow.log_model argument does and does not include.
Logging models with mlflow.log_model, including signature, input_example, and code_path arguments
Using Databricks custom containers or cluster libraries to pin consistent Python environments across nodes
Tracking and comparing runs in the MLflow experiment UI by custom metrics such as weighted_f1
Monitoring deep learning training loss and accuracy in real time with TensorBoard or MLflow metrics
Watch out for
Common Model Development exam traps
- ▸Assuming code_path bundles arbitrary dependencies; it captures code files, not installed libraries or the full environment.
- ▸Confusing cluster-scoped libraries with notebook-scoped installs, so nodes end up with mismatched Python packages.
- ▸Sorting MLflow runs by a default metric instead of the custom metric name, selecting the wrong best run.
Question index
All Model Development questions (85)
Click any question to see the full explanation, or start a practice session above.
A data scientist is using MLflow to log a model on Databricks. They want to ensure that the model can be loaded and used for inference in a different environment. Which two of the following are necessary components that must be included when logging the model to guarantee portability? (Choose two.)
Hard2A data scientist is using MLflow to track a hyperparameter tuning experiment with Spark MLlib's CrossValidator. They notice that each run in the MLflow UI shows only a single set of metrics, but they want to compare the performance of each hyperparameter combination across folds. What is the most effective way to log and visualize the per-combination and per-fold metrics in MLflow?
Hard3A data scientist is training a gradient boosting model using Spark MLlib on a large dataset in Databricks. They notice that the model's performance on a validation set is significantly worse than on the training set, and they suspect overfitting. They want to use MLflow to track hyperparameters and metrics to diagnose the issue. Which combination of MLflow logging practices will best help them identify overfitting across multiple runs?
Hard4A data scientist is using Databricks Feature Store to train a model. They define a feature table with a primary key and a timestamp key, and they want to ensure that when they create a training set, only the latest feature values as of each label event are used to avoid label leakage. Which Databricks Feature Store method should they call to create the training set with point-in-time correctness?
Hard5A data scientist is evaluating feature importance for a tree-based model trained on Databricks. They want to understand which features contribute most to the model's predictions. Which TWO methods are appropriate for extracting feature importance from a scikit-learn Random Forest model? (Choose two.)
Medium6When hyperparameter tuning using 'Hyperopt' on Databricks, what is the primary benefit of using the 'Trials' object?
Medium7A data scientist is tuning a scikit-learn GradientBoostingClassifier on Databricks. They use Hyperopt with the fmin function and the SparkTrials backend, but they notice that the best model returned by fmin is not identical to the model they get when they retrain with the same hyperparameters. They also observe that the logged metrics from each trial vary slightly even when the same hyperparameters are used. What is the most likely cause?
Medium8A data scientist is training a scikit-learn model on a large dataset using Databricks. They want to speed up hyperparameter tuning by running trials in parallel across a cluster. Which Databricks tool should they use?
Medium9A machine learning engineer is using MLflow to log a scikit-learn model. They call mlflow.sklearn.log_model(model, "model") without specifying a signature. When the model is later loaded for batch inference, the engineer observes that the model's predict method works correctly, but they cannot determine the expected input schema from the logged model. Which MLflow component is missing and would have provided this information?
Medium10When using MLflow to manage the lifecycle of a model in Databricks, why should you use the Model Registry instead of just saving model files to DBFS?
Medium11Which Databricks feature allows data scientists to automatically track parameters, metrics, and models during the training process?
Easy12Which THREE steps are essential for preparing a dataset for training using the Databricks Feature Store?
Medium13A machine learning engineer is using MLflow to log a model built with XGBoost. They want to ensure that the model's input schema is captured for validation during deployment. Which MLflow feature should they use?
Medium14What is the primary function of the 'Model Signature' in the context of Databricks MLflow?
Easy15A machine learning engineer is using MLflow to track experiments in Databricks. They want to record the hyperparameters used for each run so that they can compare runs later. Which MLflow method should they use to log a single hyperparameter?
Easy16A machine learning engineer is using MLflow on Databricks to track experiments for a fraud detection model. They notice that runs from two different team members are being logged into the same experiment, but the engineer wants to ensure that all runs from the current notebook session are automatically associated with a specific experiment. Which MLflow API call should the engineer use to set the active experiment for the current session?
Medium17When logging a model, what is the significance of the 'code_path' parameter in mlflow.log_model?
Medium18Which THREE features are provided by the MLflow Model Registry to support model governance and deployment?
Hard19A machine learning engineer is using MLflow to track a model training run on Databricks. They log a metric with mlflow.log_metric("accuracy", 0.95) and later want to retrieve it. They call mlflow.get_run(run_id) and access run.data.metrics. However, they find that the metrics dictionary is empty. What is the most likely reason?
Medium20A data scientist is training a machine learning model and wants to ensure that the code version, model parameters, and artifacts are all linked to a specific execution. Which Databricks component is designed specifically for this purpose?
Easy21A machine learning engineer is tuning a scikit-learn GradientBoostingClassifier on Databricks. They want to run 40 hyperparameter combinations, each trained on the full dataset, while keeping the driver free of model training work and collecting all results in a single MLflow parent run. Which approach should they use?
Medium22A data scientist is using Databricks Feature Store to create a training dataset for a model. They define a feature table with a primary key and a timestamp key. After creating the training set, they notice that some feature values are missing in the output. What is the most likely cause?
Easy23Which TWO of the following are common pitfalls when using 'Pandas UDFs' for distributed inference on large datasets in Databricks?
Hard24When developing a machine learning model on Databricks, why is it recommended to use 'mlflow.log_param' for tracking model configurations like learning rate?
Medium25When logging a model, you decide to store a 'data_version' tag. What is the benefit of this practice?
Easy26A data scientist is building a pipeline using MLflow. Which TWO of the following tasks should be performed to ensure experiment reproducibility and auditability?
Medium27What is the benefit of using the MLflow 'signature' when logging a model?
Easy28A data scientist is using Databricks Feature Store to build a training set for a model that predicts customer churn. The feature table contains a column `customer_id` and several features, and the label is stored in a separate Delta table. The data scientist wants to ensure that the exact same feature values used during training are available at inference time. Which approach correctly uses Databricks Feature Store to create the training set?
Medium29A data scientist is training a scikit-learn model on a Databricks cluster with autoscaling enabled. They observe that each epoch takes significantly longer than expected, and the Spark UI shows very low CPU utilization across the worker nodes. The dataset is small enough to fit in memory on a single node. What is the most likely cause of the slow training?
Medium30A data scientist is building a feature pipeline and wants to avoid recomputing expensive aggregations on every run. They need the computed feature table to be queryable by other notebooks and jobs, refreshed on a schedule, and stored in Delta Lake. Which Databricks capability should they use to define and materialize these features?
Easy31A machine learning engineer is using MLflow to track experiments for a model that uses a custom Python function to preprocess data. They want to ensure that the model can be deployed consistently across environments. Which MLflow component should they use to package the preprocessing logic along with the model?
Hard32Which THREE factors should be considered when selecting a model for deployment in a production Databricks environment?
Hard33A data scientist is training a model on a large dataset using the Databricks 'pandas_udf' functionality. They notice that the function is failing when processing specific partitions. What is the most likely cause?
Medium34Refer to the exhibit. A developer encounters this error while trying to register a model in the Unity Catalog. What does this error signify about the model deployment process?
Hard35When logging a PyTorch model to the MLflow Model Registry, which component must be explicitly defined to allow the model to be loaded in an environment where the original code structure might not exist?
Hard36When logging a machine learning model using MLflow, which component is required to capture the environment dependencies (such as library versions) to ensure the model can be reproduced in a different Databricks workspace?
Easy37A data scientist is training a scikit-learn model in a Databricks notebook and wants to automatically log parameters, metrics, and the model artifact to an MLflow experiment without writing explicit log calls. They have already installed the required libraries. Which approach should they use?
Medium38When logging a model using MLflow, what does the 'artifacts' parameter allow a user to include?
Easy39When performing hyperparameter tuning using Hyperopt with SparkTrials on Databricks, what is the primary advantage of using SparkTrials over the standard Trials object?
Medium40A data scientist is using Databricks to train a deep learning model. They need to monitor training loss and accuracy in real-time. Which tool is best suited for this task?
Medium41A data scientist is deploying a model to a Databricks Model Serving endpoint. They observe that the inference latency is high. What should they check first?
Medium42A data scientist is training a scikit-learn model on a Databricks cluster using MLflow. To enable automatic logging of parameters, metrics, and models, they call mlflow.sklearn.autolog() before fitting the model. After the run completes, they notice that the model artifact is stored in the run's artifact location but is not registered in the MLflow Model Registry. What is the most likely reason for the model not being registered?
Medium43Refer to the exhibit. What is the effect of using the 'registered_model_name' parameter in the 'log_model' function?
Medium44A data scientist is developing a model on Databricks and wants to use MLflow to track experiments. They create a new experiment using mlflow.create_experiment('my_experiment') and then run mlflow.start_run(). However, when they log parameters and metrics, they notice that the run is not associated with 'my_experiment' but with the default experiment. What is the most likely reason?
Easy45A team is transitioning their model development from a single notebook to a production-grade ML pipeline. Which Databricks feature should they use to manage and coordinate this end-to-end process?
Medium46A data scientist is using MLflow tracking on Databricks to log a model training run. They want to capture the model's hyperparameters, evaluation metrics, and the trained model artifact so that the run can be reproduced and the model can be deployed later. Which two MLflow API calls should they use to log the model artifact and its input/output schema? (Choose two.)
Medium47A data scientist is training a Random Forest model on a 500GB dataset using Databricks. They notice the training process is slow and memory-intensive on a single worker node. Which approach should they take to optimize the training process?
Medium48A machine learning engineer is using MLflow to track experiments on Databricks. They want to compare multiple runs and identify the best model based on a custom metric called 'weighted_f1'. They have logged this metric using mlflow.log_metric('weighted_f1', value) for each run. When viewing the experiment in the MLflow UI, they notice that the runs are not sorted by 'weighted_f1' and the metric does not appear in the runs table. What is the most likely cause?
Hard49A data scientist is using Databricks Feature Store to create a training dataset for a model. They define a feature table with a primary key and a timestamp column. They then create a training set using create_training_set with the feature table and a label DataFrame. They notice that the training set contains null values for some features, even though the feature table has no nulls. What is the most likely reason for the nulls in the training set?
Hard50A machine learning engineer is using Hyperopt with SparkTrials on Databricks to tune a gradient boosting model. They notice that the tuning process is taking longer than expected and want to optimize resource utilization. They have a cluster with 8 worker nodes. Which configuration should they adjust to allow SparkTrials to run more trials in parallel?
Medium51When developing a model, a data scientist uses the MLflow 'pyfunc' flavor to wrap their model. What is the primary benefit of using this approach?
Medium52Which TWO of the following practices are recommended when performing feature engineering on Databricks using Feature Store to ensure consistency between training and inference?
Medium53A machine learning engineer is using MLflow to log a model trained with a custom Python function. They want to ensure that the model can be loaded and served in a different environment. Which two of the following must be included when logging the model to ensure portability? (Choose two.)
Hard54Which of the following describes the purpose of a 'Validation Set' in the model development cycle?
Easy55When logging a model using MLflow in Databricks, which component is required to capture the environment dependencies so that the model can be accurately reproduced on a different cluster?
Easy56A machine learning engineer is training a model on a Databricks cluster and wants the training code to run inside a container that they control, with the same Python libraries available on every node. They also want the environment recorded with the MLflow run for reproducibility. Which Databricks capability should they use?
Hard57A machine learning engineer is building a scikit-learn model with hyperparameter tuning on Databricks. They want each trial to be tracked as a nested run under a single parent run in MLflow so that all trials are grouped together and the best parameters can be compared easily. Which MLflow API call should they use to start each trial run so that it is nested under the currently active run?
Medium58When logging a model in Databricks, why is it recommended to specify the `pip_requirements` or `conda_env` explicitly instead of relying on the environment's current state?
Medium59Refer to the exhibit. A model was successfully logged but fails to load in a production environment with the error shown in the exhibit. What is the most likely cause of this issue?
Hard60A machine learning engineer is using MLflow to track experiments. They call mlflow.start_run() and then log a model with mlflow.sklearn.log_model(). After the run completes, they notice that the model artifact is stored in the run's artifact location, but the run's source version and git commit are not captured. They are running from a Databricks notebook with Git integration enabled. Which action will ensure that the Git commit hash and source version are automatically logged to the MLflow run?
Hard61A data scientist is using MLflow to track experiments. They notice that all runs from a particular notebook are being logged to the default experiment instead of the experiment they intended to use. They have already called mlflow.start_run() without specifying an experiment ID. What is the most likely cause?
Medium62A machine learning engineer is using MLflow to track experiments. They want to compare multiple runs and identify the run that produced the best model based on a custom metric called 'weighted_f1'. They have logged this metric using mlflow.log_metric. Which MLflow UI feature allows them to sort and filter runs by this metric to quickly find the best run?
Hard63A data scientist is training a deep learning model on Databricks using Horovod for distributed training. They find that the model is converging slowly. What is the most likely cause related to the distributed configuration?
Medium64A data scientist is using Spark MLlib and wants to perform feature scaling on a large dataset. Which transformer should they use within a Pipeline to ensure that the scaling logic is correctly applied during both training and inference?
Medium65A data scientist is using Databricks Feature Store to create a feature table for a machine learning model. They want to ensure that the features used during training are consistent with those used during inference. Which Databricks Feature Store capability should they use?
Easy66Refer to the exhibit. A data scientist is logging a model to MLflow. Why is including an 'input_example' highly recommended in this specific code snippet?
Hard67Which technique is most effective for handling high-cardinality categorical features when training a tree-based model on Databricks?
Medium68What is the primary role of the 'Model Signature' in a Databricks ML lifecycle?
Medium69A data scientist trains a model with MLflow on Databricks and logs it using mlflow.sklearn.log_model with a registered_model_name. A downstream batch job loads the model by stage using models:/<name>/Staging. Weeks later, a colleague promotes a new version to Staging and the batch job's predictions change without any code deployment. Which change best prevents unintended downstream consumption while keeping promotion workflows intact?
Hard70When performing hyperparameter tuning using Hyperopt on Databricks, which function is primarily used to distribute the training task across the cluster?
Easy71Which feature in Databricks allows a data scientist to version and manage the lifecycle of machine learning models in a centralized repository?
Easy72Which of the following is the recommended workflow for developing a scalable model on Databricks?
Easy73A machine learning engineer needs to track model experiments in Databricks and wants to ensure that model artifacts are versioned automatically. Which approach best leverages Databricks-native capabilities for this requirement?
Medium74A data scientist is training a Random Forest model on a 500GB dataset using Databricks. They notice that the model training process is consistently running out of memory on the driver node. Which approach should be taken to resolve this memory bottleneck while maintaining model performance?
Medium75When using MLflow to track experiments, what happens if you invoke mlflow.end_run() inside a nested loop when the parent run is already active?
Hard76Refer to the exhibit. Why is including the `signature` and `input_example` in the `log_model` call considered a professional best practice?
Hard77A data scientist is training a linear regression model using scikit-learn on Databricks. They want to track the model's hyperparameters, such as fit_intercept and normalize, in MLflow. Which MLflow API call should they use to log these hyperparameters?
Easy78A data scientist is using Hyperopt with SparkTrials on a Databricks cluster to tune a scikit-learn model. They set max_evals=100 and parallelism=4. After the tuning completes, they notice that some trials failed due to memory errors on the workers. What is the most likely cause of these failures?
Medium79A data scientist is working in a Databricks notebook and wants to view the results of their MLflow runs, including metrics and parameters, directly within the notebook. Which MLflow function should they use?
Easy80A machine learning engineer is using MLflow to log a model built with XGBoost. They call mlflow.xgboost.log_model(xgb_model, 'model') and then attempt to load the model in a different environment using mlflow.pyfunc.load_model('runs:/<run_id>/model'). The load fails with an error about missing dependencies. Which action should they take to ensure the model can be loaded in the new environment?
Medium81Which TWO actions should be taken to ensure reproducibility of a Databricks ML model experiment?
Medium82A model has been trained on a dataset containing categorical features with high cardinality. Which technique is most effective for preparing these features for a linear model in a Databricks environment?
Medium83Refer to the exhibit. Why is providing an 'input_example' highly recommended during the model logging process?
Medium84A machine learning engineer is using MLflow to log a model trained with XGBoost. They want to ensure that the model can be loaded and used for inference in a different environment without requiring the original training environment. Which MLflow feature allows the model to capture its dependencies and environment?
Hard85A data scientist is preparing a dataset for training a model on Databricks. They want to split the data into training and testing sets, and they need to ensure that the split is reproducible across different runs. Which PySpark method should they use to split the DataFrame with a fixed random seed?
EasyOther domains
All Databricks-ML-Assoc exam domains
Frequently asked questions
- What does the Model Development domain cover on the Databricks-ML-Assoc exam?
- Be able to log a model with the right MLflow parameters, configure a consistent cluster environment, and compare runs by a custom metric. The single most important thing: know exactly what each mlflow.log_model argument does and does not include.
- How many questions are in this domain?
- This page lists all 85 Model Development questions in the Databricks-ML-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Model Development questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.