Be able to log a model with the right MLflow parameters, configure a consistent cluster environment, and compare runs by a custom metric. The single most important thing: know exactly what each mlflow.log_model argument does and does not include.
Start practicing
Model Development — choose a session length
Free · No account required
Domain overview
Model Development on Databricks-ML-Assoc covers building and tracking models with MLflow and cluster tooling: experiment runs, parameters and metrics, model signatures, code_path packaging, custom containers, and real-time monitoring. Questions are scenario-based, asking you to pick the correct MLflow API, cluster configuration, or monitoring tool for a stated engineering goal.
Exam objectives
Logging models with mlflow.log_model, including signature, input_example, and code_path arguments
Using Databricks custom containers or cluster libraries to pin consistent Python environments across nodes
Tracking and comparing runs in the MLflow experiment UI by custom metrics such as weighted_f1
Monitoring deep learning training loss and accuracy in real time with TensorBoard or MLflow metrics
Assuming code_path bundles arbitrary dependencies; it captures code files, not installed libraries or the full environment.
Confusing cluster-scoped libraries with notebook-scoped installs, so nodes end up with mismatched Python packages.
Sorting MLflow runs by a default metric instead of the custom metric name, selecting the wrong best run.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A data scientist is building a pipeline using MLflow. Which TWO of the following tasks should be performed to ensure experiment reproducibility and auditability?
2When performing hyperparameter tuning using Hyperopt with SparkTrials on Databricks, what is the primary advantage of using SparkTrials over the standard Trials object?
3A model has been trained on a dataset containing categorical features with high cardinality. Which technique is most effective for preparing these features for a linear model in a Databricks environment?
4Which Databricks feature allows data scientists to automatically track parameters, metrics, and models during the training process?
5Refer to the exhibit. Why is including the `signature` and `input_example` in the `log_model` call considered a professional best practice?
6A data scientist is using Databricks to train a deep learning model. They need to monitor training loss and accuracy in real-time. Which tool is best suited for this task?
7Which of the following describes the purpose of a 'Validation Set' in the model development cycle?
8A team is transitioning their model development from a single notebook to a production-grade ML pipeline. Which Databricks feature should they use to manage and coordinate this end-to-end process?
9When logging a model in Databricks, why is it recommended to specify the `pip_requirements` or `conda_env` explicitly instead of relying on the environment's current state?
10A data scientist is training a Random Forest model on a 500GB dataset using Databricks. They notice the training process is slow and memory-intensive on a single worker node. Which approach should they take to optimize the training process?
11Which TWO of the following practices are recommended when performing feature engineering on Databricks using Feature Store to ensure consistency between training and inference?
12When logging a model using MLflow in Databricks, which component is required to capture the environment dependencies so that the model can be accurately reproduced on a different cluster?
13Refer to the exhibit. A data scientist is logging a model to MLflow. Why is including an 'input_example' highly recommended in this specific code snippet?
14Which THREE steps are essential for preparing a dataset for training using the Databricks Feature Store?
15When hyperparameter tuning using 'Hyperopt' on Databricks, what is the primary benefit of using the 'Trials' object?
16A data scientist is using Spark MLlib and wants to perform feature scaling on a large dataset. Which transformer should they use within a Pipeline to ensure that the scaling logic is correctly applied during both training and inference?
17When developing a machine learning model on Databricks, why is it recommended to use 'mlflow.log_param' for tracking model configurations like learning rate?
18Refer to the exhibit. A developer encounters this error while trying to register a model in the Unity Catalog. What does this error signify about the model deployment process?
19When using MLflow to manage the lifecycle of a model in Databricks, why should you use the Model Registry instead of just saving model files to DBFS?
20Which TWO of the following are common pitfalls when using 'Pandas UDFs' for distributed inference on large datasets in Databricks?
21Which of the following is the recommended workflow for developing a scalable model on Databricks?
22Which TWO actions should be taken to ensure reproducibility of a Databricks ML model experiment?
23What is the primary role of the 'Model Signature' in a Databricks ML lifecycle?
24A machine learning engineer needs to track model experiments in Databricks and wants to ensure that model artifacts are versioned automatically. Which approach best leverages Databricks-native capabilities for this requirement?
25When performing hyperparameter tuning using Hyperopt on Databricks, which function is primarily used to distribute the training task across the cluster?
26When using MLflow to track experiments, what happens if you invoke mlflow.end_run() inside a nested loop when the parent run is already active?
27Which technique is most effective for handling high-cardinality categorical features when training a tree-based model on Databricks?
28When logging a PyTorch model to the MLflow Model Registry, which component must be explicitly defined to allow the model to be loaded in an environment where the original code structure might not exist?
29When logging a model, what is the significance of the 'code_path' parameter in mlflow.log_model?
30Refer to the exhibit. Why is providing an 'input_example' highly recommended during the model logging process?
31What is the primary function of the 'Model Signature' in the context of Databricks MLflow?
32A data scientist is training a machine learning model and wants to ensure that the code version, model parameters, and artifacts are all linked to a specific execution. Which Databricks component is designed specifically for this purpose?
33When developing a model, a data scientist uses the MLflow 'pyfunc' flavor to wrap their model. What is the primary benefit of using this approach?
34Refer to the exhibit. A model was successfully logged but fails to load in a production environment with the error shown in the exhibit. What is the most likely cause of this issue?
35When logging a model, you decide to store a 'data_version' tag. What is the benefit of this practice?
36A data scientist is training a Random Forest model on a 500GB dataset using Databricks. They notice that the model training process is consistently running out of memory on the driver node. Which approach should be taken to resolve this memory bottleneck while maintaining model performance?
37When logging a machine learning model using MLflow, which component is required to capture the environment dependencies (such as library versions) to ensure the model can be reproduced in a different Databricks workspace?
38Which feature in Databricks allows a data scientist to version and manage the lifecycle of machine learning models in a centralized repository?
39A data scientist is training a model on a large dataset using the Databricks 'pandas_udf' functionality. They notice that the function is failing when processing specific partitions. What is the most likely cause?
40Which THREE features are provided by the MLflow Model Registry to support model governance and deployment?
41What is the benefit of using the MLflow 'signature' when logging a model?
42A data scientist is deploying a model to a Databricks Model Serving endpoint. They observe that the inference latency is high. What should they check first?
43Which THREE factors should be considered when selecting a model for deployment in a production Databricks environment?
44Refer to the exhibit. What is the effect of using the 'registered_model_name' parameter in the 'log_model' function?
45A data scientist is training a deep learning model on Databricks using Horovod for distributed training. They find that the model is converging slowly. What is the most likely cause related to the distributed configuration?
46When logging a model using MLflow, what does the 'artifacts' parameter allow a user to include?
47A data scientist is training a scikit-learn model in a Databricks notebook and wants to automatically log parameters, metrics, and the model artifact to an MLflow experiment without writing explicit log calls. They have already installed the required libraries. Which approach should they use?
48A machine learning engineer is using MLflow to track experiments for a model that uses a custom Python function to preprocess data. They want to ensure that the model can be deployed consistently across environments. Which MLflow component should they use to package the preprocessing logic along with the model?
49A data scientist is working in a Databricks notebook and wants to view the results of their MLflow runs, including metrics and parameters, directly within the notebook. Which MLflow function should they use?
50A machine learning engineer is building a scikit-learn model with hyperparameter tuning on Databricks. They want each trial to be tracked as a nested run under a single parent run in MLflow so that all trials are grouped together and the best parameters can be compared easily. Which MLflow API call should they use to start each trial run so that it is nested under the currently active run?
51A data scientist is using Databricks Feature Store to train a model. They define a feature table with a primary key and a timestamp key, and they want to ensure that when they create a training set, only the latest feature values as of each label event are used to avoid label leakage. Which Databricks Feature Store method should they call to create the training set with point-in-time correctness?
52A machine learning engineer is using MLflow to log a model built with XGBoost. They want to ensure that the model's input schema is captured for validation during deployment. Which MLflow feature should they use?
53A data scientist is training a scikit-learn model on a large dataset using Databricks. They want to speed up hyperparameter tuning by running trials in parallel across a cluster. Which Databricks tool should they use?
54A data scientist is using MLflow tracking on Databricks to log a model training run. They want to capture the model's hyperparameters, evaluation metrics, and the trained model artifact so that the run can be reproduced and the model can be deployed later. Which two MLflow API calls should they use to log the model artifact and its input/output schema? (Choose two.)
55A machine learning engineer is using MLflow to log a scikit-learn model. They call mlflow.sklearn.log_model(model, "model") without specifying a signature. When the model is later loaded for batch inference, the engineer observes that the model's predict method works correctly, but they cannot determine the expected input schema from the logged model. Which MLflow component is missing and would have provided this information?
56A data scientist is training a scikit-learn model on a Databricks cluster using MLflow. To enable automatic logging of parameters, metrics, and models, they call mlflow.sklearn.autolog() before fitting the model. After the run completes, they notice that the model artifact is stored in the run's artifact location but is not registered in the MLflow Model Registry. What is the most likely reason for the model not being registered?
57A machine learning engineer is using MLflow to track experiments on Databricks. They want to compare multiple runs and identify the best model based on a custom metric called 'weighted_f1'. They have logged this metric using mlflow.log_metric('weighted_f1', value) for each run. When viewing the experiment in the MLflow UI, they notice that the runs are not sorted by 'weighted_f1' and the metric does not appear in the runs table. What is the most likely cause?
58A data scientist is training a linear regression model using scikit-learn on Databricks. They want to track the model's hyperparameters, such as fit_intercept and normalize, in MLflow. Which MLflow API call should they use to log these hyperparameters?
59A data scientist is preparing a dataset for training a model on Databricks. They want to split the data into training and testing sets, and they need to ensure that the split is reproducible across different runs. Which PySpark method should they use to split the DataFrame with a fixed random seed?
60A data scientist is developing a model on Databricks and wants to use MLflow to track experiments. They create a new experiment using mlflow.create_experiment('my_experiment') and then run mlflow.start_run(). However, when they log parameters and metrics, they notice that the run is not associated with 'my_experiment' but with the default experiment. What is the most likely reason?
61A data scientist is training a scikit-learn model on a Databricks cluster with autoscaling enabled. They observe that each epoch takes significantly longer than expected, and the Spark UI shows very low CPU utilization across the worker nodes. The dataset is small enough to fit in memory on a single node. What is the most likely cause of the slow training?
62A data scientist is using MLflow to track experiments. They notice that all runs from a particular notebook are being logged to the default experiment instead of the experiment they intended to use. They have already called mlflow.start_run() without specifying an experiment ID. What is the most likely cause?
63A data scientist is using Databricks Feature Store to build a training set for a model that predicts customer churn. The feature table contains a column `customer_id` and several features, and the label is stored in a separate Delta table. The data scientist wants to ensure that the exact same feature values used during training are available at inference time. Which approach correctly uses Databricks Feature Store to create the training set?
64A machine learning engineer is using MLflow to log a model built with XGBoost. They call mlflow.xgboost.log_model(xgb_model, 'model') and then attempt to load the model in a different environment using mlflow.pyfunc.load_model('runs:/<run_id>/model'). The load fails with an error about missing dependencies. Which action should they take to ensure the model can be loaded in the new environment?
65A machine learning engineer is using MLflow to track experiments. They want to compare multiple runs and identify the run that produced the best model based on a custom metric called 'weighted_f1'. They have logged this metric using mlflow.log_metric. Which MLflow UI feature allows them to sort and filter runs by this metric to quickly find the best run?
66A machine learning engineer is using MLflow to log a model trained with XGBoost. They want to ensure that the model can be loaded and used for inference in a different environment without requiring the original training environment. Which MLflow feature allows the model to capture its dependencies and environment?
67A machine learning engineer is using MLflow to track experiments in Databricks. They want to record the hyperparameters used for each run so that they can compare runs later. Which MLflow method should they use to log a single hyperparameter?
68A data scientist is using MLflow to log a model on Databricks. They want to ensure that the model can be loaded and used for inference in a different environment. Which two of the following are necessary components that must be included when logging the model to guarantee portability? (Choose two.)
69A machine learning engineer is using MLflow on Databricks to track experiments for a fraud detection model. They notice that runs from two different team members are being logged into the same experiment, but the engineer wants to ensure that all runs from the current notebook session are automatically associated with a specific experiment. Which MLflow API call should the engineer use to set the active experiment for the current session?
70A machine learning engineer is tuning a scikit-learn GradientBoostingClassifier on Databricks. They want to run 40 hyperparameter combinations, each trained on the full dataset, while keeping the driver free of model training work and collecting all results in a single MLflow parent run. Which approach should they use?
71A data scientist is training a gradient boosting model using Spark MLlib on a large dataset in Databricks. They notice that the model's performance on a validation set is significantly worse than on the training set, and they suspect overfitting. They want to use MLflow to track hyperparameters and metrics to diagnose the issue. Which combination of MLflow logging practices will best help them identify overfitting across multiple runs?
72A data scientist is evaluating feature importance for a tree-based model trained on Databricks. They want to understand which features contribute most to the model's predictions. Which TWO methods are appropriate for extracting feature importance from a scikit-learn Random Forest model? (Choose two.)
73A data scientist trains a model with MLflow on Databricks and logs it using mlflow.sklearn.log_model with a registered_model_name. A downstream batch job loads the model by stage using models:/<name>/Staging. Weeks later, a colleague promotes a new version to Staging and the batch job's predictions change without any code deployment. Which change best prevents unintended downstream consumption while keeping promotion workflows intact?
74A data scientist is using Hyperopt with SparkTrials on a Databricks cluster to tune a scikit-learn model. They set max_evals=100 and parallelism=4. After the tuning completes, they notice that some trials failed due to memory errors on the workers. What is the most likely cause of these failures?
75A data scientist is building a feature pipeline and wants to avoid recomputing expensive aggregations on every run. They need the computed feature table to be queryable by other notebooks and jobs, refreshed on a schedule, and stored in Delta Lake. Which Databricks capability should they use to define and materialize these features?
76A data scientist is tuning a scikit-learn GradientBoostingClassifier on Databricks. They use Hyperopt with the fmin function and the SparkTrials backend, but they notice that the best model returned by fmin is not identical to the model they get when they retrain with the same hyperparameters. They also observe that the logged metrics from each trial vary slightly even when the same hyperparameters are used. What is the most likely cause?
77A machine learning engineer is using MLflow to log a model trained with a custom Python function. They want to ensure that the model can be loaded and served in a different environment. Which two of the following must be included when logging the model to ensure portability? (Choose two.)
78A data scientist is using Databricks Feature Store to create a training dataset for a model. They define a feature table with a primary key and a timestamp key. After creating the training set, they notice that some feature values are missing in the output. What is the most likely cause?
79A machine learning engineer is using MLflow to track experiments. They call mlflow.start_run() and then log a model with mlflow.sklearn.log_model(). After the run completes, they notice that the model artifact is stored in the run's artifact location, but the run's source version and git commit are not captured. They are running from a Databricks notebook with Git integration enabled. Which action will ensure that the Git commit hash and source version are automatically logged to the MLflow run?
80A machine learning engineer is using Hyperopt with SparkTrials on Databricks to tune a gradient boosting model. They notice that the tuning process is taking longer than expected and want to optimize resource utilization. They have a cluster with 8 worker nodes. Which configuration should they adjust to allow SparkTrials to run more trials in parallel?
81A data scientist is using Databricks Feature Store to create a feature table for a machine learning model. They want to ensure that the features used during training are consistent with those used during inference. Which Databricks Feature Store capability should they use?
82A data scientist is using MLflow to track a hyperparameter tuning experiment with Spark MLlib's CrossValidator. They notice that each run in the MLflow UI shows only a single set of metrics, but they want to compare the performance of each hyperparameter combination across folds. What is the most effective way to log and visualize the per-combination and per-fold metrics in MLflow?
83A machine learning engineer is using MLflow to track a model training run on Databricks. They log a metric with mlflow.log_metric("accuracy", 0.95) and later want to retrieve it. They call mlflow.get_run(run_id) and access run.data.metrics. However, they find that the metrics dictionary is empty. What is the most likely reason?
84A machine learning engineer is training a model on a Databricks cluster and wants the training code to run inside a container that they control, with the same Python libraries available on every node. They also want the environment recorded with the MLflow run for reproducibility. Which Databricks capability should they use?
85A data scientist is using Databricks Feature Store to create a training dataset for a model. They define a feature table with a primary key and a timestamp column. They then create a training set using create_training_set with the feature table and a label DataFrame. They notice that the training set contains null values for some features, even though the feature table has no nulls. What is the most likely reason for the nulls in the training set?
Be able to log a model with the right MLflow parameters, configure a consistent cluster environment, and compare runs by a custom metric. The single most important thing: know exactly what each mlflow.log_model argument does and does not include.
The Courseiva Databricks-ML-Assoc question bank contains 85 questions in the Model Development domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Model Development domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included