Courseiva

Databricks-ML-Assoc · domain

Model Development

Model Development on Databricks-ML-Assoc covers building and tracking models with MLflow and cluster tooling: experiment runs, parameters and metrics, model signatures, code_path packaging, custom containers, and real-time monitoring. Questions are scenario-based, asking you to pick the correct MLflow API, cluster configuration, or monitoring tool for a stated engineering goal.

85 questions20 easy43 medium22 hard

Focused practice

Practice Model Development questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Model Development

Be able to log a model with the right MLflow parameters, configure a consistent cluster environment, and compare runs by a custom metric. The single most important thing: know exactly what each mlflow.log_model argument does and does not include.

Logging models with mlflow.log_model, including signature, input_example, and code_path arguments

Using Databricks custom containers or cluster libraries to pin consistent Python environments across nodes

Tracking and comparing runs in the MLflow experiment UI by custom metrics such as weighted_f1

Monitoring deep learning training loss and accuracy in real time with TensorBoard or MLflow metrics

Watch out for

Common Model Development exam traps

  • ▸Assuming code_path bundles arbitrary dependencies; it captures code files, not installed libraries or the full environment.
  • ▸Confusing cluster-scoped libraries with notebook-scoped installs, so nodes end up with mismatched Python packages.
  • ▸Sorting MLflow runs by a default metric instead of the custom metric name, selecting the wrong best run.

Question index

All Model Development questions (85)

Click any question to see the full explanation, or start a practice session above.

1

A data scientist is using MLflow to log a model on Databricks. They want to ensure that the model can be loaded and used for inference in a different environment. Which two of the following are necessary components that must be included when logging the model to guarantee portability? (Choose two.)

Hard
2

A data scientist is using MLflow to track a hyperparameter tuning experiment with Spark MLlib's CrossValidator. They notice that each run in the MLflow UI shows only a single set of metrics, but they want to compare the performance of each hyperparameter combination across folds. What is the most effective way to log and visualize the per-combination and per-fold metrics in MLflow?

Hard
3

A data scientist is training a gradient boosting model using Spark MLlib on a large dataset in Databricks. They notice that the model's performance on a validation set is significantly worse than on the training set, and they suspect overfitting. They want to use MLflow to track hyperparameters and metrics to diagnose the issue. Which combination of MLflow logging practices will best help them identify overfitting across multiple runs?

Hard
4

A data scientist is using Databricks Feature Store to train a model. They define a feature table with a primary key and a timestamp key, and they want to ensure that when they create a training set, only the latest feature values as of each label event are used to avoid label leakage. Which Databricks Feature Store method should they call to create the training set with point-in-time correctness?

Hard
5

A data scientist is evaluating feature importance for a tree-based model trained on Databricks. They want to understand which features contribute most to the model's predictions. Which TWO methods are appropriate for extracting feature importance from a scikit-learn Random Forest model? (Choose two.)

Medium
6

When hyperparameter tuning using 'Hyperopt' on Databricks, what is the primary benefit of using the 'Trials' object?

Medium
7

A data scientist is tuning a scikit-learn GradientBoostingClassifier on Databricks. They use Hyperopt with the fmin function and the SparkTrials backend, but they notice that the best model returned by fmin is not identical to the model they get when they retrain with the same hyperparameters. They also observe that the logged metrics from each trial vary slightly even when the same hyperparameters are used. What is the most likely cause?

Medium
8

A data scientist is training a scikit-learn model on a large dataset using Databricks. They want to speed up hyperparameter tuning by running trials in parallel across a cluster. Which Databricks tool should they use?

Medium
9

A machine learning engineer is using MLflow to log a scikit-learn model. They call mlflow.sklearn.log_model(model, "model") without specifying a signature. When the model is later loaded for batch inference, the engineer observes that the model's predict method works correctly, but they cannot determine the expected input schema from the logged model. Which MLflow component is missing and would have provided this information?

Medium
10

When using MLflow to manage the lifecycle of a model in Databricks, why should you use the Model Registry instead of just saving model files to DBFS?

Medium
11

Which Databricks feature allows data scientists to automatically track parameters, metrics, and models during the training process?

Easy
12

Which THREE steps are essential for preparing a dataset for training using the Databricks Feature Store?

Medium
13

A machine learning engineer is using MLflow to log a model built with XGBoost. They want to ensure that the model's input schema is captured for validation during deployment. Which MLflow feature should they use?

Medium
14

What is the primary function of the 'Model Signature' in the context of Databricks MLflow?

Easy
15

A machine learning engineer is using MLflow to track experiments in Databricks. They want to record the hyperparameters used for each run so that they can compare runs later. Which MLflow method should they use to log a single hyperparameter?

Easy
16

A machine learning engineer is using MLflow on Databricks to track experiments for a fraud detection model. They notice that runs from two different team members are being logged into the same experiment, but the engineer wants to ensure that all runs from the current notebook session are automatically associated with a specific experiment. Which MLflow API call should the engineer use to set the active experiment for the current session?

Medium
17

When logging a model, what is the significance of the 'code_path' parameter in mlflow.log_model?

Medium
18

Which THREE features are provided by the MLflow Model Registry to support model governance and deployment?

Hard
19

A machine learning engineer is using MLflow to track a model training run on Databricks. They log a metric with mlflow.log_metric("accuracy", 0.95) and later want to retrieve it. They call mlflow.get_run(run_id) and access run.data.metrics. However, they find that the metrics dictionary is empty. What is the most likely reason?

Medium
20

A data scientist is training a machine learning model and wants to ensure that the code version, model parameters, and artifacts are all linked to a specific execution. Which Databricks component is designed specifically for this purpose?

Easy
21

A machine learning engineer is tuning a scikit-learn GradientBoostingClassifier on Databricks. They want to run 40 hyperparameter combinations, each trained on the full dataset, while keeping the driver free of model training work and collecting all results in a single MLflow parent run. Which approach should they use?

Medium
22

A data scientist is using Databricks Feature Store to create a training dataset for a model. They define a feature table with a primary key and a timestamp key. After creating the training set, they notice that some feature values are missing in the output. What is the most likely cause?

Easy
23

Which TWO of the following are common pitfalls when using 'Pandas UDFs' for distributed inference on large datasets in Databricks?

Hard
24

When developing a machine learning model on Databricks, why is it recommended to use 'mlflow.log_param' for tracking model configurations like learning rate?

Medium
25

When logging a model, you decide to store a 'data_version' tag. What is the benefit of this practice?

Easy
26

A data scientist is building a pipeline using MLflow. Which TWO of the following tasks should be performed to ensure experiment reproducibility and auditability?

Medium
27

What is the benefit of using the MLflow 'signature' when logging a model?

Easy
28

A data scientist is using Databricks Feature Store to build a training set for a model that predicts customer churn. The feature table contains a column `customer_id` and several features, and the label is stored in a separate Delta table. The data scientist wants to ensure that the exact same feature values used during training are available at inference time. Which approach correctly uses Databricks Feature Store to create the training set?

Medium
29

A data scientist is training a scikit-learn model on a Databricks cluster with autoscaling enabled. They observe that each epoch takes significantly longer than expected, and the Spark UI shows very low CPU utilization across the worker nodes. The dataset is small enough to fit in memory on a single node. What is the most likely cause of the slow training?

Medium
30

A data scientist is building a feature pipeline and wants to avoid recomputing expensive aggregations on every run. They need the computed feature table to be queryable by other notebooks and jobs, refreshed on a schedule, and stored in Delta Lake. Which Databricks capability should they use to define and materialize these features?

Easy
31

A machine learning engineer is using MLflow to track experiments for a model that uses a custom Python function to preprocess data. They want to ensure that the model can be deployed consistently across environments. Which MLflow component should they use to package the preprocessing logic along with the model?

Hard
32

Which THREE factors should be considered when selecting a model for deployment in a production Databricks environment?

Hard
33

A data scientist is training a model on a large dataset using the Databricks 'pandas_udf' functionality. They notice that the function is failing when processing specific partitions. What is the most likely cause?

Medium
34

Refer to the exhibit. A developer encounters this error while trying to register a model in the Unity Catalog. What does this error signify about the model deployment process?

Hard
35

When logging a PyTorch model to the MLflow Model Registry, which component must be explicitly defined to allow the model to be loaded in an environment where the original code structure might not exist?

Hard
36

When logging a machine learning model using MLflow, which component is required to capture the environment dependencies (such as library versions) to ensure the model can be reproduced in a different Databricks workspace?

Easy
37

A data scientist is training a scikit-learn model in a Databricks notebook and wants to automatically log parameters, metrics, and the model artifact to an MLflow experiment without writing explicit log calls. They have already installed the required libraries. Which approach should they use?

Medium
38

When logging a model using MLflow, what does the 'artifacts' parameter allow a user to include?

Easy
39

When performing hyperparameter tuning using Hyperopt with SparkTrials on Databricks, what is the primary advantage of using SparkTrials over the standard Trials object?

Medium
40

A data scientist is using Databricks to train a deep learning model. They need to monitor training loss and accuracy in real-time. Which tool is best suited for this task?

Medium
41

A data scientist is deploying a model to a Databricks Model Serving endpoint. They observe that the inference latency is high. What should they check first?

Medium
42

A data scientist is training a scikit-learn model on a Databricks cluster using MLflow. To enable automatic logging of parameters, metrics, and models, they call mlflow.sklearn.autolog() before fitting the model. After the run completes, they notice that the model artifact is stored in the run's artifact location but is not registered in the MLflow Model Registry. What is the most likely reason for the model not being registered?

Medium
43

Refer to the exhibit. What is the effect of using the 'registered_model_name' parameter in the 'log_model' function?

Medium
44

A data scientist is developing a model on Databricks and wants to use MLflow to track experiments. They create a new experiment using mlflow.create_experiment('my_experiment') and then run mlflow.start_run(). However, when they log parameters and metrics, they notice that the run is not associated with 'my_experiment' but with the default experiment. What is the most likely reason?

Easy
45

A team is transitioning their model development from a single notebook to a production-grade ML pipeline. Which Databricks feature should they use to manage and coordinate this end-to-end process?

Medium
46

A data scientist is using MLflow tracking on Databricks to log a model training run. They want to capture the model's hyperparameters, evaluation metrics, and the trained model artifact so that the run can be reproduced and the model can be deployed later. Which two MLflow API calls should they use to log the model artifact and its input/output schema? (Choose two.)

Medium
47

A data scientist is training a Random Forest model on a 500GB dataset using Databricks. They notice the training process is slow and memory-intensive on a single worker node. Which approach should they take to optimize the training process?

Medium
48

A machine learning engineer is using MLflow to track experiments on Databricks. They want to compare multiple runs and identify the best model based on a custom metric called 'weighted_f1'. They have logged this metric using mlflow.log_metric('weighted_f1', value) for each run. When viewing the experiment in the MLflow UI, they notice that the runs are not sorted by 'weighted_f1' and the metric does not appear in the runs table. What is the most likely cause?

Hard
49

A data scientist is using Databricks Feature Store to create a training dataset for a model. They define a feature table with a primary key and a timestamp column. They then create a training set using create_training_set with the feature table and a label DataFrame. They notice that the training set contains null values for some features, even though the feature table has no nulls. What is the most likely reason for the nulls in the training set?

Hard
50

A machine learning engineer is using Hyperopt with SparkTrials on Databricks to tune a gradient boosting model. They notice that the tuning process is taking longer than expected and want to optimize resource utilization. They have a cluster with 8 worker nodes. Which configuration should they adjust to allow SparkTrials to run more trials in parallel?

Medium
51

When developing a model, a data scientist uses the MLflow 'pyfunc' flavor to wrap their model. What is the primary benefit of using this approach?

Medium
52

Which TWO of the following practices are recommended when performing feature engineering on Databricks using Feature Store to ensure consistency between training and inference?

Medium
53

A machine learning engineer is using MLflow to log a model trained with a custom Python function. They want to ensure that the model can be loaded and served in a different environment. Which two of the following must be included when logging the model to ensure portability? (Choose two.)

Hard
54

Which of the following describes the purpose of a 'Validation Set' in the model development cycle?

Easy
55

When logging a model using MLflow in Databricks, which component is required to capture the environment dependencies so that the model can be accurately reproduced on a different cluster?

Easy
56

A machine learning engineer is training a model on a Databricks cluster and wants the training code to run inside a container that they control, with the same Python libraries available on every node. They also want the environment recorded with the MLflow run for reproducibility. Which Databricks capability should they use?

Hard
57

A machine learning engineer is building a scikit-learn model with hyperparameter tuning on Databricks. They want each trial to be tracked as a nested run under a single parent run in MLflow so that all trials are grouped together and the best parameters can be compared easily. Which MLflow API call should they use to start each trial run so that it is nested under the currently active run?

Medium
58

When logging a model in Databricks, why is it recommended to specify the `pip_requirements` or `conda_env` explicitly instead of relying on the environment's current state?

Medium
59

Refer to the exhibit. A model was successfully logged but fails to load in a production environment with the error shown in the exhibit. What is the most likely cause of this issue?

Hard
60

A machine learning engineer is using MLflow to track experiments. They call mlflow.start_run() and then log a model with mlflow.sklearn.log_model(). After the run completes, they notice that the model artifact is stored in the run's artifact location, but the run's source version and git commit are not captured. They are running from a Databricks notebook with Git integration enabled. Which action will ensure that the Git commit hash and source version are automatically logged to the MLflow run?

Hard
61

A data scientist is using MLflow to track experiments. They notice that all runs from a particular notebook are being logged to the default experiment instead of the experiment they intended to use. They have already called mlflow.start_run() without specifying an experiment ID. What is the most likely cause?

Medium
62

A machine learning engineer is using MLflow to track experiments. They want to compare multiple runs and identify the run that produced the best model based on a custom metric called 'weighted_f1'. They have logged this metric using mlflow.log_metric. Which MLflow UI feature allows them to sort and filter runs by this metric to quickly find the best run?

Hard
63

A data scientist is training a deep learning model on Databricks using Horovod for distributed training. They find that the model is converging slowly. What is the most likely cause related to the distributed configuration?

Medium
64

A data scientist is using Spark MLlib and wants to perform feature scaling on a large dataset. Which transformer should they use within a Pipeline to ensure that the scaling logic is correctly applied during both training and inference?

Medium
65

A data scientist is using Databricks Feature Store to create a feature table for a machine learning model. They want to ensure that the features used during training are consistent with those used during inference. Which Databricks Feature Store capability should they use?

Easy
66

Refer to the exhibit. A data scientist is logging a model to MLflow. Why is including an 'input_example' highly recommended in this specific code snippet?

Hard
67

Which technique is most effective for handling high-cardinality categorical features when training a tree-based model on Databricks?

Medium
68

What is the primary role of the 'Model Signature' in a Databricks ML lifecycle?

Medium
69

A data scientist trains a model with MLflow on Databricks and logs it using mlflow.sklearn.log_model with a registered_model_name. A downstream batch job loads the model by stage using models:/<name>/Staging. Weeks later, a colleague promotes a new version to Staging and the batch job's predictions change without any code deployment. Which change best prevents unintended downstream consumption while keeping promotion workflows intact?

Hard
70

When performing hyperparameter tuning using Hyperopt on Databricks, which function is primarily used to distribute the training task across the cluster?

Easy
71

Which feature in Databricks allows a data scientist to version and manage the lifecycle of machine learning models in a centralized repository?

Easy
72

Which of the following is the recommended workflow for developing a scalable model on Databricks?

Easy
73

A machine learning engineer needs to track model experiments in Databricks and wants to ensure that model artifacts are versioned automatically. Which approach best leverages Databricks-native capabilities for this requirement?

Medium
74

A data scientist is training a Random Forest model on a 500GB dataset using Databricks. They notice that the model training process is consistently running out of memory on the driver node. Which approach should be taken to resolve this memory bottleneck while maintaining model performance?

Medium
75

When using MLflow to track experiments, what happens if you invoke mlflow.end_run() inside a nested loop when the parent run is already active?

Hard
76

Refer to the exhibit. Why is including the `signature` and `input_example` in the `log_model` call considered a professional best practice?

Hard
77

A data scientist is training a linear regression model using scikit-learn on Databricks. They want to track the model's hyperparameters, such as fit_intercept and normalize, in MLflow. Which MLflow API call should they use to log these hyperparameters?

Easy
78

A data scientist is using Hyperopt with SparkTrials on a Databricks cluster to tune a scikit-learn model. They set max_evals=100 and parallelism=4. After the tuning completes, they notice that some trials failed due to memory errors on the workers. What is the most likely cause of these failures?

Medium
79

A data scientist is working in a Databricks notebook and wants to view the results of their MLflow runs, including metrics and parameters, directly within the notebook. Which MLflow function should they use?

Easy
80

A machine learning engineer is using MLflow to log a model built with XGBoost. They call mlflow.xgboost.log_model(xgb_model, 'model') and then attempt to load the model in a different environment using mlflow.pyfunc.load_model('runs:/<run_id>/model'). The load fails with an error about missing dependencies. Which action should they take to ensure the model can be loaded in the new environment?

Medium
81

Which TWO actions should be taken to ensure reproducibility of a Databricks ML model experiment?

Medium
82

A model has been trained on a dataset containing categorical features with high cardinality. Which technique is most effective for preparing these features for a linear model in a Databricks environment?

Medium
83

Refer to the exhibit. Why is providing an 'input_example' highly recommended during the model logging process?

Medium
84

A machine learning engineer is using MLflow to log a model trained with XGBoost. They want to ensure that the model can be loaded and used for inference in a different environment without requiring the original training environment. Which MLflow feature allows the model to capture its dependencies and environment?

Hard
85

A data scientist is preparing a dataset for training a model on Databricks. They want to split the data into training and testing sets, and they need to ensure that the split is reproducible across different runs. Which PySpark method should they use to split the DataFrame with a fixed random seed?

Easy

Frequently asked questions

What does the Model Development domain cover on the Databricks-ML-Assoc exam?
Be able to log a model with the right MLflow parameters, configure a consistent cluster environment, and compare runs by a custom metric. The single most important thing: know exactly what each mlflow.log_model argument does and does not include.
How many questions are in this domain?
This page lists all 85 Model Development questions in the Databricks-ML-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Model Development questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
databricks-ml-associate DATABRICKS-ML-ASSOCIATE ml assoc model development Practice Questions