Courseiva

Databricks-ML-Assoc · domain

scenario questions

Practise Databricks Certified Machine Learning Associate scenario questions practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.

319 questions61 easy159 medium99 hard

Focused practice

Practice scenario questions questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about scenario questions

scenario questions questions test whether you can apply the concept in context, not just recognise a definition.

How the topic appears in realistic exam-style scenarios.

Which detail in the question changes the correct answer.

How to eliminate plausible but wrong options.

How to connect the question back to the wider exam objective.

Watch out for

Common scenario questions exam traps

  • ▸Answering from memory before reading the full scenario.
  • ▸Missing a constraint such as cost, availability, security, scope or command context.
  • ▸Choosing a broad answer when the question asks for the most specific fix.
  • ▸Ignoring why the wrong options are tempting.

Question index

All scenario questions questions (319)

Click any question to see the full explanation, or start a practice session above.

1

A team registered a model in Unity Catalog and wants to serve it on a Databricks Model Serving endpoint. During deployment, the build fails because the model's conda environment references a private Python package hosted on an internal PyPI mirror that the serving build cannot reach. Which approach resolves the deployment?

Hard
2

A data scientist is using MLflow Tracking to log experiments. They want to compare multiple runs of a scikit-learn model and identify the run with the lowest RMSE. Which MLflow feature should they use?

Easy
3

An ML engineer is building an automated CI/CD pipeline that, after validating a new model version, must programmatically move it to the `champion` alias in Unity Catalog so the production Model Serving endpoint begins using it. The pipeline runs in a Databricks job using a service principal. Which action accomplishes the promotion programmatically?

Hard
4

You are tracking a deep learning model experiment using MLflow. You want to ensure that the model architecture and all hyperparameters are easily reproducible. Which approach is best practice?

Medium
5

A data scientist is using MLflow to log a custom PyTorch model in a Databricks notebook. They want to register the model in the Databricks Model Registry and later serve it with MLflow model serving. Which function should they call within their MLflow run to log the model with the necessary signature and dependencies?

Medium
6

A data scientist is using MLflow to log a model on Databricks. They want to ensure that the model can be loaded and used for inference in a different environment. Which two of the following are necessary components that must be included when logging the model to guarantee portability? (Choose two.)

Hard
7

A data scientist is using MLflow to track a hyperparameter tuning experiment with Spark MLlib's CrossValidator. They notice that each run in the MLflow UI shows only a single set of metrics, but they want to compare the performance of each hyperparameter combination across folds. What is the most effective way to log and visualize the per-combination and per-fold metrics in MLflow?

Hard
8

Refer to the exhibit. A data scientist receives this error while trying to load a model. What is the most likely cause of this failure in the workflow?

Medium
9

A team deploys an MLflow pyfunc model to a Databricks Model Serving endpoint. During pre-deployment testing they call the endpoint with a small batch of records and receive an HTTP 400 error stating the request payload does not match the model signature. The model was logged with an inferred signature from a pandas DataFrame. Which action most directly resolves the mismatch?

Hard
10

A machine learning engineer is troubleshooting a Model Registry issue where models are not being transitioned correctly. Which TWO actions should the engineer take to ensure proper governance and automated testing in the Registry?

Hard
11

A data scientist is training a gradient boosting model using Spark MLlib on a large dataset in Databricks. They notice that the model's performance on a validation set is significantly worse than on the training set, and they suspect overfitting. They want to use MLflow to track hyperparameters and metrics to diagnose the issue. Which combination of MLflow logging practices will best help them identify overfitting across multiple runs?

Hard
12

A data scientist is using Databricks Feature Store to train a model. They define a feature table with a primary key and a timestamp key, and they want to ensure that when they create a training set, only the latest feature values as of each label event are used to avoid label leakage. Which Databricks Feature Store method should they call to create the training set with point-in-time correctness?

Hard
13

Which THREE actions are essential when preparing a machine learning model for deployment using the Databricks Model Registry?

Medium
14

Which Databricks feature allows you to manage the lifecycle of a model, including transitions from 'Staging' to 'Production'?

Easy
15

A data scientist is evaluating feature importance for a tree-based model trained on Databricks. They want to understand which features contribute most to the model's predictions. Which TWO methods are appropriate for extracting feature importance from a scikit-learn Random Forest model? (Choose two.)

Medium
16

When hyperparameter tuning using 'Hyperopt' on Databricks, what is the primary benefit of using the 'Trials' object?

Medium
17

An ML engineer has a scikit-learn model trained locally and wants to log it to MLflow with a signature so that Databricks Model Serving can enforce input schema validation. The engineer calls mlflow.sklearn.log_model(model, 'model') but does not use infer_signature. What is the most accurate consequence when the model is later served on Databricks Model Serving?

Medium
18

A data science team is transitioning from local model development to Databricks. They want to ensure their models are portable across different Databricks workspaces. What is the recommended practice for managing models in this environment?

Medium
19

A data scientist trains a scikit-learn model with MLflow tracking in a Databricks notebook. They call mlflow.sklearn.log_model(model, 'model') but later find that the Registered Model's schema shows no input signature, preventing automatic schema enforcement during serving. What should they have done to capture the model signature?

Medium
20

A data scientist is tuning a scikit-learn GradientBoostingClassifier on Databricks. They use Hyperopt with the fmin function and the SparkTrials backend, but they notice that the best model returned by fmin is not identical to the model they get when they retrain with the same hyperparameters. They also observe that the logged metrics from each trial vary slightly even when the same hyperparameters are used. What is the most likely cause?

Medium
21

Which TWO factors should be considered when choosing between Batch Inference and Real-time Inference for a model in Databricks?

Medium
22

A team is transitioning from local model development to production in Databricks. Which THREE practices should be implemented to ensure a successful MLOps workflow?

Medium
23

Which Databricks tool is primarily used for organizing and documenting experiments, tracking parameters, and versioning models during the machine learning lifecycle?

Medium
24

A data scientist is training a scikit-learn model on a large dataset using Databricks. They want to speed up hyperparameter tuning by running trials in parallel across a cluster. Which Databricks tool should they use?

Medium
25

A data scientist has trained a model and wants other team members to serve it through a Databricks Model Serving endpoint. The model must be discoverable by a three-level namespace and governed by Unity Catalog. What should the data scientist do first?

Easy
26

A machine learning engineer maintains a Databricks Model Serving endpoint named `recommendations-endpoint` serving a registered model `ml_team.recommender_model` at version 7. The team wants to route 10% of live traffic to a newly registered candidate version 8 while keeping version 7 serving the remaining 90%, without altering the existing client request URL. Which approach should the engineer take?

Hard
27

An ML engineer is developing a model in a Databricks notebook. They want to log the model and its dependencies to MLflow, ensuring that the exact library versions used during training are captured. They also need to register the model in the Databricks Model Registry. Which approach should they use?

Medium
28

When using the Databricks Feature Store for real-time inference, why is it recommended to use a specific online store (e.g., Cosmos DB) instead of querying the offline store?

Hard
29

An ML engineer is building a recurring batch inference pipeline in Databricks. The model is registered in Unity Catalog as a model version and must always use the version currently tagged as 'champion', which is reassigned after each retraining run. The engineer wants the inference notebook to resolve this alias at runtime rather than hard-coding a version number. Which approach should the engineer use in the notebook?

Medium
30

What is the purpose of the 'Champion' model version in the Model Registry?

Medium
31

An ML engineer is building a Databricks Job that trains a model with scikit-learn and needs to capture hyperparameters, evaluation metrics, and the resulting model artifact for each run. The team wants to compare runs visually in the workspace and later promote the best model to the Model Registry. Which MLflow capability should the engineer use to record these per-run details?

Medium
32

A team is preparing to deploy a model to a Databricks Model Serving endpoint that must scale down to zero replicas when idle yet still serve bursty traffic with acceptable cold-start latency. They also need to capture the request and response payloads for later monitoring. Which two endpoint settings or features should they configure? (Choose two.)

Hard
33

An MLOps engineer is responsible for a Databricks Model Serving endpoint that serves a mission-critical pricing model. The team wants an automated safeguard that detects when live input feature distributions drift away from the training distribution and triggers a retraining workflow, without modifying the model artifact itself. Which Databricks capability should be configured?

Hard
34

A machine learning engineer is using MLflow to log a scikit-learn model. They call mlflow.sklearn.log_model(model, "model") without specifying a signature. When the model is later loaded for batch inference, the engineer observes that the model's predict method works correctly, but they cannot determine the expected input schema from the logged model. Which MLflow component is missing and would have provided this information?

Medium
35

When using MLflow to manage the lifecycle of a model in Databricks, why should you use the Model Registry instead of just saving model files to DBFS?

Medium
36

Which Databricks feature allows data scientists to automatically track parameters, metrics, and models during the training process?

Easy
37

Which THREE steps are essential for preparing a dataset for training using the Databricks Feature Store?

Medium
38

A data scientist is building a training pipeline where raw event data lands in a Delta table. They need to transform the data, train a model, and register it to the Databricks Model Registry. The pipeline must run daily on a schedule and send an email alert if training fails. Which Databricks construct should they use to orchestrate the entire workflow?

Medium
39

A machine learning engineer is using MLflow to log a model built with XGBoost. They want to ensure that the model's input schema is captured for validation during deployment. Which MLflow feature should they use?

Medium
40

A machine learning engineer wants to register a model in the Databricks Model Registry using MLflow. They have already trained a model and logged it with MLflow. Which method should they use to register the model programmatically?

Easy
41

A data science team is using Databricks Repos to manage a machine learning project. They want to ensure that their notebooks and supporting modules are version-controlled and that they can collaborate without overwriting each other's changes. Which TWO practices should they follow? (Choose two.)

Hard
42

When deploying a model to Databricks Model Serving, what is the recommended way to handle sensitive credentials like database connection strings?

Easy
43

A data scientist has trained a scikit-learn model and logged it with MLflow. They now want to register the model in the Databricks Model Registry and transition it to the 'Production' stage. Which sequence of MLflow API calls should they use?

Medium
44

When sharing a Databricks ML experiment with another team, what is the best way to ensure they can reproduce your results exactly?

Medium
45

Which Databricks artifact should be used to encapsulate a model, its environment dependencies, and the required code to ensure consistent model behavior across different deployment environments?

Hard
46

Your team is experiencing 'data drift' in production where the model's accuracy drops over time. What is the most recommended Databricks-native approach to address this?

Medium
47

A team wants to compare multiple model runs for a fraud detection project, view their metrics side by side, and identify which run produced the best area under the ROC curve. They have already logged each run with MLflow. Which Databricks capability should they use to perform this comparison?

Easy
48

A machine learning team is using Databricks Feature Store to serve features for online inference. They need to ensure that the online store remains consistent with the offline store and supports low-latency lookups. Which two practices should they follow? (Choose two.)

Hard
49

A data scientist has trained a model using scikit-learn and wants to log it to MLflow for deployment. They need to ensure that the model can be served with the correct dependencies. Which MLflow function should they use to log the model?

Easy
50

What is the primary function of the 'Model Signature' in the context of Databricks MLflow?

Easy
51

A machine learning engineer is using MLflow to track experiments in Databricks. They want to record the hyperparameters used for each run so that they can compare runs later. Which MLflow method should they use to log a single hyperparameter?

Easy
52

An ML engineer deployed a model to a Databricks Model Serving endpoint and enabled inference tables. After a week, the team notices that some requests returned HTTP 200 but the corresponding rows in the inference table show null prediction values. They need to diagnose why predictions are missing for those requests. Which explanation is most consistent with this symptom?

Hard
53

A team is preparing to deploy a registered MLflow model to a Databricks Model Serving endpoint. They want to capture every request and response for later monitoring and debugging, and they also want the endpoint to remain available during a rolling model version update. (Choose two.)

Medium
54

A machine learning engineer is using MLflow on Databricks to track experiments for a fraud detection model. They notice that runs from two different team members are being logged into the same experiment, but the engineer wants to ensure that all runs from the current notebook session are automatically associated with a specific experiment. Which MLflow API call should the engineer use to set the active experiment for the current session?

Medium
55

A data scientist is using MLflow to track experiments on Databricks. They want to log a custom metric that is computed during model training and later compare it across runs using the MLflow UI. Which MLflow API call should they use?

Medium
56

When designing a production-grade machine learning workflow, which THREE of the following are necessary to ensure the pipeline is observable and recoverable?

Medium
57

An ML engineer is setting up a Databricks Job to retrain a production model nightly. The job must run only when upstream data validation succeeds, notify the team on failure, and avoid retraining when the input data has not changed. Which TWO capabilities of Databricks Jobs directly support these requirements? (Choose two.)

Medium
58

When logging a model, what is the significance of the 'code_path' parameter in mlflow.log_model?

Medium
59

A machine learning engineer has a model registered in Unity Catalog and wants to expose it as a REST API so an external application can send JSON payloads and receive predictions. The team has no existing serving infrastructure. Which Databricks feature should be used to create this API?

Easy
60

What is the primary benefit of using 'AutoML' in Databricks for a machine learning project?

Easy
61

A data scientist wants to compare the performance of three hyperparameter configurations for a Spark ML model. They need a central place to view metrics like RMSE and MAE across runs, and to filter runs by parameters. Which Databricks capability should they use?

Easy
62

A team is building a Feature Store in Databricks. What is the primary advantage of using the Feature Store for training models compared to using raw Delta tables?

Medium
63

A data scientist finishes a training notebook and wants to capture the source code revision and the Git repository URL on the MLflow run so reviewers can reproduce the exact code state. The repository is already connected to the Databricks workspace through Git integration. Which mechanism records this information automatically?

Easy
64

A machine learning engineer has registered a model in the Databricks Model Registry and wants to expose it as a REST API with automatic scaling and no server management. The model's Python dependencies are captured in a conda environment file logged with the run. Which Databricks capability should the engineer use to serve this model with minimal operational overhead?

Medium
65

A data scientist is tuning a scikit-learn random forest with Hyperopt in a Databricks notebook. Each trial trains on a 40 GB Delta table, and the scientist notices that every trial re-reads the full table from cloud storage, making the search slow. Which change best accelerates the hyperparameter search while preserving correctness?

Medium
66

Which THREE features are provided by the MLflow Model Registry to support model governance and deployment?

Hard
67

A data science team is deploying a model to a Databricks Model Serving endpoint. They want to enable inference logging to capture the input data and predictions for monitoring and debugging. Which TWO configurations are required to enable inference logging for a serving endpoint? (Choose two.)

Hard
68

An ML platform team is standardizing how models move from experimentation to production on Databricks. They want promotion decisions to be auditable and to prevent unvalidated models from serving live traffic. Which TWO practices align with Databricks Model Registry and Unity Catalog governance? (Choose two.)

Hard
69

A machine learning engineer is using MLflow to track a model training run on Databricks. They log a metric with mlflow.log_metric("accuracy", 0.95) and later want to retrieve it. They call mlflow.get_run(run_id) and access run.data.metrics. However, they find that the metrics dictionary is empty. What is the most likely reason?

Medium
70

Refer to the exhibit. What is the most likely cause of this error in a deployed MLflow model?

Hard
71

A data scientist is training a machine learning model and wants to ensure that the code version, model parameters, and artifacts are all linked to a specific execution. Which Databricks component is designed specifically for this purpose?

Easy
72

A machine learning engineer is tuning a scikit-learn GradientBoostingClassifier on Databricks. They want to run 40 hyperparameter combinations, each trained on the full dataset, while keeping the driver free of model training work and collecting all results in a single MLflow parent run. Which approach should they use?

Medium
73

Which approach is most efficient for deploying a high-throughput, low-latency model in Databricks?

Medium
74

A data scientist is using Databricks Feature Store to create a training dataset for a model. They define a feature table with a primary key and a timestamp key. After creating the training set, they notice that some feature values are missing in the output. What is the most likely cause?

Easy
75

Why should you use a dedicated Feature Store for your ML workflows instead of just storing features as Delta tables?

Hard
76

When designing a feature engineering pipeline in Databricks, why should you use the Feature Store instead of standard Delta tables?

Medium
77

Which Databricks feature allows you to monitor and manage the lineage of data from the source to the final model prediction?

Medium
78

Which TWO of the following are common pitfalls when using 'Pandas UDFs' for distributed inference on large datasets in Databricks?

Hard
79

Which of the following is an advantage of using Delta Lake for the data storage layer in an ML workflow?

Easy
80

A data scientist wants to track the performance of a model training run in Databricks. They use MLflow to log parameters and metrics. After the run, they need to view all runs for the experiment in a web-based interface. Which URL should they navigate to?

Easy
81

When developing a machine learning model on Databricks, why is it recommended to use 'mlflow.log_param' for tracking model configurations like learning rate?

Medium
82

When logging a model, you decide to store a 'data_version' tag. What is the benefit of this practice?

Easy
83

A machine learning engineer is using Databricks Feature Store to create a training dataset for a model that predicts customer lifetime value. The feature table includes a timestamp key. Which TWO statements are true regarding point-in-time correctness when creating the training set? (Choose two.)

Hard
84

Refer to the exhibit. A model is failing to deploy to Databricks Model Serving. The error log indicates 'Schema Mismatch'. Based on the JSON signature provided, what is the most likely cause of the deployment failure?

Hard
85

A machine learning engineer needs to track model parameters, metrics, and artifacts across distributed training runs executed on Databricks. Which component of Databricks Machine Learning should they use to manage and organize this experiment metadata?

Medium
86

A data scientist needs to perform hyperparameter tuning using Hyperopt. Which THREE components are essential to successfully implement an automated tuning run on a Databricks cluster?

Medium
87

A data scientist trains a scikit-learn model in a Databricks notebook and logs it with MLflow. She now needs to promote the exact model artifact to the Databricks Model Registry so that it can be served. Which MLflow API call accomplishes this promotion?

Medium
88

A data scientist needs to prepare a model for deployment in a highly regulated environment. Which TWO tasks must they complete to ensure the model meets auditability requirements?

Medium
89

A data scientist is building a pipeline using MLflow. Which TWO of the following tasks should be performed to ensure experiment reproducibility and auditability?

Medium
90

A machine learning engineer needs to capture every request and response payload sent to a Databricks Model Serving endpoint so that the team can later join predictions with ground-truth labels for monitoring. Which Databricks feature should they enable on the endpoint?

Medium
91

A data science team is using Databricks Jobs to orchestrate a machine learning pipeline. The pipeline includes a task that trains a model and a subsequent task that evaluates the model. The evaluation task must access the model version produced by the training task. Which mechanism should the team use to pass the model version between tasks?

Medium
92

A team deploys a model to a Databricks Model Serving endpoint and enables inference tables. After a week, they notice that the inference table contains request and response payloads but the payload columns are empty for many rows, while status codes are 200. What is the most likely explanation?

Hard
93

Refer to the exhibit. A machine learning engineer deployed an MLflow model to Databricks Model Serving, but inference requests are failing with the error shown in the exhibit. How should the engineer resolve this issue?

Hard
94

A data science team is preparing to deploy a custom machine learning model to Databricks Model Serving. Which TWO steps are required to ensure the model can successfully load and serve predictions using MLflow? Choose 2 answers.

Medium
95

Refer to the exhibit. A user attempts to load a model using the MLflow Python API, but the load fails. Based on the JSON snippet, what is the most likely issue?

Hard
96

What is the benefit of using the MLflow 'signature' when logging a model?

Easy
97

A data scientist is using Databricks Feature Store to build a training set for a model that predicts customer churn. The feature table contains a column `customer_id` and several features, and the label is stored in a separate Delta table. The data scientist wants to ensure that the exact same feature values used during training are available at inference time. Which approach correctly uses Databricks Feature Store to create the training set?

Medium
98

Which strategy is most effective for managing model drift in a production Databricks environment?

Hard
99

Which THREE of the following are benefits of using Databricks Workflows for ML model retraining?

Hard
100

A team maintains a Databricks Model Serving endpoint for a fraud model. Compliance requires that every request and response be logged to a Delta table for auditing and later analysis. The endpoint is already configured and serving traffic. What should the team do to capture this data with the least additional infrastructure?

Hard
101

A team is using Databricks Feature Store to create a feature table that will be used for both batch training and online inference. They need to ensure the feature table supports point-in-time lookups for training and low-latency reads for serving. Which TWO of the following statements are correct about meeting these requirements? (Choose two.)

Medium
102

What is the primary advantage of using Databricks Model Serving over deploying a model on a standalone web server?

Medium
103

A data scientist is training a scikit-learn model on a Databricks cluster with autoscaling enabled. They observe that each epoch takes significantly longer than expected, and the Spark UI shows very low CPU utilization across the worker nodes. The dataset is small enough to fit in memory on a single node. What is the most likely cause of the slow training?

Medium
104

A data scientist is building a feature pipeline and wants to avoid recomputing expensive aggregations on every run. They need the computed feature table to be queryable by other notebooks and jobs, refreshed on a schedule, and stored in Delta Lake. Which Databricks capability should they use to define and materialize these features?

Easy
105

A machine learning engineer is using MLflow to track experiments for a model that uses a custom Python function to preprocess data. They want to ensure that the model can be deployed consistently across environments. Which MLflow component should they use to package the preprocessing logic along with the model?

Hard
106

Which THREE factors should be considered when selecting a model for deployment in a production Databricks environment?

Hard
107

A data scientist wants to share an MLflow experiment with a teammate. What is the most direct way to ensure the teammate can access the metrics and parameters?

Medium
108

A team has an existing Databricks Model Serving endpoint serving `prod.ml.fraud_model` version 3. They register version 4, which uses a new feature set, and want to shift only 10% of traffic to version 4 while keeping version 3 for the rest. Their endpoint currently has a single served entity for version 3. What is the most appropriate approach?

Hard
109

An ML engineer needs to deploy a model for real-time inference with automatic scaling and a REST endpoint that requires token-based authentication. The model artifacts are already registered in the Databricks Model Registry. Which Databricks capability should be used?

Hard
110

Which of the following is a core characteristic of 'Model Serving' in Databricks?

Easy
111

A data scientist is training a model on a large dataset using the Databricks 'pandas_udf' functionality. They notice that the function is failing when processing specific partitions. What is the most likely cause?

Medium
112

An ML engineer is orchestrating an end-to-end machine learning pipeline using Databricks Jobs. The pipeline consists of data preparation, distributed hyperparameter tuning with Hyperopt, and model registration. The engineer needs to pass the best-performing model's run ID from the Hyperopt task to the subsequent model registration task dynamically. Which mechanism should the engineer use in Databricks Jobs to achieve this?

Medium
113

An ML engineer is using Databricks Jobs to orchestrate a machine learning pipeline that includes data ingestion, feature engineering, model training, and batch scoring. The engineer wants to ensure that the pipeline is reproducible, handles failures gracefully, and allows for easy debugging of individual tasks. Which TWO features of Databricks Jobs should the engineer leverage to meet these requirements? (Choose two.)

Hard
114

A data scientist is using Databricks AutoML to train a classification model on a dataset with a highly imbalanced target variable. They want to ensure the model evaluation focuses on the minority class. Which evaluation metric should they prioritize when interpreting AutoML results?

Hard
115

Refer to the exhibit. A developer encounters this error while trying to register a model in the Unity Catalog. What does this error signify about the model deployment process?

Hard
116

A data scientist is training a scikit-learn model on Databricks and wants to automatically log parameters, metrics, and models without writing explicit MLflow logging code. Which approach should they use?

Easy
117

When deploying a model to a production environment, why is it critical to create a dedicated 'staging' environment before the 'production' environment?

Medium
118

A data scientist deploys a model to a Databricks Model Serving endpoint and enables inference tables. After a week, they want to analyze prediction drift by joining the logged requests with ground-truth labels that arrive later. Which statement describes how they should access the inference table data for this analysis?

Hard
119

A machine learning engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model was logged with MLflow using the default signature and input example. After deployment, the engineer notices that the endpoint's REST API expects a JSON payload in a specific format. Which MLflow artifact is used by Model Serving to determine the expected input format for the endpoint?

Medium
120

When logging a PyTorch model to the MLflow Model Registry, which component must be explicitly defined to allow the model to be loaded in an environment where the original code structure might not exist?

Hard
121

A data scientist registers an MLflow model whose `conda.yaml` lists several Python packages. When they create a Databricks Model Serving endpoint from this model, the deployment fails during environment build. Which action is most likely to resolve the failure while preserving the model's dependency requirements?

Hard
122

A team observes that their Databricks Model Serving endpoint occasionally returns HTTP 429 responses during bursty traffic, even though the endpoint shows low average CPU utilization. They want to reduce these throttling errors without over-provisioning capacity. Which action is most appropriate?

Hard
123

A data scientist has a custom Python model wrapped in an MLflow pyfunc flavor and needs to serve it on Databricks Model Serving. The model's preprocessing requires a library that is not part of the default serving environment. What is the correct way to make that dependency available to the endpoint?

Easy
124

A data scientist is using MLflow to track experiments on Databricks. They notice that the metrics logged during a run are not appearing in the MLflow UI. The run is part of an experiment with many runs. What is the most likely cause for the missing metrics?

Hard
125

When configuring a Databricks Workflow to automate a machine learning pipeline, which TWO actions are necessary to ensure the pipeline is robust and manageable?

Hard
126

A data scientist wants to compare the accuracy, F1 score, and training duration of several model training runs side by side in a single table, and visually inspect how a hyperparameter affected the metric across runs. Which MLflow capability should be used?

Easy
127

Which Databricks feature should be used to provide a managed, secure, and scalable endpoint for real-time inference of models logged in the Model Registry?

Medium
128

A data scientist is training a model using MLflow on Databricks. They need to ensure that the model artifacts, environment dependencies, and signature are automatically captured to facilitate seamless deployment to Databricks Model Serving. Which command should they use within the training script?

Medium
129

A data scientist has registered a scikit-learn model in Unity Catalog as `ml_prod.churn.model_v3` and wants the Databricks Model Serving endpoint to automatically pick up newly registered model versions as they are promoted to the `champion` alias. Which configuration should the data scientist use when creating the serving endpoint?

Medium
130

Which workflow step is essential before promoting a model from 'Staging' to 'Production' in the MLflow Model Registry?

Medium
131

When logging a machine learning model using MLflow, which component is required to capture the environment dependencies (such as library versions) to ensure the model can be reproduced in a different Databricks workspace?

Easy
132

A data scientist is using Databricks Feature Store to create a feature table for a recommendation model. They want to ensure that the same feature computation logic is used both during training and at inference time to avoid training-serving skew. Which Feature Store capability directly addresses this requirement?

Easy
133

An ML engineer needs to give an external application a stable HTTPS URL to call a registered model served by Databricks Model Serving. The application must authenticate with a token and must not be able to modify the endpoint configuration. Which approach best meets these requirements?

Medium
134

A data scientist is training a scikit-learn model in a Databricks notebook and wants to automatically log parameters, metrics, and the model artifact to an MLflow experiment without writing explicit log calls. They have already installed the required libraries. Which approach should they use?

Medium
135

Refer to the exhibit. Why is the 'signature' parameter included in the log_model call?

Medium
136

A machine learning engineer is preparing to deploy a registered model to a Databricks Model Serving endpoint. Before creating the endpoint, the engineer wants to confirm the deployment prerequisites are satisfied. Which two conditions are required for a successful endpoint creation? (Choose two.)

Medium
137

When logging a model using MLflow, what does the 'artifacts' parameter allow a user to include?

Easy
138

A data scientist is training a model using MLflow on Databricks and needs to ensure that all parameters and metrics are logged for every training run. Which approach ensures the most reliable logging of artifacts and metrics during model training?

Medium
139

A team registers a model in Unity Catalog as main.ml.churn_model and wants production scoring jobs to always load the newest approved version without editing job code when a new version is promoted. The team uses the MLflow Python client inside a Databricks job. Which model URI should the scoring code use?

Hard
140

Refer to the exhibit. The logs indicate a persistent connection failure for a Databricks Model Serving endpoint. What is the most likely cause?

Hard
141

A team wants to route production traffic to a new model version while keeping risk low. They configure a Databricks Model Serving endpoint with two served entities: `champion` (entity_version 5) and `challenger` (entity_version 6). They want 95% of requests to hit `champion` and 5% to hit `challenger`. Which configuration accomplishes this?

Hard
142

A machine learning engineer registers a model in the Databricks Model Registry and wants to serve it with low-latency online inference. The model's Python dependencies include a custom private library that is not publicly available. Which deployment approach should the engineer use to ensure the private library is available at inference time?

Hard
143

Which component in Databricks is used to manage the lineage of machine learning data, ensuring that you can trace a model back to the exact version of the data it was trained on?

Medium
144

A machine learning engineer is building a training pipeline in Databricks. They want each run to record the exact Git commit hash, the versions of scikit-learn and MLflow used, and the input data path so the run can be reproduced later. Which MLflow tracking capability should they use to capture this information with the least custom code?

Medium
145

A data scientist wants to package a training script with its Python dependencies and parameters so that the same code can be rerun on a different Databricks cluster or shared with a colleague and reproduced exactly. Which MLflow component is designed for packaging and reproducing project code in this way?

Easy
146

When performing hyperparameter tuning using Hyperopt with SparkTrials on Databricks, what is the primary advantage of using SparkTrials over the standard Trials object?

Medium
147

A data scientist is using Databricks to train a deep learning model. They need to monitor training loss and accuracy in real-time. Which tool is best suited for this task?

Medium
148

A data scientist has trained a scikit-learn model and wants to log it to MLflow with a custom signature that includes input and output schema. Which MLflow method should they use to log the model along with the signature?

Medium
149

A data scientist is training a machine learning model on Databricks and needs to log parameters, metrics, and model artifacts. Which tracking component should be used to ensure the reproducibility of the experiment runs?

Medium
150

A machine learning engineer is using MLflow on Databricks to track an experiment. They want to record the model's hyperparameters and evaluation metrics, but they do not want to save the trained model artifact. Which MLflow API calls should they use?

Medium
151

A machine learning engineer is orchestrating a multi-step training workflow on Databricks using Databricks Jobs. The workflow includes data preprocessing, model training, and evaluation. The engineer needs to ensure that the evaluation step runs only if the training step completes successfully, and that the preprocessing step runs first. Which feature of Databricks Jobs should be used to define these dependencies?

Medium
152

A team has a Databricks Model Serving endpoint configured with scale-to-zero enabled and min_instances set to 0. During a load test, they observe that the first request after an idle period takes roughly 40 seconds while subsequent requests complete in under 200 milliseconds. They need to eliminate this cold-start latency for a customer-facing application without over-provisioning. Which configuration change best addresses the requirement?

Hard
153

Which Databricks ML component is best suited for managing access control for machine learning experiments and models across different teams?

Easy
154

Which THREE of the following are supported methods for serving machine learning models in Databricks?

Hard
155

What is the primary role of an 'MLflow Signature' during the model deployment phase?

Medium
156

A data scientist is deploying a model to a Databricks Model Serving endpoint. They observe that the inference latency is high. What should they check first?

Medium
157

A data scientist wants to test a newly registered model version interactively before promoting it to production. They need to send a sample request to the model and inspect the prediction and the model's input schema. Which Databricks feature should they use?

Easy
158

A data scientist is training a scikit-learn model on a Databricks cluster using MLflow. To enable automatic logging of parameters, metrics, and models, they call mlflow.sklearn.autolog() before fitting the model. After the run completes, they notice that the model artifact is stored in the run's artifact location but is not registered in the MLflow Model Registry. What is the most likely reason for the model not being registered?

Medium
159

Which technique should be used to prevent data leakage in ML workflows when performing cross-validation on time-series data?

Medium
160

A data scientist wants to track the progress of a training script that runs for several hours on a Databricks cluster. The script uses MLflow and needs to record metrics such as loss and accuracy at the end of each epoch so they can be visualized in real time. Which MLflow API call should be used inside the training loop?

Easy
161

Refer to the exhibit. What is the effect of using the 'registered_model_name' parameter in the 'log_model' function?

Medium
162

A data scientist is developing a model on Databricks and wants to use MLflow to track experiments. They create a new experiment using mlflow.create_experiment('my_experiment') and then run mlflow.start_run(). However, when they log parameters and metrics, they notice that the run is not associated with 'my_experiment' but with the default experiment. What is the most likely reason?

Easy
163

You are building a pipeline where a feature table must be updated daily. Which Databricks construct is the most appropriate for orchestrating this periodic feature engineering job?

Medium
164

A data scientist is working in a Databricks notebook and wants to use MLflow to log a trained scikit-learn model. They want to ensure that the model can be loaded later for inference. What is the correct MLflow function to log the model?

Easy
165

A data scientist is training a machine learning model on Databricks and needs to ensure that every experiment run is automatically tracked, including parameters, metrics, and model artifacts. Which component should the scientist use to achieve this with minimal code changes?

Medium
166

A data scientist has registered a model in the Databricks Model Registry. They want to transition the model from 'Staging' to 'Production' but need to ensure that only specific users can perform this transition. Which Databricks feature should they use to enforce this access control?

Medium
167

A team is transitioning their model development from a single notebook to a production-grade ML pipeline. Which Databricks feature should they use to manage and coordinate this end-to-end process?

Medium
168

A data scientist has registered a scikit-learn model in Unity Catalog and now wants to serve it behind a Databricks Model Serving endpoint. The model's MLflow signature records a pandas DataFrame input with three named columns. The team wants the endpoint to reject malformed requests automatically rather than silently scoring them. Which action should the data scientist take?

Medium
169

A team queries a Databricks Model Serving endpoint through the serving client and receives an error indicating the endpoint is not ready. They confirmed the endpoint exists. Which condition most directly explains why requests fail until it clears?

Medium
170

When deploying a model to a production endpoint, what is the best practice for handling dependencies?

Medium
171

A data scientist has trained a model and wants to deploy it for real-time inference with automatic scaling and without managing infrastructure. Which Databricks feature should they use?

Easy
172

A data scientist needs to track parameters, metrics, and model artifacts during training on Databricks. Which component is the primary tool for managing the entire lifecycle of these ML experiments?

Medium
173

Why is it important to use a 'Feature Store' rather than joining raw tables directly in the training notebook?

Medium
174

A data scientist is monitoring model drift in Databricks. Which TWO approaches are recommended to detect performance degradation in a production model?

Hard
175

A team is deploying a batch scoring pipeline that loads a registered MLflow model and runs predictions over a large Delta table using Spark. They want the scoring job to reuse the model's training-time preprocessing and to remain reproducible months later. Which two practices should they follow? (Choose two.)

Hard
176

A data scientist is using MLflow to track experiments on Databricks. They want to compare multiple runs and identify the best performing model based on a custom metric. Which TWO features of MLflow can be used to achieve this? (Choose two.)

Hard
177

A data scientist is using MLflow on Databricks to log a scikit-learn model. They call mlflow.sklearn.log_model(model, 'model') and then inspect the run. They notice the model artifact is stored, but the run does not appear in the Models page of the workspace. They did not call any model registration function. What is the most likely reason the model is not listed in the Models page?

Medium
178

Which Databricks component is specifically designed to manage the full lifecycle of machine learning models, including registration, versioning, and stage transitions?

Easy
179

What is the primary function of the 'Model Signatures' in MLflow?

Easy
180

When designing an ML workflow, what is the primary benefit of using MLflow Projects over executing raw scripts?

Easy
181

Which THREE factors should be considered when choosing the 'workload size' (e.g., Small, Medium, Large) for a Databricks Model Serving endpoint?

Hard
182

A data scientist is using MLflow tracking on Databricks to log a model training run. They want to capture the model's hyperparameters, evaluation metrics, and the trained model artifact so that the run can be reproduced and the model can be deployed later. Which two MLflow API calls should they use to log the model artifact and its input/output schema? (Choose two.)

Medium
183

A data scientist is training a Random Forest model on a 500GB dataset using Databricks. They notice the training process is slow and memory-intensive on a single worker node. Which approach should they take to optimize the training process?

Medium
184

A data scientist needs to deploy a model to Databricks Model Serving. Which component is strictly required to be logged in MLflow to enable the 'Model Serving' feature?

Medium
185

A machine learning engineer is using MLflow to track experiments on Databricks. They want to compare multiple runs and identify the best model based on a custom metric called 'weighted_f1'. They have logged this metric using mlflow.log_metric('weighted_f1', value) for each run. When viewing the experiment in the MLflow UI, they notice that the runs are not sorted by 'weighted_f1' and the metric does not appear in the runs table. What is the most likely cause?

Hard
186

You are training a scikit-learn model with MLflow in a Databricks notebook. The model's preprocessing includes a custom Python function that you wrote in the notebook. You need to register the model to the Databricks Model Registry and later deploy it with Model Serving, ensuring the preprocessing is applied automatically at inference. Which approach should you use?

Medium
187

A fraud detection team wants their Databricks Model Serving endpoint to log every request and response payload to a Unity Catalog Delta table so analysts can later join predictions with ground-truth labels. Which endpoint capability should they enable?

Medium
188

A platform team is rolling out a new Databricks Model Serving endpoint for a churn model. They must ensure the endpoint can be queried by an external application and that only authorized callers can invoke it. Which TWO actions should they take? (Choose two.)

Medium
189

A machine learning team is using Databricks Feature Store to serve features for a real-time model. They have a feature table that is updated daily with new data. To ensure the online store always has the latest feature values for low-latency inference, which approach should they take?

Hard
190

A data scientist is using Databricks Feature Store to create a training dataset for a model. They define a feature table with a primary key and a timestamp column. They then create a training set using create_training_set with the feature table and a label DataFrame. They notice that the training set contains null values for some features, even though the feature table has no nulls. What is the most likely reason for the nulls in the training set?

Hard
191

When evaluating a machine learning model, what is the main purpose of creating a separate evaluation dataset in Databricks?

Medium
192

A machine learning engineer is using Hyperopt with SparkTrials on Databricks to tune a gradient boosting model. They notice that the tuning process is taking longer than expected and want to optimize resource utilization. They have a cluster with 8 worker nodes. Which configuration should they adjust to allow SparkTrials to run more trials in parallel?

Medium
193

A team trains a model with Databricks Feature Store features and logs the training set using feature_lookups. At inference time, they want the model to automatically retrieve the same feature values from the online store so the serving endpoint does not require the caller to supply those features. What must the team do when logging the model so this automatic lookup works?

Hard
194

A data scientist is using MLflow Tracking in Databricks to compare multiple runs of a hyperparameter tuning experiment. They want to quickly identify the run with the lowest validation loss and then register that model version in the MLflow Model Registry. Which MLflow UI feature allows sorting runs by a specific metric to find the best run?

Easy
195

An ML engineer is configuring a Databricks Job to retrain a model daily. The job must run a notebook that reads from a feature table, trains a model, and registers it to the Model Registry. The engineer wants to ensure that the job fails immediately if the model's accuracy drops below a threshold. Which approach should they use?

Hard
196

Which TWO of the following are primary benefits of using the Databricks Feature Store for machine learning workflows?

Medium
197

An ML engineer is configuring a Databricks Job to automate nightly retraining of a model. The job must (1) run only after the upstream feature engineering job succeeds, and (2) notify the team via email if the training task fails. Which TWO configurations satisfy these requirements? (Choose two.)

Medium
198

An ML engineer runs an automated hyperparameter sweep with MLflow on a Databricks cluster. The sweep launches 200 runs, and the engineer wants to retrieve, in a notebook, the run ID of the single run that achieved the highest validation accuracy so it can be registered. Which approach correctly identifies that run?

Hard
199

A data scientist is using Databricks Feature Store to build training sets and wants to ensure the features used at training time are consistent with those served at inference time. Which TWO practices help guarantee this consistency? (Choose two.)

Medium
200

When developing a model, a data scientist uses the MLflow 'pyfunc' flavor to wrap their model. What is the primary benefit of using this approach?

Medium
201

Which TWO of the following practices are recommended when performing feature engineering on Databricks using Feature Store to ensure consistency between training and inference?

Medium
202

Which Databricks component should be used to track parameters, code versions, metrics, and output files when running machine learning experiments?

Easy
203

A machine learning engineer is training a model using Databricks AutoML. They notice that the generated notebook includes a step that uses Hyperopt for hyperparameter tuning, but the tuning process is taking too long. They want to reduce the search space without sacrificing model performance significantly. Which Hyperopt configuration change should they make?

Hard
204

Which TWO of the following are benefits of using the Databricks Feature Store for machine learning workflows?

Hard
205

A machine learning engineer is using MLflow to log a model trained with a custom Python function. They want to ensure that the model can be loaded and served in a different environment. Which two of the following must be included when logging the model to ensure portability? (Choose two.)

Hard
206

When using Databricks Model Serving, what is the primary benefit of using a 'Provisioned Throughput' endpoint over a 'Serverless' endpoint?

Medium
207

A machine learning team is using Databricks Jobs to orchestrate a multi-step ML pipeline: data ingestion, feature engineering, model training, and batch inference. They need to ensure that if the model training step fails, the batch inference step does not run, and that the entire pipeline can be retried from the failed step. Which Databricks Jobs feature should they use to achieve this?

Hard
208

A machine learning engineer is configuring a Databricks Model Serving endpoint for a model that requires GPU acceleration. They set the workload size to 'GPU_Medium' but the endpoint fails to deploy. Which of the following is the most likely cause of the failure?

Medium
209

An ML engineer registers a model in Unity Catalog and needs a stable pointer that always resolves to the version currently approved for production, without modifying calling code each time a new version is promoted. Which Model Registry feature should the engineer configure to provide this stable reference?

Hard
210

Refer to the exhibit. What is the impact of setting 'auto_capture_request_payload' to true?

Medium
211

Which Databricks feature provides a managed environment specifically optimized for machine learning libraries like TensorFlow, PyTorch, and XGBoost?

Easy
212

What is the primary benefit of using a 'Job Cluster' instead of an 'All-Purpose Cluster' for automated ML workflows?

Easy
213

Which of the following describes the purpose of a 'Validation Set' in the model development cycle?

Easy
214

A data scientist has registered a scikit-learn model in the Databricks Model Registry as `prod.churn_model`. The production endpoint serving this model must automatically roll back to the previously served version if the newly deployed version's error rate exceeds a threshold within one hour of deployment. Which Databricks feature should the data scientist configure to meet this requirement?

Medium
215

A data scientist needs to run the same feature-engineering notebook against three different parameter sets in a Databricks Job. She wants each parameter set to execute independently and in parallel, with separate logs. Which Job feature should she use?

Easy
216

A user is experiencing 'Out of Memory' (OOM) errors during the evaluation phase of a large XGBoost model on Databricks. What is the most effective way to address this while utilizing the distributed nature of Databricks?

Medium
217

When logging a model using MLflow in Databricks, which component is required to capture the environment dependencies so that the model can be accurately reproduced on a different cluster?

Easy
218

A team is standing up a real-time Databricks Model Serving endpoint for a fraud model. Requests will carry several numeric features, and the team wants the endpoint to reject malformed payloads with a clear client error rather than silently scoring them, and to avoid cold-start latency during business hours. Which two actions should the team take? (Choose two.)

Hard
219

A team serves a model with Databricks Model Serving and wants to send production traffic to a newly registered model version while keeping the ability to revert instantly if quality degrades. They prefer not to edit the endpoint configuration to switch versions. Which approach best fits this requirement?

Hard
220

When deploying a model using Model Serving, how does Databricks ensure that the environment remains consistent between the training workspace and the serving environment?

Medium
221

A machine learning engineer is training a model on a Databricks cluster and wants the training code to run inside a container that they control, with the same Python libraries available on every node. They also want the environment recorded with the MLflow run for reproducibility. Which Databricks capability should they use?

Hard
222

A machine learning engineer has an existing Databricks Model Serving endpoint named churn-endpoint serving version 3 of a model. The team has validated version 5 and wants to direct live traffic to it while keeping the deployment reversible if quality degrades. What is the most appropriate action?

Hard
223

A machine learning engineer is preparing to deploy a scikit-learn model as a Databricks Model Serving endpoint. The model expects a pandas DataFrame with specific column names and types. Which two actions should the engineer take to ensure the endpoint correctly validates and processes inference requests? (Choose two.)

Medium
224

A team operates a Databricks Model Serving endpoint with min_instances set to 0 and max_instances set to 4. During a nightly batch job, the endpoint receives a burst of requests and some clients observe elevated latency. The team wants to keep costs low during idle periods while reducing cold-start latency during bursts. Which configuration change best achieves this?

Hard
225

A team's Model Serving endpoint occasionally returns errors when the upstream feature store is slow. They want the endpoint to retry transient failures and reduce cold-start latency for bursty traffic. Which combination of endpoint settings best addresses both concerns?

Hard
226

A machine learning engineer is building a scikit-learn model with hyperparameter tuning on Databricks. They want each trial to be tracked as a nested run under a single parent run in MLflow so that all trials are grouped together and the best parameters can be compared easily. Which MLflow API call should they use to start each trial run so that it is nested under the currently active run?

Medium
227

When logging a model in Databricks, why is it recommended to specify the `pip_requirements` or `conda_env` explicitly instead of relying on the environment's current state?

Medium
228

A machine learning engineer is deploying an MLflow model to a Databricks Model Serving endpoint. The model was trained on a Spark DataFrame and logged with MLflow using the default signature detection. During testing, the endpoint returns predictions, but the engineer notices the input schema shown in the Serving UI does not match the actual DataFrame column types used during training. Which MLflow logging step should the engineer verify first to resolve this schema mismatch?

Medium
229

A data scientist is using MLflow to train a scikit-learn model on Databricks. They call mlflow.sklearn.autolog() before fitting the model. After the run completes, they need to retrieve the automatically logged model and load it for batch inference in a separate notebook. Which approach correctly retrieves the logged model for loading?

Medium
230

A data scientist is preparing a feature table in Databricks Feature Store. To ensure the feature table can be used for online inference with low latency, which step is mandatory?

Medium
231

A data engineer is working with a large Delta table and wants to optimize it for machine learning feature engineering. They frequently filter data by a column named 'event_date' and join on a column named 'user_id'. Which Delta Lake feature should they use to improve query performance?

Easy
232

Refer to the exhibit. A model was successfully logged but fails to load in a production environment with the error shown in the exhibit. What is the most likely cause of this issue?

Hard
233

A team owns a Databricks Model Serving endpoint that receives sporadic bursts of traffic. They want to reduce cold-start latency during bursts while keeping cost predictable, and they also need to capture the request payloads and predictions for later monitoring. Which two configuration choices should they make? (Choose two.)

Hard
234

A machine learning engineer is using MLflow to track experiments. They call mlflow.start_run() and then log a model with mlflow.sklearn.log_model(). After the run completes, they notice that the model artifact is stored in the run's artifact location, but the run's source version and git commit are not captured. They are running from a Databricks notebook with Git integration enabled. Which action will ensure that the Git commit hash and source version are automatically logged to the MLflow run?

Hard
235

A data scientist is training a model using MLflow on Databricks. They need to ensure that the model artifacts and metrics are logged automatically without adding manual logging code to the training script. Which approach should they use?

Medium
236

A machine learning engineer is using Databricks Feature Store to create a training dataset. They want to ensure that the features used during training are exactly the same as those served at inference time. Which Feature Store capability should they rely on?

Hard
237

A data scientist is using MLflow to track experiments. They notice that all runs from a particular notebook are being logged to the default experiment instead of the experiment they intended to use. They have already called mlflow.start_run() without specifying an experiment ID. What is the most likely cause?

Medium
238

A data scientist is comparing multiple hyperparameter configurations for a model and wants to view the resulting metrics side by side in a single interface, sort runs by accuracy, and drill into individual run details. Which MLflow component provides this capability?

Easy
239

Refer to the exhibit. A data scientist is attempting to deploy a model using the MLflow client. The error above occurs during the deployment script. What is the most likely cause of this failure?

Medium
240

Refer to the exhibit. A user attempts to update a model stage in the Model Registry and receives this error. What is the most appropriate action to resolve this?

Hard
241

A machine learning engineer is using MLflow to track experiments. They want to compare multiple runs and identify the run that produced the best model based on a custom metric called 'weighted_f1'. They have logged this metric using mlflow.log_metric. Which MLflow UI feature allows them to sort and filter runs by this metric to quickly find the best run?

Hard
242

A data scientist is training a deep learning model on Databricks using Horovod for distributed training. They find that the model is converging slowly. What is the most likely cause related to the distributed configuration?

Medium
243

A data scientist registered a model in Unity Catalog and now wants to serve it as a low-latency REST endpoint for an application. They need automatic scaling, a secure endpoint URL, and the ability to update the served model version without redeploying infrastructure. Which Databricks capability should they use?

Hard
244

A data scientist is using Spark MLlib and wants to perform feature scaling on a large dataset. Which transformer should they use within a Pipeline to ensure that the scaling logic is correctly applied during both training and inference?

Medium
245

A team has several model versions registered in Unity Catalog. They want to serve a specific version through a Databricks Model Serving endpoint and later promote a newer version without changing the endpoint URL used by applications. Which approach should they use?

Hard
246

An ML engineer trains a scikit-learn model on a Spark DataFrame in a Databricks notebook using MLflow autologging. The run logs parameters and metrics, but the engineer later opens the MLflow run and cannot find any input dataset lineage. Which action should the engineer take to ensure the training dataset is recorded with the run in the MLflow UI?

Medium
247

When a data scientist needs to share an MLflow experiment with a team member, what is the best practice for ensuring collaborative access within Databricks?

Medium
248

Refer to the exhibit. A data scientist is attempting to log a model to an S3 bucket via MLflow, but they receive the error shown. What is the most likely root cause?

Hard
249

A data scientist is using Databricks Feature Store to create a feature table for a machine learning model. They want to ensure that the features used during training are consistent with those used during inference. Which Databricks Feature Store capability should they use?

Easy
250

Refer to the exhibit. A Databricks job failed to start, returning the error shown. The job depends on MLflow for tracking. What is the most likely cause of this failure?

Hard
251

Refer to the exhibit. A data scientist is logging a model to MLflow. Why is including an 'input_example' highly recommended in this specific code snippet?

Hard
252

A data scientist wants to speed up the process of finding the optimal hyperparameters for a machine learning model. Which Databricks-supported library is optimized for distributed hyperparameter tuning?

Medium
253

A data scientist is using Databricks Feature Store to create a feature table for a recommendation model. They want to ensure that the feature table can be used for both training and batch scoring, and that it supports point-in-time correctness. Which two actions must they take when creating the feature table? (Choose two.)

Hard
254

What is the primary technical limitation when deploying an MLflow model that has custom Python dependencies not included in the standard Databricks Runtime?

Hard
255

A data scientist is using Databricks AutoML to train a classification model on a dataset with a binary target. They want to understand which features contributed most to the model's predictions and need a human-readable summary. Which AutoML output should they examine?

Medium
256

Which technique is most effective for handling high-cardinality categorical features when training a tree-based model on Databricks?

Medium
257

A team is preparing to deploy an MLflow model to a Databricks Model Serving endpoint and wants to diagnose why requests are failing before contacting support. Which two actions allow them to inspect the endpoint's behavior and errors? (Choose two.)

Medium
258

A machine learning engineer needs to track hyperparameter tuning runs and log artifacts using MLflow inside a Databricks Notebook. Which approach should be used to ensure runs are automatically nested under a parent run?

Medium
259

A machine learning engineer is deploying a model using MLflow Model Serving on Databricks. They want to ensure the endpoint can handle bursts of traffic and automatically scale. Which configuration should they set when creating the served model?

Medium
260

Which of the following is a primary benefit of using a model serving endpoint versus a batch inference job?

Medium
261

What is the primary role of the 'Model Signature' in a Databricks ML lifecycle?

Medium
262

An ML engineer registers a model to the Databricks Model Registry and moves it to the 'Production' stage. A downstream batch scoring job references the model as models:/churn_model/Production. A data scientist then registers a new model version and transitions it to 'Production'. What happens to the downstream batch scoring job the next time it runs?

Hard
263

A data scientist registers a model in the Databricks Model Registry and wants to record the model's intended input and output schema so that a serving endpoint can validate incoming requests. Which action accomplishes this when logging the model?

Easy
264

A machine learning engineer is using Databricks Feature Store to create a feature table for a model that predicts customer churn. The feature table includes customer demographics and transaction history. The engineer wants to ensure that the model can access the latest feature values during online inference. What should the engineer do?

Medium
265

A data scientist trains a model with MLflow on Databricks and logs it using mlflow.sklearn.log_model with a registered_model_name. A downstream batch job loads the model by stage using models:/<name>/Staging. Weeks later, a colleague promotes a new version to Staging and the batch job's predictions change without any code deployment. Which change best prevents unintended downstream consumption while keeping promotion workflows intact?

Hard
266

A data scientist has trained a scikit-learn model and wants to expose it for real-time inference through Databricks Model Serving. The model is currently logged as an MLflow run artifact but has not been registered anywhere. What must the data scientist do before creating the serving endpoint?

Easy
267

An ML engineer runs a hyperparameter tuning job on Databricks using Hyperopt with SparkTrials. The objective function trains a model and returns a validation metric. The engineer notices that each trial logs to the same MLflow run, making it impossible to compare trials. What should the engineer do to ensure each trial appears as a separate MLflow run?

Hard
268

Which approach is most efficient for handling high-cardinality categorical features in a machine learning model while maintaining compatibility with standard Databricks model serving?

Hard
269

A data scientist trains a scikit-learn model in a Databricks notebook and calls mlflow.sklearn.log_model with the registered_model_name argument set to 'churn_model'. Later, a colleague needs to know which source notebook and Git commit produced run ID 3f8a1c. Where can this information be retrieved?

Medium
270

A data scientist registers a scikit-learn model in the Unity Catalog model registry with the name `prod.ml_team.fraud_detector`. They now want to serve it with Databricks Model Serving. Which value should be supplied as the model identifier when creating the serving endpoint?

Medium
271

When performing hyperparameter tuning using Hyperopt on Databricks, which function is primarily used to distribute the training task across the cluster?

Easy
272

An ML engineer wants to package a training project so it can be run reproducibly from a Databricks Job across environments, with its Python dependencies and entry point defined. Which MLflow capability should be used?

Medium
273

A data scientist trains a model with a feature engineering pipeline and wants batch scoring to happen nightly on a Delta table using Databricks, producing predictions that downstream dashboards read. The scoring job must scale with data volume and be re-runnable if it fails. Which approach best fits these requirements?

Medium
274

A data scientist registers a scikit-learn model in Unity Catalog as `prod.ml.churn_model`. They then create a Databricks Model Serving endpoint via the REST API using `served_entities` with `entity_name` set to `prod.ml.churn_model` and `entity_version` set to `"3"`. The endpoint creation fails. What is the most likely cause?

Medium
275

Which THREE actions are best practice when deploying a machine learning model using Databricks Model Serving?

Hard
276

A data scientist is using MLflow on Databricks to track experiments. After running several training jobs, they notice that the run metrics are recorded but the model artifact is missing when they view the run details. They logged the model using the default MLflow API. What is the most likely cause?

Medium
277

Which feature in Databricks allows a data scientist to version and manage the lifecycle of machine learning models in a centralized repository?

Easy
278

Why should you use an inference table in Databricks Model Serving?

Medium
279

An ML engineer has registered a model in the Databricks Model Registry. The model must be deployed to a REST endpoint that automatically scales with traffic and provides a stable serving environment. Which Databricks capability should they use?

Hard
280

Which of the following is the recommended workflow for developing a scalable model on Databricks?

Easy
281

A data scientist wants to package a training script so that it can be run repeatedly with different hyperparameters and shared with colleagues who use different Python library versions. The script must create a reproducible environment. Which MLflow component should they use to define the project and its dependencies?

Medium
282

A machine learning engineer needs to track model experiments in Databricks and wants to ensure that model artifacts are versioned automatically. Which approach best leverages Databricks-native capabilities for this requirement?

Medium
283

A data scientist is training a Random Forest model on a 500GB dataset using Databricks. They notice that the model training process is consistently running out of memory on the driver node. Which approach should be taken to resolve this memory bottleneck while maintaining model performance?

Medium
284

What is the primary function of the 'Model Registry' in the Databricks ML ecosystem?

Medium
285

You are using Databricks Feature Store to create a feature table that will be used for both batch training and online inference. The feature table must be refreshed daily with new data, and the online store must serve the latest feature values within minutes of the refresh. Which configuration should you use?

Medium
286

You are deploying a model using MLflow, and you want to log the custom pre-processing logic alongside the model so it is automatically applied during inference. How should you achieve this?

Hard
287

When using MLflow to track experiments, what happens if you invoke mlflow.end_run() inside a nested loop when the parent run is already active?

Hard
288

A data scientist needs to track parameters, metrics, and model artifacts during training on Databricks. Which approach is the industry-standard best practice to ensure reproducibility and lineage?

Medium
289

A machine learning engineer is using MLflow Tracking on Databricks to compare multiple runs. They want to programmatically retrieve the best run based on a metric called 'rmse' from an experiment. Which MLflow API call should they use?

Medium
290

Refer to the exhibit. Why is including the `signature` and `input_example` in the `log_model` call considered a professional best practice?

Hard
291

A data scientist has registered a model in the Databricks Model Registry and wants to deploy it as a REST API endpoint for real-time inference. Which Databricks feature should they use?

Easy
292

A data scientist is training a linear regression model using scikit-learn on Databricks. They want to track the model's hyperparameters, such as fit_intercept and normalize, in MLflow. Which MLflow API call should they use to log these hyperparameters?

Easy
293

You are building a pipeline in Databricks and need to ensure that a training job only runs after the upstream data preparation job has successfully completed. Which Databricks feature should you use?

Medium
294

Refer to the exhibit. A user attempts to transition a model to production, but the code fails with the provided error. What is the most likely cause?

Hard
295

Refer to the exhibit. A machine learning engineer wants to promote this model to Production. Which MLflow action should be performed to achieve this while ensuring the model is ready for deployment?

Medium
296

An ML engineer is configuring a Databricks Job to retrain a model daily. The job must first run a notebook that creates a feature table, then run a notebook that trains and registers the model. The engineer wants to ensure the training notebook only runs if the feature table creation succeeds. Which Databricks Workflows feature should they use to define this dependency?

Hard
297

A machine learning engineer has registered a scikit-learn model in Unity Catalog as `prod.ml.risk_model` with version 3. They want to serve it through a Databricks Model Serving endpoint that always uses this exact model version, even after new versions are registered. Which endpoint configuration should they use?

Medium
298

A data scientist is using Hyperopt with SparkTrials on a Databricks cluster to tune a scikit-learn model. They set max_evals=100 and parallelism=4. After the tuning completes, they notice that some trials failed due to memory errors on the workers. What is the most likely cause of these failures?

Medium
299

A data scientist is working in a Databricks notebook and wants to view the results of their MLflow runs, including metrics and parameters, directly within the notebook. Which MLflow function should they use?

Easy
300

A machine learning engineer is using MLflow to log a model built with XGBoost. They call mlflow.xgboost.log_model(xgb_model, 'model') and then attempt to load the model in a different environment using mlflow.pyfunc.load_model('runs:/<run_id>/model'). The load fails with an error about missing dependencies. Which action should they take to ensure the model can be loaded in the new environment?

Medium
301

Which TWO actions should be taken to ensure reproducibility of a Databricks ML model experiment?

Medium
302

A data scientist has registered a scikit-learn model in Unity Catalog as `ml_team.prod.churn_model` and wants it served by a Databricks Model Serving endpoint that performs online inference for a web app. The endpoint must automatically pick up new model versions as they are promoted, without the data scientist editing the endpoint each time. Which approach should the data scientist use to configure the served entity?

Medium
303

An ML engineer is using Databricks Jobs to orchestrate a pipeline that includes a notebook for feature engineering and a notebook for model training. The training notebook must run only after the feature engineering notebook completes successfully, and both must run on a schedule. Which configuration in Databricks Jobs achieves this dependency?

Hard
304

A team has registered a model in Unity Catalog and wants to serve it with Databricks Model Serving. Their security policy requires that the endpoint access the model without embedding long-lived credentials, and they want the endpoint to use a dedicated service principal for accessing downstream feature tables. Which configuration should they apply to meet these requirements?

Hard
305

Which approach is most efficient for tracking hyperparameter tuning metrics across thousands of runs on Databricks?

Medium
306

A model has been trained on a dataset containing categorical features with high cardinality. Which technique is most effective for preparing these features for a linear model in a Databricks environment?

Medium
307

A team has just registered a new version of a fraud detection model in the workspace Model Registry. Before routing production traffic to it, they want to send a small percentage of live requests to the new version and compare its predictions against the current production model. Which Databricks Model Serving feature should they use?

Easy
308

A team wants to monitor a production Model Serving endpoint for data drift and to capture the exact request payloads and responses for later auditing. They want this logging to happen automatically without adding code to the client application. What should they configure on the endpoint?

Medium
309

A data scientist is performing hyperparameter tuning for a scikit-learn model using MLflow on a Databricks cluster. They want to parallelize the tuning to reduce total runtime. Which Databricks-supported method allows them to run multiple trials concurrently with minimal code changes?

Medium
310

An ML engineer is using Databricks Jobs to orchestrate a multi-step ML pipeline. The pipeline includes a task that trains a model and logs it to MLflow, followed by a task that registers the model in the Model Registry. The engineer wants to ensure that the model is only registered if its accuracy exceeds a threshold. Which approach best implements this conditional logic?

Hard
311

An ML engineer is using MLflow Tracking to compare multiple runs of a hyperparameter tuning experiment. The engineer wants to quickly identify the run that achieved the best validation accuracy and then promote that run's model to the Model Registry. Which MLflow feature allows the engineer to view and compare runs in a centralized UI?

Easy
312

A machine learning engineer is training a scikit-learn model on a Databricks cluster and wants every hyperparameter, evaluation metric, and the fitted estimator itself to be captured automatically without writing custom logging code. The engineer uses MLflow with autologging enabled. Which statement best describes what MLflow autologging records for this scikit-learn run?

Medium
313

A machine learning engineer has a registered model and wants to expose it as a REST endpoint that their application can call for real-time predictions. They need the endpoint to be created and managed natively within Databricks, with the ability to enable scale-to-zero during idle periods. Which Databricks capability should they use?

Easy
314

Refer to the exhibit. Why is providing an 'input_example' highly recommended during the model logging process?

Medium
315

A machine learning engineer is orchestrating a multi-step training workflow with Databricks Jobs. The workflow must retrain a model, evaluate it, and only register it if a quality threshold is met, otherwise stop without registering. Which approach best implements this conditional logic within the job?

Hard
316

A machine learning engineer is using MLflow to log a model trained with XGBoost. They want to ensure that the model can be loaded and used for inference in a different environment without requiring the original training environment. Which MLflow feature allows the model to capture its dependencies and environment?

Hard
317

When preparing data for machine learning in Databricks, which feature of Delta Lake is most beneficial for managing large-scale datasets during the training process?

Easy
318

A machine learning engineer is deploying a model to a Databricks Model Serving endpoint and needs to send feature values that are computed by a separate upstream pipeline. The engineer wants the endpoint to accept a JSON payload describing a single record with named fields. Which approach correctly describes how the client should format the request?

Medium
319

A data scientist is preparing a dataset for training a model on Databricks. They want to split the data into training and testing sets, and they need to ensure that the split is reproducible across different runs. Which PySpark method should they use to split the DataFrame with a fixed random seed?

Easy

Frequently asked questions

What does the scenario questions domain cover on the Databricks-ML-Assoc exam?
scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
How many questions are in this domain?
This page lists all 319 scenario questions questions in the Databricks-ML-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only scenario questions questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.