Courseiva

Databricks-ML-Pro · domain

Model Deployment

This domain covers taking registered models from the Databricks Model Registry into production: creating Model Serving endpoints, configuring custom dependencies, sizing and scaling compute, controlling versions and traffic, and diagnosing latency or throughput problems. Questions are scenario-based, asking you to pick the correct configuration, command, or diagnostic step rather than recall a definition.

61 questions11 easy26 medium24 hard

Focused practice

Practice Model Deployment questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Model Deployment

Be able to deploy a registered MLflow model to a Databricks Model Serving endpoint, declare custom dependencies, and tune workload size and traffic. The key skill is diagnosing latency by separating model inference from preprocessing, package loading, and endpoint scaling.

Creating and updating Model Serving endpoints with the Databricks SDK, REST API, or MLflow

Installing custom Python packages from private PyPI or artifact stores via environment specs

Using MLflow model signatures and requirements to define serving inputs and dependencies

Configuring traffic splitting, scale-to-zero, and workload sizing for latency SLAs

Watch out for

Common Model Deployment exam traps

  • ▸Assuming the default Databricks Runtime already contains a private or niche package instead of declaring it in the endpoint environment spec.
  • ▸Blaming the model for latency when the bottleneck is endpoint sizing, cold starts, or an oversized preprocessing pipeline.
  • ▸Editing endpoint config without understanding how traffic percentages and served model versions interact during a rollout.

Question index

All Model Deployment questions (61)

Click any question to see the full explanation, or start a practice session above.

1

Refer to the exhibit. The JSON configuration represents an existing Databricks Model Serving endpoint. You need to update this endpoint to support a traffic split between version 5 and version 6 for A/B testing. Which update strategy is correct?

Hard
2

A team is deploying a model to Databricks Model Serving and wants to implement a canary release strategy to gradually shift traffic from the current model version to a new version. Which TWO configurations are required to achieve this? (Choose two.)

Medium
3

An ML engineer is deploying a model to Databricks Model Serving that requires a custom Python package not available in the default environment. The model was logged with MLflow and includes the package in its conda environment. What must the engineer ensure for the endpoint to successfully load the model?

Medium
4

A team has deployed a model to a Databricks Model Serving endpoint. They want to monitor the endpoint's performance and detect data drift over time. Which Databricks feature should they use to automatically track inference data and compute drift metrics?

Easy
5

A data science team has deployed a model to Databricks Model Serving and wants to ensure that the endpoint can handle sudden spikes in traffic without manual intervention. Which feature should they configure?

Easy
6

A machine learning engineer needs to deploy a custom scikit-learn model to a Databricks Model Serving endpoint with a strict response time SLA of under 50 milliseconds. The model includes an extensive text-cleaning pipeline that utilizes heavy regex matching. How should the engineer package the model to ensure maximum inference efficiency and meet the low-latency requirement?

Medium
7

An ML engineer is deploying a model that includes a custom Python class for preprocessing. During deployment to a Model Serving endpoint, the model fails to load with a 'ModuleNotFoundError'. What is the most likely cause of this error despite having the class in the training notebook?

Hard
8

A team is deploying a scikit-learn model to Databricks Model Serving and wants to minimize cold-start latency so that the first request after a period of inactivity is still fast. Which TWO actions help achieve this? (Choose two.)

Medium
9

A team is deploying a model to Databricks Model Serving that requires a specific version of a Python library that conflicts with the version pre-installed in the serving environment. They include the library version in the model's requirements.txt. However, upon deployment, the endpoint fails to start, and logs indicate a dependency conflict. What is the most likely cause of this failure?

Hard
10

When using the Unity Catalog Model Registry, what are the primary advantages of using 'Aliases' over 'Versions' when calling a model from a production application? (Select TWO)

Medium
11

An ML engineer has deployed a model to Databricks Model Serving and wants to update the endpoint to serve a new model version without changing the endpoint URL or causing downtime. Which approach is correct?

Medium
12

When deploying a model to a Databricks Model Serving endpoint, what is the purpose of the 'Small', 'Medium', and 'Large' workload size settings?

Easy
13

An ML engineer is deploying a model to Databricks Model Serving that requires a custom Python package. The package is not available in the default environment and must be installed from a private PyPI repository. Which method ensures the package is available to the model at serving time?

Hard
14

A data scientist is deploying a model to Databricks Model Serving. The model was trained using a scikit-learn pipeline that includes a custom transformer. The custom transformer is defined in a Python module that is not part of the model artifact. What should the data scientist do to ensure the model can be served successfully?

Medium
15

An ML engineer is updating a model serving endpoint to use a new model version. They want to gradually shift traffic from the old version to the new version to monitor performance before full rollout. Which feature of Databricks Model Serving should they use?

Medium
16

A team is deploying a model that requires custom Python libraries not available in the default Databricks Runtime. Which TWO methods can be used to ensure these dependencies are available in the Model Serving environment? (Select TWO)

Hard
17

An ML engineer is transitioning a model from the Workspace Model Registry to the Unity Catalog (UC) Model Registry. Which TWO statements describe benefits or requirements of using Unity Catalog for model management? (Select TWO)

Medium
18

A data science team is preparing to deploy a high-throughput recommendation model using Databricks Model Serving. Which TWO factors must be considered to optimize endpoint latency and resource utilization? (Choose two)

Hard
19

A data scientist has deployed a model to Databricks Model Serving and wants to monitor its performance over time. They need to track prediction drift and data quality issues. Which Databricks feature should they use to automatically capture inference logs and compute metrics?

Medium
20

An ML engineer is deploying a model to Databricks Model Serving and needs to ensure the endpoint can handle sudden spikes in traffic without downtime. The model has a large memory footprint and takes several seconds to load. Which TWO configurations should the engineer implement to achieve this? (Choose two.)

Hard
21

When deploying a model to Databricks Model Serving, you notice that inference latency is higher than expected. Which diagnostic approach is most effective for identifying the bottleneck?

Medium
22

An ML engineer is deploying a model to a Databricks Model Serving endpoint. The model's inference function logs predictions to a Delta table for monitoring. During testing, they notice that the logging adds significant latency. They need to reduce the impact on inference latency. Which approach should they take?

Hard
23

A data scientist has trained a model and registered it in Unity Catalog. They now need to deploy it for real-time inference with automatic scaling and a REST API endpoint. Which Databricks feature should they use?

Easy
24

An ML engineer is deploying a model to Databricks Model Serving and needs to ensure that the endpoint can handle traffic spikes while minimizing costs during idle periods. The engineer considers enabling scale-to-zero and configuring autoscaling. Which TWO statements about these features are correct? (Choose two.)

Medium
25

An ML engineer needs to deploy a model to Databricks Model Serving that requires a specific version of a Python library. The library is available on PyPI. Where should the engineer specify this dependency?

Easy
26

A company has deployed a model to a Databricks Model Serving endpoint. The model's predictions must be logged to a Delta table for monitoring and auditing. The ML engineer wants to enable inference logging without modifying the model's code. Which approach achieves this with minimal effort?

Hard
27

A data scientist has deployed a model to a Databricks Model Serving endpoint. The endpoint is configured with scale-to-zero enabled and a workload size of Small. After a period of inactivity, the endpoint scales down to zero. A client application sends a request to the endpoint after this idle period. What happens to the first request?

Easy
28

An ML engineer has deployed a model to Databricks Model Serving and wants to monitor the endpoint's performance over time. They need to track the number of requests, latency, and error rates. Which Databricks feature provides these metrics out-of-the-box?

Easy
29

An ML engineer is updating a production model serving endpoint to use a new model version. The endpoint currently serves version 1 with the 'Champion' alias. The engineer wants to test version 2 with a small percentage of live traffic before full rollout. Which deployment strategy should they use in Databricks Model Serving?

Hard
30

A fraud detection model is deployed to a Databricks Model Serving endpoint. The team wants to test a new model version without affecting existing predictions. They need to send a copy of live traffic to the new version and log its predictions for comparison, while the current version continues to serve all responses. Which feature should they use?

Hard
31

An ML engineer wants to ensure that only models that have passed a specific validation suite can be assigned the 'Champion' alias in Unity Catalog. What is the recommended way to automate this process?

Medium
32

An ML engineer is deploying a model to Databricks Model Serving that uses a custom Python function as a pre-processing step. The function relies on a global variable defined in a separate module. After deployment, the endpoint returns errors indicating the global variable is not defined. The engineer confirmed the module is included in the model's conda environment. What is the most likely cause?

Hard
33

An ML engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model's inference function requires access to an external feature store table for real-time feature lookup. Which approach allows the model to retrieve these features during serving while maintaining low latency and avoiding per-request authentication complexity?

Medium
34

A financial services company uses Databricks Model Serving to deploy a real-time fraud detection model. The endpoint is configured with scale-to-zero enabled. During a period of no traffic, the endpoint scales down to zero. When a sudden burst of requests arrives, the first few requests experience high latency. Which mechanism is responsible for this behavior?

Hard
35

A data scientist has a model registered in Unity Catalog and wants to let an external application score it over HTTPS without embedding Databricks credentials in the application. The application's identity is already a service principal in the workspace. Which approach should be used to authenticate calls to the Model Serving endpoint?

Easy
36

A data scientist needs to perform batch inference on a large dataset stored in Delta Lake using a model registered in the Unity Catalog. Which approach is most efficient for leveraging Spark's distributed computing capabilities while using the MLflow model?

Medium
37

A fraud detection team has a model registered in Unity Catalog as main.ml.fraud_model. They need to serve it in real time, but compliance requires that every scoring request automatically generate an audit record in a Delta table, and that the model only be promoted to production after a human reviews the audit logs from a canary period. Which deployment configuration satisfies these requirements?

Medium
38

A machine learning engineer needs to deploy a custom scikit-learn model to a Databricks Model Serving endpoint with maximum throughput and minimum latency. The model requires an external preprocessing Python script during inference. Which deployment approach best leverages MLflow and Databricks architecture?

Medium
39

Which statement best describes the role of the 'Signature' in an MLflow model when deploying to a Databricks Model Serving endpoint?

Medium
40

A data scientist has registered a model in Unity Catalog and wants to deploy it to a Databricks Model Serving endpoint. What is the simplest way to create the endpoint?

Easy
41

A machine learning engineer has deployed a model to a Databricks Model Serving endpoint. The model requires a custom Python package that is not available in the default environment. The engineer has already logged the model with MLflow and included the package in the conda environment. However, upon deployment, the endpoint fails to start. What is the most likely cause?

Hard
42

A data scientist has registered a model in Unity Catalog and wants to deploy it to a Databricks Model Serving endpoint. The model version is 3 and the model name is 'fraud_model'. Which identifier should be used to reference this model version when creating the endpoint?

Easy
43

A machine learning engineer needs to deploy a custom PyTorch model to a Databricks Model Serving endpoint. The model requires a custom pre-processing step that is not part of the standard MLflow transformers or pyfunc flavor. Which deployment approach ensures the custom logic executes reliably within the serverless serving container?

Medium
44

An ML engineer needs to deploy a model to Databricks Model Serving that requires a custom Python package. The package is not available in the default Databricks Runtime and must be installed from a private PyPI repository. Which approach should be used to include this package in the model's environment?

Medium
45

A data science team is using Databricks Model Serving to deploy a model that must process sensitive data. They need to ensure that all inference requests are logged for auditing purposes, including the input data and predictions. Which approach should they take?

Medium
46

A data science team is deploying a model to Databricks Model Serving and needs to ensure that the endpoint can handle sudden spikes in traffic without dropping requests. They want to configure auto-scaling appropriately. Which TWO parameters should they adjust to control the scaling behavior? (Choose two.)

Medium
47

Refer to the exhibit. A machine learning team has updated their model serving endpoint configuration as shown in the JSON. Which deployment strategy is being implemented, and what is the primary risk associated with this specific configuration?

Hard
48

An ML engineer is deploying a model to Databricks Model Serving that uses a custom transformer requiring a GPU. The endpoint must handle high throughput with low latency. Which workload type and configuration should be selected?

Hard
49

A machine learning engineer has a model registered in Unity Catalog as prod.ml.iris_model. They need to deploy it to a real-time serving endpoint that automatically scales based on traffic and provides a REST API for predictions. The model's signature is logged. Which deployment method should they use?

Medium
50

A machine learning engineer needs to deploy a custom PyTorch model to a Databricks Model Serving endpoint. The model requires custom post-processing logic and loading auxiliary tokenizer files alongside the serialized weights. Which approach provides the correct mechanism to package and serve this custom artifact?

Medium
51

An ML engineer has registered a scikit-learn model in the Unity Catalog Model Registry and wants to serve it as a real-time endpoint using Databricks Model Serving. The model's MLflow signature expects a JSON payload with an array of records. The engineer creates a serving endpoint with a workload size of Medium and configures the served entity to use the latest model version. Which additional configuration is required to enable automatic payload logging to a Delta table for monitoring?

Medium
52

An ML engineer is deploying a model to Databricks Model Serving and needs to enable automatic scaling based on traffic. The model has variable inference latency and the team wants to optimize cost while maintaining performance. Which TWO configurations are required to achieve this? (Choose two.)

Hard
53

A team has deployed a model to Databricks Model Serving and enabled inference tables. They notice that the inference table contains request and response payloads but no ground truth labels. They want to automatically join ground truth labels for monitoring. What should they do?

Hard
54

Refer to the exhibit. An administrator notices that the cost for this specific endpoint is higher than expected even when there is no traffic. Based on the exhibit, what is the most likely cause of the high idle cost?

Hard
55

An ML engineer has a Databricks Model Serving endpoint that is currently serving a registered model version. A new model version is registered in Unity Catalog and must be rolled out to the endpoint without any downtime. Which approach should the engineer use to safely transition traffic to the new model version?

Hard
56

A machine learning engineer is deploying a model to Databricks Model Serving and wants to implement a blue-green deployment strategy. They have registered two model versions in Unity Catalog: version 1 (current production) and version 2 (new candidate). They want to route 10% of traffic to version 2 for testing while keeping 90% on version 1. Which feature should they use to achieve this?

Hard
57

An ML engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model expects a single feature vector of 10 float values per request. The endpoint must return predictions in under 100 ms. Which approach should the engineer use to minimize per-request overhead?

Medium
58

A financial institution deploys a credit scoring model using Databricks Model Serving. The model must log all incoming requests and outgoing responses to a Delta table for auditing. The ML engineer needs to enable this logging with minimal performance impact. Which solution should they implement?

Hard
59

An ML engineer needs to deploy a model to Databricks Model Serving. The model was logged with MLflow and registered in Unity Catalog. The engineer wants to ensure that only the latest version of the model is served and that the endpoint can be updated without downtime. Which approach should they use?

Easy
60

An ML engineer is deploying a model to Databricks Model Serving and wants to implement A/B testing between two model versions. The engineer needs to route a percentage of traffic to each version and collect performance metrics. Which feature of Databricks Model Serving should the engineer use?

Hard
61

A team has deployed a model to Databricks Model Serving and wants to enable autoscaling to handle variable traffic. They configure the endpoint with scale_to_zero_enabled set to true and a min_provisioned_concurrency of 0. After deployment, they notice that the endpoint takes several seconds to respond to the first request after a period of inactivity. What is the cause of this latency?

Hard

Frequently asked questions

What does the Model Deployment domain cover on the Databricks-ML-Pro exam?
Be able to deploy a registered MLflow model to a Databricks Model Serving endpoint, declare custom dependencies, and tune workload size and traffic. The key skill is diagnosing latency by separating model inference from preprocessing, package loading, and endpoint scaling.
How many questions are in this domain?
This page lists all 61 Model Deployment questions in the Databricks-ML-Pro question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Model Deployment questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
databricks-ml-professional DATABRICKS-ML-PROFESSIONAL ml pro model deployment Practice Questions