Be able to deploy a registered MLflow model to a Databricks Model Serving endpoint, declare custom dependencies, and tune workload size and traffic. The key skill is diagnosing latency by separating model inference from preprocessing, package loading, and endpoint scaling.
Start practicing
Model Deployment — choose a session length
Free · No account required
Domain overview
This domain covers taking registered models from the Databricks Model Registry into production: creating Model Serving endpoints, configuring custom dependencies, sizing and scaling compute, controlling versions and traffic, and diagnosing latency or throughput problems. Questions are scenario-based, asking you to pick the correct configuration, command, or diagnostic step rather than recall a definition.
Exam objectives
Creating and updating Model Serving endpoints with the Databricks SDK, REST API, or MLflow
Installing custom Python packages from private PyPI or artifact stores via environment specs
Using MLflow model signatures and requirements to define serving inputs and dependencies
Configuring traffic splitting, scale-to-zero, and workload sizing for latency SLAs
Assuming the default Databricks Runtime already contains a private or niche package instead of declaring it in the endpoint environment spec.
Blaming the model for latency when the bottleneck is endpoint sizing, cold starts, or an oversized preprocessing pipeline.
Editing endpoint config without understanding how traffic percentages and served model versions interact during a rollout.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A machine learning engineer needs to deploy a custom scikit-learn model to a Databricks Model Serving endpoint with a strict response time SLA of under 50 milliseconds. The model includes an extensive text-cleaning pipeline that utilizes heavy regex matching. How should the engineer package the model to ensure maximum inference efficiency and meet the low-latency requirement?
2A machine learning engineer needs to deploy a custom PyTorch model to a Databricks Model Serving endpoint. The model requires a custom pre-processing step that is not part of the standard MLflow transformers or pyfunc flavor. Which deployment approach ensures the custom logic executes reliably within the serverless serving container?
3A data science team is preparing to deploy a high-throughput recommendation model using Databricks Model Serving. Which TWO factors must be considered to optimize endpoint latency and resource utilization? (Choose two)
4A machine learning engineer needs to deploy a custom PyTorch model to a Databricks Model Serving endpoint. The model requires custom post-processing logic and loading auxiliary tokenizer files alongside the serialized weights. Which approach provides the correct mechanism to package and serve this custom artifact?
5A machine learning engineer needs to deploy a custom scikit-learn model to a Databricks Model Serving endpoint with maximum throughput and minimum latency. The model requires an external preprocessing Python script during inference. Which deployment approach best leverages MLflow and Databricks architecture?
6An ML engineer is transitioning a model from the Workspace Model Registry to the Unity Catalog (UC) Model Registry. Which TWO statements describe benefits or requirements of using Unity Catalog for model management? (Select TWO)
7Refer to the exhibit. A machine learning team has updated their model serving endpoint configuration as shown in the JSON. Which deployment strategy is being implemented, and what is the primary risk associated with this specific configuration?
8A data scientist needs to perform batch inference on a large dataset stored in Delta Lake using a model registered in the Unity Catalog. Which approach is most efficient for leveraging Spark's distributed computing capabilities while using the MLflow model?
9When deploying a model to a Databricks Model Serving endpoint, what is the purpose of the 'Small', 'Medium', and 'Large' workload size settings?
10An ML engineer wants to ensure that only models that have passed a specific validation suite can be assigned the 'Champion' alias in Unity Catalog. What is the recommended way to automate this process?
11A team is deploying a model that requires custom Python libraries not available in the default Databricks Runtime. Which TWO methods can be used to ensure these dependencies are available in the Model Serving environment? (Select TWO)
12Which statement best describes the role of the 'Signature' in an MLflow model when deploying to a Databricks Model Serving endpoint?
13Refer to the exhibit. An administrator notices that the cost for this specific endpoint is higher than expected even when there is no traffic. Based on the exhibit, what is the most likely cause of the high idle cost?
14When using the Unity Catalog Model Registry, what are the primary advantages of using 'Aliases' over 'Versions' when calling a model from a production application? (Select TWO)
15An ML engineer is deploying a model that includes a custom Python class for preprocessing. During deployment to a Model Serving endpoint, the model fails to load with a 'ModuleNotFoundError'. What is the most likely cause of this error despite having the class in the training notebook?
16Refer to the exhibit. The JSON configuration represents an existing Databricks Model Serving endpoint. You need to update this endpoint to support a traffic split between version 5 and version 6 for A/B testing. Which update strategy is correct?
17When deploying a model to Databricks Model Serving, you notice that inference latency is higher than expected. Which diagnostic approach is most effective for identifying the bottleneck?
18An ML engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model's inference function requires access to an external feature store table for real-time feature lookup. Which approach allows the model to retrieve these features during serving while maintaining low latency and avoiding per-request authentication complexity?
19A financial services company uses Databricks Model Serving to deploy a real-time fraud detection model. The endpoint is configured with scale-to-zero enabled. During a period of no traffic, the endpoint scales down to zero. When a sudden burst of requests arrives, the first few requests experience high latency. Which mechanism is responsible for this behavior?
20A data science team is deploying a model to Databricks Model Serving and needs to ensure that the endpoint can handle sudden spikes in traffic without dropping requests. They want to configure auto-scaling appropriately. Which TWO parameters should they adjust to control the scaling behavior? (Choose two.)
21An ML engineer has deployed a model to Databricks Model Serving and wants to monitor the endpoint's performance over time. They need to track the number of requests, latency, and error rates. Which Databricks feature provides these metrics out-of-the-box?
22An ML engineer is deploying a model to Databricks Model Serving and wants to implement A/B testing between two model versions. The engineer needs to route a percentage of traffic to each version and collect performance metrics. Which feature of Databricks Model Serving should the engineer use?
23A team is deploying a model to Databricks Model Serving that requires a specific version of a Python library that conflicts with the version pre-installed in the serving environment. They include the library version in the model's requirements.txt. However, upon deployment, the endpoint fails to start, and logs indicate a dependency conflict. What is the most likely cause of this failure?
24An ML engineer is deploying a model to Databricks Model Serving that requires a custom Python package. The package is not available in the default environment and must be installed from a private PyPI repository. Which method ensures the package is available to the model at serving time?
25A data scientist has registered a model in Unity Catalog and wants to deploy it to a Databricks Model Serving endpoint. What is the simplest way to create the endpoint?
26A machine learning engineer has a model registered in Unity Catalog as prod.ml.iris_model. They need to deploy it to a real-time serving endpoint that automatically scales based on traffic and provides a REST API for predictions. The model's signature is logged. Which deployment method should they use?
27A team has deployed a model to Databricks Model Serving and wants to enable autoscaling to handle variable traffic. They configure the endpoint with scale_to_zero_enabled set to true and a min_provisioned_concurrency of 0. After deployment, they notice that the endpoint takes several seconds to respond to the first request after a period of inactivity. What is the cause of this latency?
28A team is deploying a model to Databricks Model Serving and wants to implement a canary release strategy to gradually shift traffic from the current model version to a new version. Which TWO configurations are required to achieve this? (Choose two.)
29An ML engineer needs to deploy a model to Databricks Model Serving that requires a custom Python package. The package is not available in the default Databricks Runtime and must be installed from a private PyPI repository. Which approach should be used to include this package in the model's environment?
30An ML engineer is deploying a model to a Databricks Model Serving endpoint. The model's inference function logs predictions to a Delta table for monitoring. During testing, they notice that the logging adds significant latency. They need to reduce the impact on inference latency. Which approach should they take?
31An ML engineer is deploying a model to Databricks Model Serving that uses a custom transformer requiring a GPU. The endpoint must handle high throughput with low latency. Which workload type and configuration should be selected?
32A machine learning engineer has deployed a model to a Databricks Model Serving endpoint. The model requires a custom Python package that is not available in the default environment. The engineer has already logged the model with MLflow and included the package in the conda environment. However, upon deployment, the endpoint fails to start. What is the most likely cause?
33A financial institution deploys a credit scoring model using Databricks Model Serving. The model must log all incoming requests and outgoing responses to a Delta table for auditing. The ML engineer needs to enable this logging with minimal performance impact. Which solution should they implement?
34An ML engineer needs to deploy a model to Databricks Model Serving. The model was logged with MLflow and registered in Unity Catalog. The engineer wants to ensure that only the latest version of the model is served and that the endpoint can be updated without downtime. Which approach should they use?
35An ML engineer is updating a model serving endpoint to use a new model version. They want to gradually shift traffic from the old version to the new version to monitor performance before full rollout. Which feature of Databricks Model Serving should they use?
36A data scientist is deploying a model to Databricks Model Serving. The model was trained using a scikit-learn pipeline that includes a custom transformer. The custom transformer is defined in a Python module that is not part of the model artifact. What should the data scientist do to ensure the model can be served successfully?
37A team has deployed a model to Databricks Model Serving and enabled inference tables. They notice that the inference table contains request and response payloads but no ground truth labels. They want to automatically join ground truth labels for monitoring. What should they do?
38An ML engineer has registered a scikit-learn model in the Unity Catalog Model Registry and wants to serve it as a real-time endpoint using Databricks Model Serving. The model's MLflow signature expects a JSON payload with an array of records. The engineer creates a serving endpoint with a workload size of Medium and configures the served entity to use the latest model version. Which additional configuration is required to enable automatic payload logging to a Delta table for monitoring?
39A data scientist has deployed a model to a Databricks Model Serving endpoint. The endpoint is configured with scale-to-zero enabled and a workload size of Small. After a period of inactivity, the endpoint scales down to zero. A client application sends a request to the endpoint after this idle period. What happens to the first request?
40A data scientist has trained a model and registered it in Unity Catalog. They now need to deploy it for real-time inference with automatic scaling and a REST API endpoint. Which Databricks feature should they use?
41A company has deployed a model to a Databricks Model Serving endpoint. The model's predictions must be logged to a Delta table for monitoring and auditing. The ML engineer wants to enable inference logging without modifying the model's code. Which approach achieves this with minimal effort?
42An ML engineer is deploying a model to Databricks Model Serving and needs to enable automatic scaling based on traffic. The model has variable inference latency and the team wants to optimize cost while maintaining performance. Which TWO configurations are required to achieve this? (Choose two.)
43An ML engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model expects a single feature vector of 10 float values per request. The endpoint must return predictions in under 100 ms. Which approach should the engineer use to minimize per-request overhead?
44An ML engineer is deploying a model to Databricks Model Serving that uses a custom Python function as a pre-processing step. The function relies on a global variable defined in a separate module. After deployment, the endpoint returns errors indicating the global variable is not defined. The engineer confirmed the module is included in the model's conda environment. What is the most likely cause?
45A data scientist has deployed a model to Databricks Model Serving and wants to monitor its performance over time. They need to track prediction drift and data quality issues. Which Databricks feature should they use to automatically capture inference logs and compute metrics?
46A fraud detection model is deployed to a Databricks Model Serving endpoint. The team wants to test a new model version without affecting existing predictions. They need to send a copy of live traffic to the new version and log its predictions for comparison, while the current version continues to serve all responses. Which feature should they use?
47A data science team is using Databricks Model Serving to deploy a model that must process sensitive data. They need to ensure that all inference requests are logged for auditing purposes, including the input data and predictions. Which approach should they take?
48A team has deployed a model to a Databricks Model Serving endpoint. They want to monitor the endpoint's performance and detect data drift over time. Which Databricks feature should they use to automatically track inference data and compute drift metrics?
49An ML engineer is updating a production model serving endpoint to use a new model version. The endpoint currently serves version 1 with the 'Champion' alias. The engineer wants to test version 2 with a small percentage of live traffic before full rollout. Which deployment strategy should they use in Databricks Model Serving?
50An ML engineer is deploying a model to Databricks Model Serving and needs to ensure that the endpoint can handle traffic spikes while minimizing costs during idle periods. The engineer considers enabling scale-to-zero and configuring autoscaling. Which TWO statements about these features are correct? (Choose two.)
51An ML engineer is deploying a model to Databricks Model Serving that requires a custom Python package not available in the default environment. The model was logged with MLflow and includes the package in its conda environment. What must the engineer ensure for the endpoint to successfully load the model?
52An ML engineer needs to deploy a model to Databricks Model Serving that requires a specific version of a Python library. The library is available on PyPI. Where should the engineer specify this dependency?
53A data science team has deployed a model to Databricks Model Serving and wants to ensure that the endpoint can handle sudden spikes in traffic without manual intervention. Which feature should they configure?
54An ML engineer is deploying a model to Databricks Model Serving and needs to ensure the endpoint can handle sudden spikes in traffic without downtime. The model has a large memory footprint and takes several seconds to load. Which TWO configurations should the engineer implement to achieve this? (Choose two.)
55A data scientist has registered a model in Unity Catalog and wants to deploy it to a Databricks Model Serving endpoint. The model version is 3 and the model name is 'fraud_model'. Which identifier should be used to reference this model version when creating the endpoint?
56A machine learning engineer is deploying a model to Databricks Model Serving and wants to implement a blue-green deployment strategy. They have registered two model versions in Unity Catalog: version 1 (current production) and version 2 (new candidate). They want to route 10% of traffic to version 2 for testing while keeping 90% on version 1. Which feature should they use to achieve this?
57An ML engineer has deployed a model to Databricks Model Serving and wants to update the endpoint to serve a new model version without changing the endpoint URL or causing downtime. Which approach is correct?
58A fraud detection team has a model registered in Unity Catalog as main.ml.fraud_model. They need to serve it in real time, but compliance requires that every scoring request automatically generate an audit record in a Delta table, and that the model only be promoted to production after a human reviews the audit logs from a canary period. Which deployment configuration satisfies these requirements?
59A team is deploying a scikit-learn model to Databricks Model Serving and wants to minimize cold-start latency so that the first request after a period of inactivity is still fast. Which TWO actions help achieve this? (Choose two.)
60A data scientist has a model registered in Unity Catalog and wants to let an external application score it over HTTPS without embedding Databricks credentials in the application. The application's identity is already a service principal in the workspace. Which approach should be used to authenticate calls to the Model Serving endpoint?
61An ML engineer has a Databricks Model Serving endpoint that is currently serving a registered model version. A new model version is registered in Unity Catalog and must be rolled out to the endpoint without any downtime. Which approach should the engineer use to safely transition traffic to the new model version?
Be able to deploy a registered MLflow model to a Databricks Model Serving endpoint, declare custom dependencies, and tune workload size and traffic. The key skill is diagnosing latency by separating model inference from preprocessing, package loading, and endpoint scaling.
The Courseiva Databricks-ML-Pro question bank contains 61 questions in the Model Deployment domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Model Deployment domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included