Courseiva

Databricks-ML-Assoc · domain

Model Deployment

This domain covers deploying MLflow models on Databricks: Model Registry stages and versions, Model Serving endpoints, traffic splitting, and batch inference via jobs. Questions test whether you can configure endpoints, promote validated model versions safely, validate request schemas, and diagnose serving errors from exhibits.

72 questions10 easy33 medium29 hard

Focused practice

Practice Model Deployment questions

Scored sessions drawing only from this domain — pick a length below.

Start 20-question practice test →

What this domain covers

What to know about Model Deployment

Be able to register a model version, create or update a Model Serving endpoint, and split traffic between versions for safe rollout. The most important thing: configure served entities and traffic percentages explicitly, and log a model signature so the endpoint validates request schemas.

Creating and updating Databricks Model Serving endpoints for registered MLflow model versions

Using traffic splitting to route live requests between model versions during rollout

Recording input and output schema signatures so serving endpoints validate incoming requests

Choosing Model Serving endpoints versus batch inference jobs for latency and throughput needs

Watch out for

Common Model Deployment exam traps

  • ▸Assuming updating an endpoint automatically shifts all traffic; you must explicitly configure served entities and traffic percentages.
  • ▸Forgetting that schema enforcement depends on a logged model signature, not just registry metadata or comments.
  • ▸Confusing batch inference jobs with real-time serving, then misreading error exhibits caused by request format or schema mismatch.

Question index

All Model Deployment questions (72)

Click any question to see the full explanation, or start a practice session above.

1

A team registered a model in Unity Catalog and wants to serve it on a Databricks Model Serving endpoint. During deployment, the build fails because the model's conda environment references a private Python package hosted on an internal PyPI mirror that the serving build cannot reach. Which approach resolves the deployment?

Hard
2

An ML engineer is building an automated CI/CD pipeline that, after validating a new model version, must programmatically move it to the `champion` alias in Unity Catalog so the production Model Serving endpoint begins using it. The pipeline runs in a Databricks job using a service principal. Which action accomplishes the promotion programmatically?

Hard
3

A team deploys an MLflow pyfunc model to a Databricks Model Serving endpoint. During pre-deployment testing they call the endpoint with a small batch of records and receive an HTTP 400 error stating the request payload does not match the model signature. The model was logged with an inferred signature from a pandas DataFrame. Which action most directly resolves the mismatch?

Hard
4

Which Databricks feature allows you to manage the lifecycle of a model, including transitions from 'Staging' to 'Production'?

Easy
5

A data scientist has trained a model and wants other team members to serve it through a Databricks Model Serving endpoint. The model must be discoverable by a three-level namespace and governed by Unity Catalog. What should the data scientist do first?

Easy
6

A machine learning engineer maintains a Databricks Model Serving endpoint named `recommendations-endpoint` serving a registered model `ml_team.recommender_model` at version 7. The team wants to route 10% of live traffic to a newly registered candidate version 8 while keeping version 7 serving the remaining 90%, without altering the existing client request URL. Which approach should the engineer take?

Hard
7

What is the purpose of the 'Champion' model version in the Model Registry?

Medium
8

A team is preparing to deploy a model to a Databricks Model Serving endpoint that must scale down to zero replicas when idle yet still serve bursty traffic with acceptable cold-start latency. They also need to capture the request and response payloads for later monitoring. Which two endpoint settings or features should they configure? (Choose two.)

Hard
9

An MLOps engineer is responsible for a Databricks Model Serving endpoint that serves a mission-critical pricing model. The team wants an automated safeguard that detects when live input feature distributions drift away from the training distribution and triggers a retraining workflow, without modifying the model artifact itself. Which Databricks capability should be configured?

Hard
10

When deploying a model to Databricks Model Serving, what is the recommended way to handle sensitive credentials like database connection strings?

Easy
11

An ML engineer deployed a model to a Databricks Model Serving endpoint and enabled inference tables. After a week, the team notices that some requests returned HTTP 200 but the corresponding rows in the inference table show null prediction values. They need to diagnose why predictions are missing for those requests. Which explanation is most consistent with this symptom?

Hard
12

A team is preparing to deploy a registered MLflow model to a Databricks Model Serving endpoint. They want to capture every request and response for later monitoring and debugging, and they also want the endpoint to remain available during a rolling model version update. (Choose two.)

Medium
13

A machine learning engineer has a model registered in Unity Catalog and wants to expose it as a REST API so an external application can send JSON payloads and receive predictions. The team has no existing serving infrastructure. Which Databricks feature should be used to create this API?

Easy
14

A machine learning engineer has registered a model in the Databricks Model Registry and wants to expose it as a REST API with automatic scaling and no server management. The model's Python dependencies are captured in a conda environment file logged with the run. Which Databricks capability should the engineer use to serve this model with minimal operational overhead?

Medium
15

A data science team is deploying a model to a Databricks Model Serving endpoint. They want to enable inference logging to capture the input data and predictions for monitoring and debugging. Which TWO configurations are required to enable inference logging for a serving endpoint? (Choose two.)

Hard
16

Refer to the exhibit. What is the most likely cause of this error in a deployed MLflow model?

Hard
17

A machine learning engineer needs to capture every request and response payload sent to a Databricks Model Serving endpoint so that the team can later join predictions with ground-truth labels for monitoring. Which Databricks feature should they enable on the endpoint?

Medium
18

A team deploys a model to a Databricks Model Serving endpoint and enables inference tables. After a week, they notice that the inference table contains request and response payloads but the payload columns are empty for many rows, while status codes are 200. What is the most likely explanation?

Hard
19

Which strategy is most effective for managing model drift in a production Databricks environment?

Hard
20

A team maintains a Databricks Model Serving endpoint for a fraud model. Compliance requires that every request and response be logged to a Delta table for auditing and later analysis. The endpoint is already configured and serving traffic. What should the team do to capture this data with the least additional infrastructure?

Hard
21

A team has an existing Databricks Model Serving endpoint serving `prod.ml.fraud_model` version 3. They register version 4, which uses a new feature set, and want to shift only 10% of traffic to version 4 while keeping version 3 for the rest. Their endpoint currently has a single served entity for version 3. What is the most appropriate approach?

Hard
22

When deploying a model to a production environment, why is it critical to create a dedicated 'staging' environment before the 'production' environment?

Medium
23

A data scientist deploys a model to a Databricks Model Serving endpoint and enables inference tables. After a week, they want to analyze prediction drift by joining the logged requests with ground-truth labels that arrive later. Which statement describes how they should access the inference table data for this analysis?

Hard
24

A machine learning engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model was logged with MLflow using the default signature and input example. After deployment, the engineer notices that the endpoint's REST API expects a JSON payload in a specific format. Which MLflow artifact is used by Model Serving to determine the expected input format for the endpoint?

Medium
25

A data scientist registers an MLflow model whose `conda.yaml` lists several Python packages. When they create a Databricks Model Serving endpoint from this model, the deployment fails during environment build. Which action is most likely to resolve the failure while preserving the model's dependency requirements?

Hard
26

A team observes that their Databricks Model Serving endpoint occasionally returns HTTP 429 responses during bursty traffic, even though the endpoint shows low average CPU utilization. They want to reduce these throttling errors without over-provisioning capacity. Which action is most appropriate?

Hard
27

A data scientist has a custom Python model wrapped in an MLflow pyfunc flavor and needs to serve it on Databricks Model Serving. The model's preprocessing requires a library that is not part of the default serving environment. What is the correct way to make that dependency available to the endpoint?

Easy
28

A data scientist has registered a scikit-learn model in Unity Catalog as `ml_prod.churn.model_v3` and wants the Databricks Model Serving endpoint to automatically pick up newly registered model versions as they are promoted to the `champion` alias. Which configuration should the data scientist use when creating the serving endpoint?

Medium
29

An ML engineer needs to give an external application a stable HTTPS URL to call a registered model served by Databricks Model Serving. The application must authenticate with a token and must not be able to modify the endpoint configuration. Which approach best meets these requirements?

Medium
30

A machine learning engineer is preparing to deploy a registered model to a Databricks Model Serving endpoint. Before creating the endpoint, the engineer wants to confirm the deployment prerequisites are satisfied. Which two conditions are required for a successful endpoint creation? (Choose two.)

Medium
31

A team wants to route production traffic to a new model version while keeping risk low. They configure a Databricks Model Serving endpoint with two served entities: `champion` (entity_version 5) and `challenger` (entity_version 6). They want 95% of requests to hit `champion` and 5% to hit `challenger`. Which configuration accomplishes this?

Hard
32

A team has a Databricks Model Serving endpoint configured with scale-to-zero enabled and min_instances set to 0. During a load test, they observe that the first request after an idle period takes roughly 40 seconds while subsequent requests complete in under 200 milliseconds. They need to eliminate this cold-start latency for a customer-facing application without over-provisioning. Which configuration change best addresses the requirement?

Hard
33

What is the primary role of an 'MLflow Signature' during the model deployment phase?

Medium
34

A data scientist wants to test a newly registered model version interactively before promoting it to production. They need to send a sample request to the model and inspect the prediction and the model's input schema. Which Databricks feature should they use?

Easy
35

A data scientist has registered a scikit-learn model in Unity Catalog and now wants to serve it behind a Databricks Model Serving endpoint. The model's MLflow signature records a pandas DataFrame input with three named columns. The team wants the endpoint to reject malformed requests automatically rather than silently scoring them. Which action should the data scientist take?

Medium
36

A team queries a Databricks Model Serving endpoint through the serving client and receives an error indicating the endpoint is not ready. They confirmed the endpoint exists. Which condition most directly explains why requests fail until it clears?

Medium
37

When deploying a model to a production endpoint, what is the best practice for handling dependencies?

Medium
38

A team is deploying a batch scoring pipeline that loads a registered MLflow model and runs predictions over a large Delta table using Spark. They want the scoring job to reuse the model's training-time preprocessing and to remain reproducible months later. Which two practices should they follow? (Choose two.)

Hard
39

Which THREE factors should be considered when choosing the 'workload size' (e.g., Small, Medium, Large) for a Databricks Model Serving endpoint?

Hard
40

A data scientist needs to deploy a model to Databricks Model Serving. Which component is strictly required to be logged in MLflow to enable the 'Model Serving' feature?

Medium
41

A fraud detection team wants their Databricks Model Serving endpoint to log every request and response payload to a Unity Catalog Delta table so analysts can later join predictions with ground-truth labels. Which endpoint capability should they enable?

Medium
42

A platform team is rolling out a new Databricks Model Serving endpoint for a churn model. They must ensure the endpoint can be queried by an external application and that only authorized callers can invoke it. Which TWO actions should they take? (Choose two.)

Medium
43

When using Databricks Model Serving, what is the primary benefit of using a 'Provisioned Throughput' endpoint over a 'Serverless' endpoint?

Medium
44

A machine learning engineer is configuring a Databricks Model Serving endpoint for a model that requires GPU acceleration. They set the workload size to 'GPU_Medium' but the endpoint fails to deploy. Which of the following is the most likely cause of the failure?

Medium
45

Refer to the exhibit. What is the impact of setting 'auto_capture_request_payload' to true?

Medium
46

A data scientist has registered a scikit-learn model in the Databricks Model Registry as `prod.churn_model`. The production endpoint serving this model must automatically roll back to the previously served version if the newly deployed version's error rate exceeds a threshold within one hour of deployment. Which Databricks feature should the data scientist configure to meet this requirement?

Medium
47

A team is standing up a real-time Databricks Model Serving endpoint for a fraud model. Requests will carry several numeric features, and the team wants the endpoint to reject malformed payloads with a clear client error rather than silently scoring them, and to avoid cold-start latency during business hours. Which two actions should the team take? (Choose two.)

Hard
48

A team serves a model with Databricks Model Serving and wants to send production traffic to a newly registered model version while keeping the ability to revert instantly if quality degrades. They prefer not to edit the endpoint configuration to switch versions. Which approach best fits this requirement?

Hard
49

When deploying a model using Model Serving, how does Databricks ensure that the environment remains consistent between the training workspace and the serving environment?

Medium
50

A machine learning engineer has an existing Databricks Model Serving endpoint named churn-endpoint serving version 3 of a model. The team has validated version 5 and wants to direct live traffic to it while keeping the deployment reversible if quality degrades. What is the most appropriate action?

Hard
51

A machine learning engineer is preparing to deploy a scikit-learn model as a Databricks Model Serving endpoint. The model expects a pandas DataFrame with specific column names and types. Which two actions should the engineer take to ensure the endpoint correctly validates and processes inference requests? (Choose two.)

Medium
52

A team operates a Databricks Model Serving endpoint with min_instances set to 0 and max_instances set to 4. During a nightly batch job, the endpoint receives a burst of requests and some clients observe elevated latency. The team wants to keep costs low during idle periods while reducing cold-start latency during bursts. Which configuration change best achieves this?

Hard
53

A team's Model Serving endpoint occasionally returns errors when the upstream feature store is slow. They want the endpoint to retry transient failures and reduce cold-start latency for bursty traffic. Which combination of endpoint settings best addresses both concerns?

Hard
54

A machine learning engineer is deploying an MLflow model to a Databricks Model Serving endpoint. The model was trained on a Spark DataFrame and logged with MLflow using the default signature detection. During testing, the endpoint returns predictions, but the engineer notices the input schema shown in the Serving UI does not match the actual DataFrame column types used during training. Which MLflow logging step should the engineer verify first to resolve this schema mismatch?

Medium
55

A team owns a Databricks Model Serving endpoint that receives sporadic bursts of traffic. They want to reduce cold-start latency during bursts while keeping cost predictable, and they also need to capture the request payloads and predictions for later monitoring. Which two configuration choices should they make? (Choose two.)

Hard
56

Refer to the exhibit. A user attempts to update a model stage in the Model Registry and receives this error. What is the most appropriate action to resolve this?

Hard
57

A team has several model versions registered in Unity Catalog. They want to serve a specific version through a Databricks Model Serving endpoint and later promote a newer version without changing the endpoint URL used by applications. Which approach should they use?

Hard
58

A team is preparing to deploy an MLflow model to a Databricks Model Serving endpoint and wants to diagnose why requests are failing before contacting support. Which two actions allow them to inspect the endpoint's behavior and errors? (Choose two.)

Medium
59

Which of the following is a primary benefit of using a model serving endpoint versus a batch inference job?

Medium
60

A data scientist registers a model in the Databricks Model Registry and wants to record the model's intended input and output schema so that a serving endpoint can validate incoming requests. Which action accomplishes this when logging the model?

Easy
61

A data scientist has trained a scikit-learn model and wants to expose it for real-time inference through Databricks Model Serving. The model is currently logged as an MLflow run artifact but has not been registered anywhere. What must the data scientist do before creating the serving endpoint?

Easy
62

A data scientist registers a scikit-learn model in the Unity Catalog model registry with the name `prod.ml_team.fraud_detector`. They now want to serve it with Databricks Model Serving. Which value should be supplied as the model identifier when creating the serving endpoint?

Medium
63

A data scientist trains a model with a feature engineering pipeline and wants batch scoring to happen nightly on a Delta table using Databricks, producing predictions that downstream dashboards read. The scoring job must scale with data volume and be re-runnable if it fails. Which approach best fits these requirements?

Medium
64

A data scientist registers a scikit-learn model in Unity Catalog as `prod.ml.churn_model`. They then create a Databricks Model Serving endpoint via the REST API using `served_entities` with `entity_name` set to `prod.ml.churn_model` and `entity_version` set to `"3"`. The endpoint creation fails. What is the most likely cause?

Medium
65

Why should you use an inference table in Databricks Model Serving?

Medium
66

A machine learning engineer has registered a scikit-learn model in Unity Catalog as `prod.ml.risk_model` with version 3. They want to serve it through a Databricks Model Serving endpoint that always uses this exact model version, even after new versions are registered. Which endpoint configuration should they use?

Medium
67

A data scientist has registered a scikit-learn model in Unity Catalog as `ml_team.prod.churn_model` and wants it served by a Databricks Model Serving endpoint that performs online inference for a web app. The endpoint must automatically pick up new model versions as they are promoted, without the data scientist editing the endpoint each time. Which approach should the data scientist use to configure the served entity?

Medium
68

A team has registered a model in Unity Catalog and wants to serve it with Databricks Model Serving. Their security policy requires that the endpoint access the model without embedding long-lived credentials, and they want the endpoint to use a dedicated service principal for accessing downstream feature tables. Which configuration should they apply to meet these requirements?

Hard
69

A team has just registered a new version of a fraud detection model in the workspace Model Registry. Before routing production traffic to it, they want to send a small percentage of live requests to the new version and compare its predictions against the current production model. Which Databricks Model Serving feature should they use?

Easy
70

A team wants to monitor a production Model Serving endpoint for data drift and to capture the exact request payloads and responses for later auditing. They want this logging to happen automatically without adding code to the client application. What should they configure on the endpoint?

Medium
71

A machine learning engineer has a registered model and wants to expose it as a REST endpoint that their application can call for real-time predictions. They need the endpoint to be created and managed natively within Databricks, with the ability to enable scale-to-zero during idle periods. Which Databricks capability should they use?

Easy
72

A machine learning engineer is deploying a model to a Databricks Model Serving endpoint and needs to send feature values that are computed by a separate upstream pipeline. The engineer wants the endpoint to accept a JSON payload describing a single record with named fields. Which approach correctly describes how the client should format the request?

Medium

Frequently asked questions

What does the Model Deployment domain cover on the Databricks-ML-Assoc exam?
Be able to register a model version, create or update a Model Serving endpoint, and split traffic between versions for safe rollout. The most important thing: configure served entities and traffic percentages explicitly, and log a model signature so the endpoint validates request schemas.
How many questions are in this domain?
This page lists all 72 Model Deployment questions in the Databricks-ML-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
What is the best way to practise this domain?
Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
Can I practise only Model Deployment questions?
Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.
databricks-ml-associate DATABRICKS-ML-ASSOCIATE ml assoc model deployment Practice Questions