Courseiva

CCNA Model Deployment Questions

72 questions · Model Deployment · All types, answers revealed

1
MCQhard

A team registered a model in Unity Catalog and wants to serve it on a Databricks Model Serving endpoint. During deployment, the build fails because the model's conda environment references a private Python package hosted on an internal PyPI mirror that the serving build cannot reach. Which approach resolves the deployment?

A.Add the internal PyPI mirror URL as an environment variable on the serving endpoint configuration.
B.Convert the private package into a wheel, log it as a model artifact, and reference it in the model's conda environment via a pip requirement pointing to the artifact path.
C.Increase the endpoint's workload size so the build has more memory for dependency resolution.
D.Remove the private package from the conda environment and rely on the package being preinstalled in the serving base image.
AnswerB

Logging the private wheel alongside the model and referencing it with a relative pip path in the conda environment file lets the serving build resolve the dependency from the model artifacts themselves, avoiding external network access. This is the documented pattern for private libraries and keeps the environment reproducible and self-contained for deployment.

Why this answer

The robust fix is to make the model environment self-contained by logging the private wheel as part of the model artifacts and pointing the conda pip section at that relative path. The serving build then installs the dependency from artifacts it already has, with no need to reach an internal index. This preserves reproducibility and avoids exposing private infrastructure to the serving build.

Exam trap

The trap here is treating a dependency-resolution failure as a compute sizing or environment-variable problem instead of a package-availability problem.

2
MCQhard

An ML engineer is building an automated CI/CD pipeline that, after validating a new model version, must programmatically move it to the `champion` alias in Unity Catalog so the production Model Serving endpoint begins using it. The pipeline runs in a Databricks job using a service principal. Which action accomplishes the promotion programmatically?

A.Write the model version number into a Delta table that the endpoint reads at request time to decide which version to load.
B.Call the MLflow client method `set_registered_model_alias` with the model name, alias `champion`, and the target version, using a client connected to the Databricks registry URI.
C.Delete the current `champion` alias and recreate it by registering the model again under the same name with a new version number.
D.Invoke the endpoint's REST API with an update payload that includes the new model version number, forcing the endpoint to reload the model.
AnswerB

The MLflow client exposes `set_registered_model_alias`, which updates the alias pointer in the registry. When the client is configured with the Databricks tracking URI and the service principal's credentials, the call moves `champion` to the validated version, and the serving endpoint that resolves that alias picks it up. This is the supported programmatic promotion path.

Why this answer

Programmatic promotion in Unity Catalog is done through the MLflow client's alias APIs: `set_registered_model_alias` moves the alias to the validated version, and any serving endpoint resolving that alias adopts the change. Endpoint inference calls, custom Delta lookup tables, and delete-and-recreate maneuvers do not perform an atomic alias update and either fail outright or introduce availability risk.

Exam trap

The trap here is conflating endpoint inference APIs with registry management APIs, leading to the belief that a prediction call can change which model version is served.

3
MCQhard

A team deploys an MLflow pyfunc model to a Databricks Model Serving endpoint. During pre-deployment testing they call the endpoint with a small batch of records and receive an HTTP 400 error stating the request payload does not match the model signature. The model was logged with an inferred signature from a pandas DataFrame. Which action most directly resolves the mismatch?

A.Re-log the model with an explicit input_example and signature, then register the new version and update the served entity to that version.
B.Enable inference tables on the endpoint so failed requests are logged and the signature is auto-corrected.
C.Increase the endpoint's workload size from Small to Large so the request payload can be accepted.
D.Add a `Content-Type: text/plain` header to the request so the endpoint parses the payload differently.
AnswerA

Re-logging with an explicit signature and input_example ensures the logged schema matches the payload the endpoint expects, and registering a new version plus updating the served entity points the endpoint at the corrected model. This directly addresses the signature mismatch reported in the 400 error.

Why this answer

A signature mismatch means the logged model's expected schema differs from the payload the endpoint receives. The reliable fix is to re-log the model with an explicit signature and input example, register the new version, and repoint the served entity. Compute scaling, content-type changes, and inference tables address different concerns and do not alter schema validation.

Exam trap

The trap here is treating a signature validation error as a capacity or networking problem instead of a schema mismatch that must be fixed by re-logging the model.

4
MCQeasy

Which Databricks feature allows you to manage the lifecycle of a model, including transitions from 'Staging' to 'Production'?

A.Databricks SQL
B.MLflow Model Registry
C.Delta Live Tables
D.Unity Catalog
AnswerB

The MLflow Model Registry is the purpose-built service in Databricks for managing the entire model lifecycle. It allows for versioning, metadata tracking, and stage transitions, providing an organized approach to moving models from initial experimentation through staging to final deployment in a production serving endpoint.

Why this answer

The MLflow Model Registry provides a centralized store for managing models throughout their lifecycle. It allows teams to track model versions, apply transitions (e.g., Staging to Production), and manage metadata. This central control is fundamental for MLOps, as it provides a clear record of which model is currently active in the production environment and ensures that only validated models are promoted to downstream systems.

Exam trap

Candidates often confuse workspace git repositories or MLflow tracking servers with the specific component dedicated to lifecycle stage transitions like 'Staging' to 'Production'.

5
MCQeasy

A data scientist has trained a model and wants other team members to serve it through a Databricks Model Serving endpoint. The model must be discoverable by a three-level namespace and governed by Unity Catalog. What should the data scientist do first?

A.Export the model to a DBFS path and point the serving endpoint at the DBFS file path.
B.Log the model with MLflow and register it to a Unity Catalog model using `mlflow.register_model` with a three-level name.
C.Package the model as a Docker image and push it to a container registry for the endpoint to pull.
D.Create a Databricks job that runs the model's predict method on a schedule and writes results to a Delta table.
AnswerB

Registering the model to Unity Catalog under a three-level namespace like `catalog.schema.model_name` makes it discoverable and governed, and provides the versioned artifact that Model Serving references. This is the prerequisite step before creating an endpoint. Without a registered Unity Catalog model version, the serving configuration has no valid entity to load.

Why this answer

Serving a model through Databricks Model Serving requires a registered model version. Registering to Unity Catalog with a three-level name provides governance and discoverability, and the endpoint then references that catalog, schema, and model name with a version. Exporting to DBFS, scheduling batch jobs, or using custom containers do not establish the Unity Catalog model entry required here.

Exam trap

The trap here is thinking the endpoint can point directly at a serialized model file, when it must reference a registered model version in Unity Catalog or the workspace Model Registry.

6
MCQhard

A machine learning engineer maintains a Databricks Model Serving endpoint named `recommendations-endpoint` serving a registered model `ml_team.recommender_model` at version 7. The team wants to route 10% of live traffic to a newly registered candidate version 8 while keeping version 7 serving the remaining 90%, without altering the existing client request URL. Which approach should the engineer take?

A.Configure traffic splitting on the existing endpoint by setting `traffic_config` with `routes` assigning 10% to version 8 and 90% to version 7.
B.Update the endpoint's `config` to point only at version 8 and rely on client-side retries to fall back to version 7 when errors occur.
C.Overwrite version 7 in the Model Registry with the artifacts of version 8, then restart the endpoint so it serves the updated model.
D.Create a second endpoint for version 8 and use an external load balancer to split traffic between the two endpoint URLs.
AnswerA

Databricks Model Serving supports traffic splitting on an existing endpoint through the `traffic_config` field, where each entry in `routes` specifies a served model name and a percentage. Setting version 8 to 10% and version 7 to 90% preserves the same endpoint URL and load-balances live requests accordingly, enabling safe canary-style evaluation without touching client code.

Why this answer

Databricks Model Serving endpoints expose a `traffic_config` setting whose `routes` array maps served model versions to percentage weights. Adjusting those weights on the existing endpoint shifts a defined share of live requests to the candidate version while the endpoint URL and client integration remain unchanged. This supports progressive rollout and A/B comparison, and the split can be reverted instantly by restoring the previous weights.

Exam trap

The trap here is assuming a new model version requires a brand-new endpoint, when Databricks Model Serving can split traffic between versions on the same endpoint through `traffic_config`.

7
MCQmedium

What is the purpose of the 'Champion' model version in the Model Registry?

A.To serve as the model used by all users by default
B.To identify the model that took the longest to train
C.To represent a model that is currently being tested
D.To store the source code for the model
AnswerA

The Champion model is intended to be the standard version that applications and users should interact with. It serves as the primary version for production workloads, ensuring that all traffic is routed to the most stable and performant model version currently vetted by the organization's governance processes.

Why this answer

The 'Champion' model is the designated version that currently handles the primary traffic for a deployment. By explicitly labeling a version as the Champion, teams can clearly identify the production-ready model. This provides a single source of truth for deployment automation, ensuring that when an endpoint is configured to serve the 'Champion', it is always running the intended, validated model without manual configuration changes.

Exam trap

Candidates often confuse the 'Champion' version with the latest trained model version, assuming newest always means default, whereas 'Champion' must be explicitly designated for production traffic.

8
Multi-Selecthard

A team is preparing to deploy a model to a Databricks Model Serving endpoint that must scale down to zero replicas when idle yet still serve bursty traffic with acceptable cold-start latency. They also need to capture the request and response payloads for later monitoring. Which two endpoint settings or features should they configure? (Choose two.)

Select 2 answers
A.Configure a provisioned throughput endpoint to reserve dedicated capacity for the model.
B.Increase the workload size to Large to reduce cold-start latency when scaling up from zero.
C.Enable scale-to-zero on the served entity so the endpoint can drop to zero replicas during idle periods.
D.Enable inference tables on the endpoint to log request and response payloads to a Delta table.
E.Set the endpoint's minimum replicas to one so that a warm replica is always available.
AnswersC, D

Scale-to-zero lets the endpoint release all compute when no requests arrive, which directly satisfies the requirement to scale down to zero replicas when idle. When traffic resumes, the endpoint scales back up, incurring a cold start. Because the scenario demands scaling to zero, this setting is necessary; without it the endpoint would keep at least one replica running and continue consuming capacity during idle periods.

Why this answer

The scenario has two distinct requirements: scaling to zero when idle and capturing payloads for monitoring. Scale-to-zero satisfies the first by releasing all compute during idle periods, and inference tables satisfy the second by logging request and response payloads to a Delta table. Keeping a minimum replica or reserving provisioned throughput would prevent scaling to zero, and workload size does not address either requirement.

Exam trap

The trap here is assuming that reducing cold-start latency requires keeping a warm replica, when that choice directly prevents the endpoint from ever scaling down to zero as required.

9
MCQhard

An MLOps engineer is responsible for a Databricks Model Serving endpoint that serves a mission-critical pricing model. The team wants an automated safeguard that detects when live input feature distributions drift away from the training distribution and triggers a retraining workflow, without modifying the model artifact itself. Which Databricks capability should be configured?

A.Increase the endpoint's workload size so it can handle higher request volume.
B.Add a custom preprocessing step to the model that rejects requests outside the training range.
C.Enable inference tables on the endpoint and build a Databricks SQL alert over the logged payloads.
D.Set the model version's stage to Production and rely on stage transitions to signal drift.
AnswerC

Inference tables persist the request and response payloads of the endpoint to a Delta table, which can be queried continuously. Comparing logged feature values against training statistics in SQL and raising an alert on the drift metric gives an automated trigger for retraining, and it requires no change to the model artifact. This directly matches the requirement for a non-invasive drift safeguard.

Why this answer

Inference tables capture the actual request payloads hitting the endpoint into a queryable Delta table, which is the foundation for computing drift metrics against training baselines. Pairing that with a scheduled query or SQL alert creates an automated signal that can start a retraining job, all without touching the model artifact, unlike request rejection, capacity changes or registry stage bookkeeping.

Exam trap

The trap here is confusing endpoint scaling or registry stage changes with monitoring, when only logged live payloads can reveal distribution shift.

10
MCQeasy

When deploying a model to Databricks Model Serving, what is the recommended way to handle sensitive credentials like database connection strings?

A.Include credentials as environment variables in the model's conda.yaml file.
B.Store the credentials in an encrypted text file inside the model package.
C.Use Databricks Secrets to reference credentials within the serving environment.
D.Use a global variable within the model inference script.
AnswerC

Databricks Secrets allow you to reference sensitive information securely during deployment. This approach ensures that credentials remain encrypted at rest and in transit, and are only accessible by authorized users or service principals, aligning with enterprise security standards for cloud-based machine learning deployment workflows and infrastructure.

Why this answer

Using Databricks Secrets is the industry standard for managing sensitive information securely. By decoupling credentials from the model code or deployment configuration, you prevent hard-coding secrets which could be inadvertently exposed. This practice is vital for maintaining security compliance and protecting proprietary data access, as it enables centralized management, auditing, and rotation of keys without necessitating code changes in the model repository.

Exam trap

Candidates frequently choose environment variables or hardcoded strings for credentials, overlooking the security risks and the native Databricks Secrets utility designed for this.

11
MCQhard

An ML engineer deployed a model to a Databricks Model Serving endpoint and enabled inference tables. After a week, the team notices that some requests returned HTTP 200 but the corresponding rows in the inference table show null prediction values. They need to diagnose why predictions are missing for those requests. Which explanation is most consistent with this symptom?

A.The endpoint's autoscaling reduced replicas, causing dropped responses.
B.The model returned a response shape that the serving layer could not map to the prediction column.
C.The endpoint's service principal lacked SELECT permission on the inference table.
D.The inference table was configured with a sampling fraction that excluded those rows.
AnswerB

Inference tables log request and response payloads, and the serving layer extracts predictions based on the expected response structure. If the model returns a dict or array that does not match the expected prediction field, the endpoint can still return HTTP 200 while the logged prediction column is null. This mismatch between the model's output shape and the logging schema is a common cause of missing predictions in inference tables.

Why this answer

Inference tables capture request and response payloads for each served request, and the serving layer derives the prediction column from the model's response structure. When a model returns output in a format the logging layer does not recognize as a prediction, the request still succeeds with HTTP 200 but the logged prediction is null. Sampling, autoscaling, and read permissions produce different symptoms such as missing rows or errors, not present rows with null predictions.

Exam trap

The trap here is assuming a successful HTTP status guarantees a fully populated inference table row, when response-shape mismatches can leave the prediction column null.

12
Multi-Selectmedium

A team is preparing to deploy a registered MLflow model to a Databricks Model Serving endpoint. They want to capture every request and response for later monitoring and debugging, and they also want the endpoint to remain available during a rolling model version update. (Choose two.)

Select 2 answers
A.Enable verbose logging in the MLflow model's Python code and write logs to DBFS.
B.Enable inference tables on the serving endpoint.
C.Delete the existing endpoint and create a new one pointing at the new model version.
D.Configure the endpoint with zero-downtime updates by using a new model version in the same endpoint configuration.
E.Set the endpoint's `min_instances` to 0 to reduce cost during the update.
AnswersB, D

Inference tables automatically log request payloads and response payloads from a Model Serving endpoint into a Delta table in Unity Catalog. Enabling them satisfies the requirement to capture every request and response for monitoring and debugging, and the logged data can be queried directly with SQL for drift analysis or troubleshooting.

Why this answer

Capturing every request and response is achieved by enabling inference tables, which log payloads to a Delta table. Keeping the endpoint available during a model version change is achieved by updating the served version in place, which triggers a rolling update rather than endpoint deletion. Together these satisfy both the observability and availability requirements without downtime or manual log plumbing.

Exam trap

The trap here is assuming that deleting and recreating the endpoint is required to change model versions, when Model Serving supports in-place version updates with rolling rollout.

13
MCQeasy

A machine learning engineer has a model registered in Unity Catalog and wants to expose it as a REST API so an external application can send JSON payloads and receive predictions. The team has no existing serving infrastructure. Which Databricks feature should be used to create this API?

A.A Databricks job that runs a notebook on a schedule and writes predictions to a table.
B.A Databricks SQL warehouse with a query that invokes the model as a user-defined function.
C.A Databricks Model Serving endpoint created from the registered model.
D.An MLflow experiment tracking run that logs the model's parameters and metrics.
AnswerC

Model Serving provisions a managed, autoscaling HTTP endpoint for a registered model and returns predictions from JSON request payloads. It requires no infrastructure work, handles scaling and versioning, and supports the request-response pattern the external application needs. Creating the endpoint from the Unity Catalog model is the direct way to obtain a REST API.

Why this answer

Model Serving is the Databricks capability that turns a registered model into a managed REST endpoint with autoscaling and request-response semantics. Scheduled notebook jobs, experiment tracking runs and SQL warehouses serve other purposes, so they cannot provide the low-friction HTTP API the external application needs for on-demand scoring.

Exam trap

The trap here is equating any way of running a model, such as a scheduled job or SQL function, with a serving endpoint that exposes an HTTP API.

14
MCQmedium

A machine learning engineer has registered a model in the Databricks Model Registry and wants to expose it as a REST API with automatic scaling and no server management. The model's Python dependencies are captured in a conda environment file logged with the run. Which Databricks capability should the engineer use to serve this model with minimal operational overhead?

A.A Databricks job that runs the model on a schedule and writes predictions to a Delta table
B.Databricks Model Serving with a model version URI from the Model Registry
C.A SQL warehouse configured with the model's conda environment installed on all nodes
D.An all-purpose cluster running an MLflow model server process started manually
AnswerB

Databricks Model Serving directly consumes a registered model version URI, builds the serving environment from the logged conda dependencies, and exposes a REST endpoint with managed scaling. It requires no cluster or server administration, matching the requirement for minimal operational overhead while providing low-latency inference for the registered model.

Why this answer

Databricks Model Serving is the managed capability that takes a registered model version URI, reconstructs the environment from logged dependencies, and publishes a scalable REST endpoint without server administration. The other choices describe batch, interactive, or SQL compute that cannot deliver managed low-latency REST inference for a registered MLflow model.

Exam trap

The trap here is assuming any Databricks compute can host a model endpoint, when only Model Serving provides managed REST inference for registered model versions.

15
Multi-Selecthard

A data science team is deploying a model to a Databricks Model Serving endpoint. They want to enable inference logging to capture the input data and predictions for monitoring and debugging. Which TWO configurations are required to enable inference logging for a serving endpoint? (Choose two.)

Select 2 answers
A.Set the endpoint's workload size to at least 'Medium'.
B.Ensure the model signature includes a field for logging.
C.Specify a Delta table to store the inference logs.
D.Enable inference logging in the endpoint configuration.
E.Configure the model to output logs in JSON format.
AnswersC, D

Inference logging requires a target Delta table where the logs will be written. The table must be specified in the endpoint configuration, and the endpoint's service principal must have write permissions to it. This table stores the request and response data for monitoring.

Why this answer

To enable inference logging for a Model Serving endpoint, you must explicitly enable it in the endpoint configuration and specify a Delta table to store the logs. The endpoint's identity needs write access to that table. Other factors like workload size or model signature do not affect logging.

This allows capturing request and response data for monitoring and debugging.

Exam trap

The trap here is assuming that model-level configurations like output format or signature are needed for inference logging, when actually it is a serving endpoint feature requiring explicit enablement and a storage table.

16
MCQhard

Refer to the exhibit. What is the most likely cause of this error in a deployed MLflow model?

A.The serving endpoint has reached its maximum memory capacity.
B.The client application is sending data that does not conform to the defined input schema.
C.The model has been corrupted during the transition to Production.
D.The serving endpoint is not authenticated to access the data.
AnswerB

The error message explicitly points to a schema mismatch where a 'string' was provided instead of a 'double'. This is a direct violation of the model signature, which governs the expected input data types. The client application must be updated to cast the data correctly before sending the request.

Why this answer

The error explicitly indicates a schema mismatch between the client's payload and the model's expected input signature. This frequently occurs when downstream applications send data without proper type casting. Validating the input data before sending it to the serving endpoint is a critical step in production MLOps to prevent type-related failures and ensure robust service interaction between disparate software components in a distributed architecture.

Exam trap

Candidates often blame the model code or the infrastructure environment for schema errors. They overlook the client-side payload, which is the most frequent source of input type mismatches in production.

17
MCQmedium

A machine learning engineer needs to capture every request and response payload sent to a Databricks Model Serving endpoint so that the team can later join predictions with ground-truth labels for monitoring. Which Databricks feature should they enable on the endpoint?

A.Model signature enforcement
B.Inference tables
C.Endpoint access logs in the workspace audit log
D.MLflow model logging to the run's artifact store
AnswerB

Inference tables automatically log request and response payloads from a Model Serving endpoint into a Delta table in Unity Catalog. Each row records the timestamp, request, response, and metadata, enabling downstream joins with ground-truth labels and monitoring of data drift or model quality. Enabling inference tables is the supported way to capture payloads for later analysis.

Why this answer

Inference tables are the Databricks Model Serving feature that persists request and response payloads to a Delta table. Enabling them allows teams to join predictions with ground-truth labels over time, monitor drift, and debug production behavior. Signature enforcement, MLflow artifact logging, and audit logs serve different purposes and do not capture payloads.

Exam trap

The trap here is assuming that enabling the model signature or auditing the endpoint is equivalent to logging inference payloads, when only inference tables persist request and response data.

18
MCQhard

A team deploys a model to a Databricks Model Serving endpoint and enables inference tables. After a week, they notice that the inference table contains request and response payloads but the payload columns are empty for many rows, while status codes are 200. What is the most likely explanation?

A.The endpoint was serving a different model version than the one with the inference table enabled.
B.The model signature was not logged, so Databricks could not map payloads to columns.
C.The endpoint's autoscaling scaled to zero, so requests were not logged.
D.The inference table's payload logging was disabled or the payload size exceeded the logging limit, so only metadata was recorded.
AnswerD

Inference tables can be configured to log metadata only, and payloads that exceed the maximum logged size are truncated or omitted. When status codes are 200 but payload columns are empty, the most likely cause is that payload logging is disabled in the endpoint configuration or that requests exceeded the size threshold, leaving only metadata rows.

Why this answer

Inference tables have a payload logging setting that can be enabled or disabled, and payloads above a size threshold are not fully recorded. When status codes are successful but payload columns are empty, the cause is almost always that payload logging is off or the requests exceeded the size limit. Reviewing the endpoint's inference table configuration and the payload sizes resolves the issue.

Exam trap

The trap here is assuming that any inference table row includes payloads, when payload logging is a separate setting and large payloads may be truncated or omitted.

19
MCQhard

Which strategy is most effective for managing model drift in a production Databricks environment?

A.Manually check the model accuracy once a year
B.Re-deploy the model with new data every hour
C.Log inference data to Delta tables for analysis
D.Use a static model that never changes
AnswerC

Logging inference inputs and predictions to Delta tables creates an audit trail that enables monitoring. This data can be analyzed to measure drift in input features or model outputs. This is the industry-standard approach in Databricks for building a feedback loop that informs when a model needs to be updated.

Why this answer

Managing model drift requires active monitoring of input data and output predictions. Setting up Databricks Model Serving with feature tables allows for logging of inference requests and responses to a Delta table. By comparing these production distributions against the training data using tools like Great Expectations or custom SQL queries, teams can detect shifts in data patterns that necessitate retraining or model adjustments to maintain performance.

Exam trap

Candidates often select real-time model re-training or offline batch metrics aggregation, missing that logging inference data to Delta tables is the core requirement for drift analysis.

20
MCQhard

A team maintains a Databricks Model Serving endpoint for a fraud model. Compliance requires that every request and response be logged to a Delta table for auditing and later analysis. The endpoint is already configured and serving traffic. What should the team do to capture this data with the least additional infrastructure?

A.Configure the endpoint to emit metrics to a monitoring dashboard and export the dashboard data nightly
B.Write a client wrapper that logs each request and response to a Delta table before returning the prediction
C.Attach an event log delivery to the serving endpoint and parse the logs into a Delta table with a scheduled job
D.Enable inference tables on the endpoint so requests and responses are automatically logged to a Delta table
AnswerD

Inference tables are a built-in Model Serving feature that automatically logs the request payloads and model responses to a Delta table in Unity Catalog. Enabling them requires no extra pipelines or clusters and directly satisfies the auditing requirement, making it the least-infrastructure solution for capturing endpoint traffic.

Why this answer

Inference tables are the native Model Serving mechanism that records request and response payloads into a Delta table, giving complete server-side auditing without extra infrastructure. Client-side logging, dashboard metrics, and event logs each capture only partial or operational data, so they cannot satisfy the requirement to log every request and response for the fraud endpoint.

Exam trap

The trap here is confusing operational endpoint logs and metrics with inference tables, which are the feature that actually captures request and response payloads.

21
MCQhard

A team has an existing Databricks Model Serving endpoint serving `prod.ml.fraud_model` version 3. They register version 4, which uses a new feature set, and want to shift only 10% of traffic to version 4 while keeping version 3 for the rest. Their endpoint currently has a single served entity for version 3. What is the most appropriate approach?

A.Register version 4 under a new model name and point the existing served entity at the new model name.
B.Add a second served entity for version 4 to the same endpoint and configure traffic splitting between the two served entities.
C.Update the existing served entity's entity_version to 4 and rely on the endpoint's automatic gradual rollout.
D.Create a second endpoint for version 4 and configure a load balancer outside Databricks to split traffic 10/90 between endpoints.
AnswerB

Databricks Model Serving supports multiple served entities on one endpoint with configurable traffic percentages. Adding a served entity for version 4 and setting its traffic share to 10% achieves the canary rollout while version 3 continues to receive the remaining traffic.

Why this answer

Traffic splitting in Databricks Model Serving is achieved by defining multiple served entities on a single endpoint and assigning each a traffic percentage. Adding a served entity for version 4 at 10% and leaving version 3 at 90% implements the desired canary rollout. Swapping the served entity version, using an external load balancer, or renaming the model do not provide the controlled split.

Exam trap

The trap here is assuming that updating a served entity to a new version performs a gradual rollout, when in fact it replaces the served version entirely.

22
MCQmedium

When deploying a model to a production environment, why is it critical to create a dedicated 'staging' environment before the 'production' environment?

A.To increase the cost of the deployment infrastructure.
B.To perform end-to-end integration testing in a production-like environment.
C.To hide the model from the model registry.
D.To store backup copies of the training dataset.
AnswerB

Staging environments are designed to replicate the production environment as closely as possible. This allows for testing the entire pipeline, including data connectivity, API endpoints, and model performance, ensuring that any issues are detected and resolved before the model is exposed to actual production traffic and users.

Why this answer

A staging environment acts as a vital quality gate that mirrors the production configuration. By testing the model in a staging environment, developers can identify integration issues, resource bottlenecks, and environmental discrepancies without impacting actual end-users. This practice reduces risk, ensures the model behaves predictably under load, and allows for thorough regression testing, which is essential for maintaining high availability and reliability in mission-critical machine learning applications.

Exam trap

Candidates often think staging environments are meant solely for hyperparameter tuning or final model training, ignoring their role as integration testing gates.

23
MCQhard

A data scientist deploys a model to a Databricks Model Serving endpoint and enables inference tables. After a week, they want to analyze prediction drift by joining the logged requests with ground-truth labels that arrive later. Which statement describes how they should access the inference table data for this analysis?

A.Read the endpoint's driver logs from the cluster's log delivery location, parse the JSON entries, and join them with the labels table.
B.Query the inference table using the Delta table name shown in the endpoint's Serving UI, then join it with a labels table on a request identifier.
C.Call the endpoint's `/metrics` REST API to export inference records, then load them into a Spark DataFrame for joining with labels.
D.Enable model monitoring on the endpoint and rely on its automatically generated drift metrics, which replace the need to join with ground-truth labels.
AnswerB

Inference tables are stored as Delta tables in Unity Catalog or the Hive metastore, and the endpoint's UI exposes the fully qualified table name. Users can query them with SQL or Spark and join to downstream label tables using the request ID column that Databricks includes. This enables drift analysis and model monitoring. The table contains request payloads, responses, and metadata such as timestamps and model version.

Why this answer

Inference tables persist request and response payloads as Delta tables, with a fully qualified name visible in the endpoint UI. Analysts can query them directly and join with delayed ground-truth labels using the included request ID and timestamps. This supports drift and accuracy analysis without custom logging.

Lakehouse Monitoring can build on these tables for automated metrics.

Exam trap

The trap here is assuming inference data is only available through monitoring dashboards or REST metrics, rather than as a queryable Delta table.

24
MCQmedium

A machine learning engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model was logged with MLflow using the default signature and input example. After deployment, the engineer notices that the endpoint's REST API expects a JSON payload in a specific format. Which MLflow artifact is used by Model Serving to determine the expected input format for the endpoint?

A.The conda.yaml file, which lists the Python dependencies.
B.The requirements.txt file, which pins package versions.
C.The MLmodel file's signature field, which contains the input schema.
D.The input_example.json file, which provides a sample input.
AnswerC

The MLmodel file is generated by MLflow and includes a signature field that defines the input and output schema. Model Serving reads this to validate and parse incoming requests, ensuring the JSON payload matches the expected columns and types. This enables automatic request validation and correct inference.

Why this answer

Model Serving relies on the MLflow model's signature, stored in the MLmodel file, to understand the expected input schema. This signature includes column names, data types, and shapes, allowing the endpoint to validate and parse JSON requests correctly. Without a signature, the endpoint may not enforce input validation, leading to errors.

The signature is essential for seamless deployment.

Exam trap

The trap here is assuming that the input example or dependency files define the API input format, when actually the model signature in the MLmodel file is the authoritative source.

25
MCQhard

A data scientist registers an MLflow model whose `conda.yaml` lists several Python packages. When they create a Databricks Model Serving endpoint from this model, the deployment fails during environment build. Which action is most likely to resolve the failure while preserving the model's dependency requirements?

A.Remove all pinned package versions from the `conda.yaml` so the environment builder can pick any compatible release.
B.Increase the endpoint's workload size to Large so the environment builder has more memory to resolve the dependency graph.
C.Add the missing or incompatible packages to the endpoint's environment by specifying them in the model's `requirements.txt` or adjusting the `conda.yaml` to versions available in the Databricks runtime.
D.Convert the model to a different flavor, such as replacing scikit-learn with a PyTorch implementation, so the dependencies are no longer needed.
AnswerC

Model Serving builds the environment from the dependency specification captured with the model. If a package is missing or pinned to a version that cannot be resolved, the build fails. Correcting the `requirements.txt` or aligning `conda.yaml` entries with versions available in the Databricks runtime lets the builder resolve the environment while still installing the libraries the model needs, preserving the model's required dependencies.

Why this answer

Model Serving reconstructs the model's environment from the dependency files captured at logging time. When a required package is absent or a pinned version cannot be resolved, the build fails. Adjusting the `requirements.txt` or `conda.yaml` to include the necessary packages at versions compatible with the Databricks runtime lets the environment build succeed while keeping the dependencies the model genuinely needs.

Exam trap

The trap here is assuming that enlarging the endpoint's workload size will fix any deployment error, when environment build failures stem from the dependency specification rather than compute capacity.

26
MCQhard

A team observes that their Databricks Model Serving endpoint occasionally returns HTTP 429 responses during bursty traffic, even though the endpoint shows low average CPU utilization. They want to reduce these throttling errors without over-provisioning capacity. Which action is most appropriate?

A.Switch the endpoint to a smaller workload size so individual replicas handle requests faster.
B.Enable scale-to-zero and rely on autoscaling to handle bursts.
C.Increase the endpoint's `max_instances` and tune the autoscaling concurrency target so replicas scale out faster during bursts.
D.Increase the endpoint's `min_instances` so that more replicas are always warm.
AnswerC

HTTP 429 responses from Model Serving typically indicate that the per-replica request queue is saturated even though average CPU is low, because bursts arrive faster than replicas can accept them. Raising `max_instances` allows more replicas to spin up, and tuning the concurrency target makes the autoscaler react sooner, absorbing bursts before requests are rejected.

Why this answer

Throttling at low average CPU points to concurrency saturation during short bursts. Allowing the endpoint to scale to more replicas and configuring autoscaling to react to concurrency sooner lets the endpoint absorb spikes before requests are rejected. This addresses the root cause without permanently over-provisioning capacity during steady-state periods.

Exam trap

The trap here is equating low average CPU with sufficient capacity, when 429s during bursts are caused by concurrency queue saturation rather than sustained compute pressure.

27
MCQeasy

A data scientist has a custom Python model wrapped in an MLflow pyfunc flavor and needs to serve it on Databricks Model Serving. The model's preprocessing requires a library that is not part of the default serving environment. What is the correct way to make that dependency available to the endpoint?

A.Upload the library wheel to DBFS and reference its path in the endpoint's environment variables.
B.Log the model with the required library listed in the model's conda environment or pip requirements so it is captured in the model artifact.
C.Add the library as a cluster library on the all-purpose cluster where the model was trained.
D.Install the library interactively on the driver node before creating the endpoint.
AnswerB

MLflow records dependencies when the model is logged, and Model Serving rebuilds the environment from those captured requirements. Adding the library to the model's conda environment or pip requirements at log time ensures the endpoint installs it during container build. This is the supported, reproducible mechanism for custom dependencies in served models.

Why this answer

Model Serving reconstructs the runtime environment from the dependency specification stored with the logged model. To bring in a library beyond the default environment, the dependency must be captured in the model's conda environment or pip requirements when the model is logged, which guarantees the endpoint installs it reproducibly. Cluster-level installs and storage paths do not reach the serving container.

Exam trap

The trap here is assuming the serving environment inherits packages from the cluster that trained the model, when the endpoint instead rebuilds dependencies solely from what was recorded with the logged model.

28
MCQmedium

A data scientist has registered a scikit-learn model in Unity Catalog as `ml_prod.churn.model_v3` and wants the Databricks Model Serving endpoint to automatically pick up newly registered model versions as they are promoted to the `champion` alias. Which configuration should the data scientist use when creating the serving endpoint?

A.Serve the model using the alias-qualified path `ml_prod.churn.model_v3@champion`.
B.Export the model with `mlflow.sklearn.save_model` to a DBFS path and point the endpoint at that directory.
C.Serve the model by its explicit version number, `ml_prod.churn.model_v3/3`, and re-create the endpoint after each promotion.
D.Serve the model by its registered name only, `ml_prod.churn.model_v3`, and rely on Databricks to always use the latest version.
AnswerA

Databricks Model Serving supports Unity Catalog alias-qualified model paths. Pointing the endpoint at `model_v3@champion` causes the serving infrastructure to resolve the alias at request time, so when a new version is promoted to the `champion` alias, the endpoint automatically serves the updated artifact without endpoint reconfiguration or downtime.

Why this answer

Alias-qualified paths in Unity Catalog let a serving endpoint follow a movable pointer. When a new model version is registered and the `champion` alias is reassigned to it, the endpoint resolves the alias and begins serving the new artifact. This enables zero-touch promotion workflows, keeps governance in Unity Catalog, and avoids endpoint recreation or explicit version pinning.

Exam trap

The trap here is assuming that a bare registered model name automatically resolves to the latest version, when Databricks Model Serving requires an explicit version or alias-qualified path.

29
MCQmedium

An ML engineer needs to give an external application a stable HTTPS URL to call a registered model served by Databricks Model Serving. The application must authenticate with a token and must not be able to modify the endpoint configuration. Which approach best meets these requirements?

A.Grant the external application a service principal with a token and assign it only the permissions required to query the serving endpoint.
B.Expose the endpoint through a public network with no authentication so the external application can call it directly with the URL alone.
C.Embed the model's artifacts in the external application and call the model locally instead of using the serving endpoint.
D.Share the serving endpoint's URL and provide the external application with a Databricks personal access token belonging to a workspace admin.
AnswerA

Databricks Model Serving endpoints expose a stable HTTPS URL that clients can call with bearer-token authentication. Using a dedicated service principal scoped to query permissions satisfies authentication without granting configuration rights, honoring least privilege. This gives the external application a durable, callable endpoint while ensuring it cannot alter the endpoint or reach unrelated workspace resources, which is exactly what the scenario requires.

Why this answer

Model Serving endpoints are invoked over a stable HTTPS URL using bearer-token authentication. Issuing the external application a service principal token scoped to query the endpoint satisfies the authentication requirement while preventing configuration changes, since the principal lacks management permissions. Sharing an admin token over-privileges the client, and the remaining choices either bypass the endpoint or remove authentication entirely.

Exam trap

The trap here is treating any working token as acceptable, when the requirement that the client cannot modify the endpoint makes a narrowly scoped service principal the only suitable credential.

30
Multi-Selectmedium

A machine learning engineer is preparing to deploy a registered model to a Databricks Model Serving endpoint. Before creating the endpoint, the engineer wants to confirm the deployment prerequisites are satisfied. Which two conditions are required for a successful endpoint creation? (Choose two.)

Select 2 answers
A.The model version has a logged MLflow model artifact with a valid model signature
B.The model version is transitioned to the Production stage in the Model Registry
C.The endpoint name matches the registered model name exactly
D.A dedicated all-purpose cluster is running and attached to the endpoint
E.The serving environment can be reconstructed from the model's logged dependencies
AnswersA, E

Model Serving requires a logged MLflow model artifact it can load, and a model signature defines the expected input and output schema. Without a signature, the endpoint cannot validate incoming request payloads, so a properly logged model with a signature is a genuine prerequisite for successful endpoint creation.

Why this answer

A successful endpoint deployment depends on a loadable MLflow model artifact with a defined signature and a reconstructable dependency environment, because the service must build a container and validate request schemas. Registry stage, attached clusters, and name matching are not technical prerequisites and do not affect whether the endpoint can be created and brought to a ready state.

Exam trap

The trap here is treating Model Registry stages as deployment gates, when serving actually keys off the model version URI and logged artifacts.

31
MCQhard

A team wants to route production traffic to a new model version while keeping risk low. They configure a Databricks Model Serving endpoint with two served entities: `champion` (entity_version 5) and `challenger` (entity_version 6). They want 95% of requests to hit `champion` and 5% to hit `challenger`. Which configuration accomplishes this?

A.Set `traffic_config` with routes assigning `traffic_percentage` 95 to `champion` and 5 to `challenger`.
B.Deploy two separate endpoints and use a client-side load balancer to send 95% of calls to the champion endpoint.
C.Set `scale_to_zero_enabled` to true only on `challenger` so it receives proportionally fewer requests.
D.Set `workload_size` to `Small` on `champion` and `Large` on `challenger` to bias traffic toward the larger entity.
AnswerA

Databricks Model Serving supports `traffic_config` with a `routes` array, where each entry names a served entity and specifies an integer `traffic_percentage`. Summing to 100 across routes enables canary or A/B routing. Setting 95 for champion and 5 for challenger delivers the desired split. This is the supported mechanism for gradual rollout without redeploying the endpoint.

Why this answer

Databricks Model Serving uses `traffic_config` with a `routes` list, where each route references a served entity by name and assigns an integer `traffic_percentage`. Summing route percentages to 100 achieves the desired 95/5 split. Other configuration fields like workload size and scale-to-zero govern resources and scaling behavior, not request distribution.

Exam trap

The trap here is confusing resource-allocation settings such as workload size or scale-to-zero with traffic-routing settings, which are configured separately in `traffic_config`.

32
MCQhard

A team has a Databricks Model Serving endpoint configured with scale-to-zero enabled and min_instances set to 0. During a load test, they observe that the first request after an idle period takes roughly 40 seconds while subsequent requests complete in under 200 milliseconds. They need to eliminate this cold-start latency for a customer-facing application without over-provisioning. Which configuration change best addresses the requirement?

A.Increase max_instances so more replicas can absorb the burst.
B.Increase the workload size from Small to Large.
C.Enable route optimization in the endpoint configuration.
D.Set min_instances to 1 so at least one instance stays warm.
AnswerD

Setting min_instances to 1 keeps one replica provisioned at all times, so the container and model are already loaded and the first request after an idle period is served without a cold start. This is the documented way to trade a small amount of always-on compute for predictable low latency, and it does not require scaling the endpoint beyond what the workload needs during quiet periods.

Why this answer

Cold-start latency occurs because a scaled-to-zero endpoint must provision a container and load the model before serving the first request. Keeping one replica always available by setting min_instances to 1 removes that provisioning and loading delay while still allowing the endpoint to scale up under load. Increasing max_instances, changing workload size, or enabling routing features does not prevent the endpoint from scaling to zero during idle periods.

Exam trap

The trap here is confusing max_instances, which governs burst capacity, with min_instances, which governs whether any replica stays warm when traffic drops to zero.

33
MCQmedium

What is the primary role of an 'MLflow Signature' during the model deployment phase?

A.To encrypt the model artifacts for security
B.To define the input and output schema for validation
C.To determine the optimal number of cluster nodes
D.To compress the model for faster deployments
AnswerB

The signature specifies the required input features and the output format. This metadata is used by the serving endpoint to validate incoming JSON requests, ensuring they match the expected schema. This prevents invalid data from causing failures during the inference process, which is critical for production stability.

Why this answer

The MLflow signature acts as a contract between the model and the application that consumes it. By defining expected input types and shapes, the signature enables automatic validation of incoming payloads before they reach the model. This prevents runtime errors, such as type mismatches or missing features, and ensures that the serving endpoint remains robust and reliable when receiving production traffic from various sources.

Exam trap

Candidates often mistake the signature for a security feature or an authentication method. They fail to recognize it is primarily a schema contract for data validation.

34
MCQeasy

A data scientist wants to test a newly registered model version interactively before promoting it to production. They need to send a sample request to the model and inspect the prediction and the model's input schema. Which Databricks feature should they use?

A.The Databricks Jobs UI run history for the training job.
B.The Serving tab of the registered model in Unity Catalog, which provides a test request interface.
C.The `mlflow.pyfunc.load_model` function in a notebook.
D.The Model Registry's version description field.
AnswerB

The Serving tab in the model version UI provides an interactive interface where a user can send a sample request to the model and view the response. It also displays the model signature, including input and output schema, so the data scientist can confirm expected fields before promoting the version to production.

Why this answer

The Serving tab on a registered model version in Unity Catalog offers a built-in test interface. It lets the user submit a sample request, see the prediction, and review the model signature that defines expected inputs and outputs. This provides a quick validation step before the version is promoted or attached to a production endpoint.

Exam trap

The trap here is confusing notebook-based model loading with the serving test interface, when only the Serving tab exercises the actual serving path and shows the signature.

35
MCQmedium

A data scientist has registered a scikit-learn model in Unity Catalog and now wants to serve it behind a Databricks Model Serving endpoint. The model's MLflow signature records a pandas DataFrame input with three named columns. The team wants the endpoint to reject malformed requests automatically rather than silently scoring them. Which action should the data scientist take?

A.Enable inference tables on the endpoint so malformed records are flagged in the Delta table and excluded from scoring.
B.Register the model with a new signature that uses a tensor spec instead of a DataFrame spec so the endpoint validates each element.
C.Add a Python preprocessing script that parses the raw JSON and raises an error for unexpected fields, then deploy that script as the endpoint's model.
D.Create the endpoint with the registered model version; Model Serving enforces the logged MLflow signature and schema validation by default.
AnswerD

Model Serving reads the MLflow signature stored with the registered model version and uses it to validate incoming payloads against the expected DataFrame schema. Requests whose columns or types do not match are rejected before inference runs, satisfying the requirement without extra configuration.

Why this answer

When a model is logged with an MLflow signature, Databricks Model Serving uses that signature to validate incoming requests. A pandas DataFrame signature with named columns causes the endpoint to check column names and types, rejecting payloads that do not conform. No custom preprocessing or inference table configuration is required, because validation happens at the serving layer before the model is invoked.

Exam trap

The trap here is assuming that schema validation requires custom code or inference tables, when Model Serving already enforces the logged MLflow signature automatically.

36
MCQmedium

A team queries a Databricks Model Serving endpoint through the serving client and receives an error indicating the endpoint is not ready. They confirmed the endpoint exists. Which condition most directly explains why requests fail until it clears?

A.The endpoint is in a state such as Updating or Not Ready while its configuration is being applied or a replica is still building.
B.The caller's personal access token has expired.
C.The model version referenced by the endpoint has been deleted from the registry.
D.The serving client is using the wrong workspace URL for the endpoint.
AnswerA

Endpoints move through provisioning and updating states, and during those transitions the service may reject requests because no ready replica can serve them. Once the endpoint reaches Ready, traffic is accepted. This is the most direct explanation for an existing endpoint returning a not-ready error, and it clears without any client-side change.

Why this answer

When an endpoint is being created or its configuration is being applied, it passes through transitional states and cannot serve requests until at least one replica is ready. A not-ready error for an endpoint known to exist points to that provisioning or updating state, which resolves on its own once the deployment completes and the endpoint reports Ready.

Exam trap

The trap here is attributing a not-ready response to an authentication or registry problem, when it actually reflects the endpoint's own lifecycle state while replicas are still being provisioned.

37
MCQmedium

When deploying a model to a production endpoint, what is the best practice for handling dependencies?

A.Manually install dependencies in the cluster terminal
B.Include a requirements.txt file in the model artifact
C.Use the 'latest' tag for all library imports
D.Only use libraries that come pre-installed in the Databricks Runtime
AnswerB

Including a requirements.txt file or letting MLflow capture the environment ensures that the serving environment mirrors the training environment. This is the standard method to maintain consistency, allowing the Databricks serving infrastructure to install the correct package versions during the initial container build and deployment process.

Why this answer

The best practice is to ensure all dependencies are explicitly defined in the MLflow model artifact. Using custom environment files or letting MLflow infer them ensures that the production container is built using the exact versions of packages used in training. This minimizes the risk of production-level errors caused by version mismatches, ensuring stability and reproducibility across the model lifecycle in the Databricks ecosystem.

Exam trap

Candidates often assume that the environment is handled automatically by the cluster. They forget that production endpoints need explicit, version-controlled dependency manifests like requirements.txt to ensure consistency.

38
Multi-Selecthard

A team is deploying a batch scoring pipeline that loads a registered MLflow model and runs predictions over a large Delta table using Spark. They want the scoring job to reuse the model's training-time preprocessing and to remain reproducible months later. Which two practices should they follow? (Choose two.)

Select 2 answers
A.Reference the model by its numeric version when loading it in the batch job, so the exact artifact used in validation is always retrieved.
B.Load the model using the `latest` alias so the batch job always benefits from the most recent improvements automatically.
C.Log the model with `mlflow.pyfunc.log_model` (or framework flavor) including the model signature and input example so the schema and preprocessing are captured with the artifact.
D.Store the preprocessing code in a separate notebook and import it at scoring time so the batch job and training job share the same functions.
E.Convert the model to a pickle file, store it in a Delta table as a binary column, and load it via Spark at scoring time.
AnswersA, C

Loading a specific numeric version pins the job to an immutable artifact, guaranteeing that reruns months later use the identical weights and preprocessing. This is the standard way to achieve reproducibility for batch scoring, and it complements the signature metadata. Relying on a moving alias would instead silently change behavior whenever the alias is reassigned, breaking reproducibility.

Why this answer

Reproducible batch scoring requires an immutable artifact reference and self-describing metadata. Logging the model with a signature and input example embeds schema and preprocessing, while loading a specific numeric version pins the exact artifact so reruns match validation. Moving aliases, external notebook imports, and raw pickles all introduce drift or lose the metadata needed to reproduce training-time behavior.

Exam trap

The trap here is treating an alias like a version pin, when aliases intentionally move and therefore cannot provide the immutability that reproducible batch scoring requires.

39
Multi-Selecthard

Which THREE factors should be considered when choosing the 'workload size' (e.g., Small, Medium, Large) for a Databricks Model Serving endpoint?

Select 3 answers
A.The memory requirements of the model artifact
B.The total number of users in the workspace
C.Expected request latency targets
D.The volume of expected incoming traffic
E.The color scheme of the MLflow UI
AnswersA, C, D

The model artifact size and its runtime memory consumption directly dictate the minimum compute resources needed. If the model requires more RAM than the chosen workload size provides, the endpoint will fail to load or experience frequent crashes. Assessing memory usage during the testing phase is critical for size selection.

Why this answer

Selecting the correct workload size is a trade-off between latency requirements, compute costs, and the model's resource footprint. Larger models with high parameter counts often require 'Large' configurations to prevent memory exhaustion or high tail latency. Conversely, smaller models may perform adequately on 'Small' instances, which are cost-effective.

Monitoring the endpoint's resource utilization metrics is key to right-sizing the deployment and ensuring optimal performance within the budget constraints.

Exam trap

Candidates often select cost or model accuracy instead of operational metrics like memory footprint, traffic volume, and latency targets when sizing endpoints.

40
MCQmedium

A data scientist needs to deploy a model to Databricks Model Serving. Which component is strictly required to be logged in MLflow to enable the 'Model Serving' feature?

A.The model's training accuracy metrics
B.The model signature defining input and output schema
C.A Unity Catalog registered function
D.A dedicated high-concurrency cluster
AnswerB

The model signature provides the necessary schema metadata for the serving endpoint. This allows Databricks to enforce input validation for all incoming REST API requests. By defining the signature during the log_model call, you ensure the serving container knows how to translate JSON payloads into the correct format.

Why this answer

Databricks Model Serving requires a model artifact to be logged with a signature. The signature defines the expected input schema, which allows the serving endpoint to validate incoming requests. Without a signature, the model cannot be correctly parsed by the serving infrastructure, preventing the deployment from starting.

This ensures that the production inference endpoint operates with expected data formats, maintaining reliability in downstream applications.

Exam trap

Candidates often believe that logging the model object alone is sufficient for serving. They overlook the signature, which is mandatory for the serving endpoint to validate input schemas.

41
MCQmedium

A fraud detection team wants their Databricks Model Serving endpoint to log every request and response payload to a Unity Catalog Delta table so analysts can later join predictions with ground-truth labels. Which endpoint capability should they enable?

A.Delta Live Tables pipelines attached to the endpoint's event stream.
B.Model monitoring with a Lakehouse Monitor created on the endpoint's metrics.
C.Verbose cluster logs enabled on the underlying serving compute.
D.Inference tables, configured with the endpoint's request and response logging.
AnswerD

Inference tables capture the request payload, the response, and associated metadata into a Delta table in Unity Catalog. Enabling this on the endpoint gives the fraud team an auditable record of each scored transaction that can be joined later with confirmed fraud labels. This is the native, governed mechanism for payload-level observability on Model Serving endpoints.

Why this answer

Inference tables are the supported way to persist per-request payloads and responses from a Model Serving endpoint into a governed Delta table. Because the table lives in Unity Catalog, downstream jobs can join captured predictions with delayed fraud labels to measure real-world performance. This closes the loop between online scoring and offline evaluation without custom logging code in the model.

Exam trap

The trap here is confusing monitoring tools that summarize metrics with the payload capture mechanism that records raw requests and responses.

42
Multi-Selectmedium

A platform team is rolling out a new Databricks Model Serving endpoint for a churn model. They must ensure the endpoint can be queried by an external application and that only authorized callers can invoke it. Which TWO actions should they take? (Choose two.)

Select 2 answers
A.Expose the endpoint only through a cluster-scoped init script that writes the scoring URL into the application's configuration.
B.Grant the calling principal permission on the served model or endpoint and issue a Databricks personal access token or OAuth token for authentication.
C.Disable authentication on the endpoint so the external application can call the REST API without credentials.
D.Enable inference tables on the endpoint so that requests are logged and therefore implicitly authenticated.
E.Create a service principal, grant it appropriate permissions, and have the external application authenticate as that service principal.
AnswersB, E

Model Serving endpoints are protected by Databricks authentication and Unity Catalog or workspace permissions. Granting the caller permission and supplying a bearer token ensures requests are authenticated and authorized. This combination is the supported way to expose an endpoint to an external application while enforcing access control, satisfying both the reachability and the authorization requirements in the scenario.

Why this answer

Securing an endpoint for external use requires both an authenticated identity and a permission grant. A service principal with an OAuth token is the preferred machine identity, and granting it permission on the endpoint or served model enforces authorization. Using a personal access token for a human or service identity works similarly.

Together these actions let the external application reach the endpoint while ensuring unauthorized callers are rejected.

Exam trap

The trap here is assuming that enabling logging or another observability feature provides access control, when authentication and permissions must be configured explicitly.

43
MCQmedium

When using Databricks Model Serving, what is the primary benefit of using a 'Provisioned Throughput' endpoint over a 'Serverless' endpoint?

A.It is always the cheapest option for low-traffic models.
B.It provides guaranteed performance and lower latency for heavy workloads.
C.It eliminates the need to provide a model signature.
D.It is the only way to deploy non-Python models.
AnswerB

Provisioned Throughput reserves dedicated resources for the model, which eliminates the variability associated with shared serverless infrastructure. This results in consistent, predictable performance and lower latency, making it the ideal choice for production applications that must handle high-volume traffic with strict performance and reliability requirements.

Why this answer

Provisioned Throughput is designed for high-performance use cases requiring guaranteed performance, such as low-latency requirements or high-throughput scenarios that cannot tolerate fluctuations in resource availability. By reserving dedicated capacity, it provides the predictable performance needed for enterprise-level applications, ensuring that critical models meet their SLAs despite external load variations. This is a crucial choice for models that require consistent performance regardless of traffic volume patterns.

Exam trap

Candidates often select serverless endpoints for everything, forgetting that Provisioned Throughput provides guaranteed performance and lower latency for heavy enterprise workloads.

44
MCQmedium

A machine learning engineer is configuring a Databricks Model Serving endpoint for a model that requires GPU acceleration. They set the workload size to 'GPU_Medium' but the endpoint fails to deploy. Which of the following is the most likely cause of the failure?

A.The model was not logged with the 'mlflow.pyfunc' flavor.
B.The model's signature does not include GPU-specific metadata.
C.The endpoint was configured with 'min_instances' set to 0.
D.The Databricks workspace does not have GPU instances enabled or available.
AnswerD

Model Serving GPU workload sizes require that the workspace has GPU instances enabled and sufficient quota. If GPU instances are not available or the quota is exceeded, the endpoint deployment will fail. This is a common configuration issue when attempting to use GPU-accelerated serving.

Why this answer

GPU workload sizes for Model Serving require the workspace to have GPU instances enabled and available. If the workspace lacks GPU capacity or the user lacks permissions, deployment fails. Other factors like flavor or signature are unrelated to GPU provisioning.

Ensuring GPU availability is a prerequisite for GPU-accelerated endpoints.

Exam trap

The trap here is focusing on model artifacts like signature or flavor, when the issue is actually about the underlying compute infrastructure and quota for GPU instances.

45
MCQmedium

Refer to the exhibit. What is the impact of setting 'auto_capture_request_payload' to true?

A.It automatically retrains the model whenever new data is captured.
B.It stores incoming request data in a table for analysis and monitoring.
C.It limits the endpoint to only accept JSON payloads.
D.It encrypts the payload before it reaches the model.
AnswerB

This feature facilitates the collection of inference inputs and outputs into a structured format, such as a Delta table. This allows data scientists to monitor the model's performance in production, conduct drift analysis, and use the data for future fine-tuning or retraining efforts, which is vital for maintaining model accuracy.

Why this answer

Enabling 'auto_capture_request_payload' is a powerful feature for debugging, auditing, and continuous model improvement. It allows teams to log the exact data received by the endpoint, enabling them to analyze prediction inputs, perform retrospective bias analysis, and verify data quality. This is a critical component of a robust MLOps strategy, as it provides the raw data needed to identify the causes of poor predictions and justify model decisions for compliance.

Exam trap

Candidates often assume this setting automatically triggers model retraining or performance alerts. It only logs data; it does not perform automated analytical processing or model updates on the captured payload.

46
MCQmedium

A data scientist has registered a scikit-learn model in the Databricks Model Registry as `prod.churn_model`. The production endpoint serving this model must automatically roll back to the previously served version if the newly deployed version's error rate exceeds a threshold within one hour of deployment. Which Databricks feature should the data scientist configure to meet this requirement?

A.Use MLflow Model Registry webhooks to trigger a rollback when the endpoint's error rate exceeds the threshold.
B.Set up a Databricks Lakehouse Monitoring quality monitor on the endpoint's inference table with a custom metric and an alert that triggers a rollback via a job.
C.Configure the endpoint to serve the new version with a small percentage of traffic and manually promote it after validating the error rate.
D.Enable automatic model version rollback on the serving endpoint by defining a rollback policy with the error-rate threshold.
AnswerB

This is the correct approach because Databricks Lakehouse Monitoring can compute custom metrics from the endpoint's inference table, and alerts can trigger a Databricks job. That job can then update the endpoint configuration to revert to the previous model version. This provides the required automatic, threshold-based rollback within the specified time frame using native Databricks capabilities.

Why this answer

Automatic rollback based on runtime error rate is not a built-in Model Serving feature. The viable path is to monitor the endpoint's inference table with Lakehouse Monitoring, define a custom metric for error rate, and use an alert to trigger a job that reconfigures the endpoint to the previous version. This leverages Databricks-native monitoring and orchestration to meet the requirement.

Exam trap

The trap here is assuming Model Serving includes built-in automatic rollback policies based on error-rate thresholds, when in fact rollback must be orchestrated through monitoring and jobs.

47
Multi-Selecthard

A team is standing up a real-time Databricks Model Serving endpoint for a fraud model. Requests will carry several numeric features, and the team wants the endpoint to reject malformed payloads with a clear client error rather than silently scoring them, and to avoid cold-start latency during business hours. Which two actions should the team take? (Choose two.)

Select 2 answers
A.Configure a non-zero minimum number of serving instances so replicas stay warm.
B.Set the endpoint's scale-to-zero behaviour so idle capacity is released between requests.
C.Register the model in the workspace Model Registry instead of Unity Catalog.
D.Enable inference tables to persist the request and response payloads.
E.Log the model with an MLflow model signature that declares the input schema and types.
AnswersA, E

Keeping a minimum instance count above zero holds provisioned, model-loaded replicas ready at all times. Requests arriving during business hours are handled immediately rather than waiting for a container to start and load artifacts, which removes the cold-start delay. This is the standard control for latency-sensitive endpoints that must respond predictably even after periods of low traffic.

Why this answer

Declaring an MLflow model signature gives the endpoint a schema contract so invalid payloads are refused before scoring, and holding a minimum instance count above zero keeps replicas loaded so requests are not delayed by provisioning. Together these satisfy both the fail-fast validation goal and the warm-capacity latency goal, whereas scale-to-zero, inference tables and registry choice affect cost or governance only.

Exam trap

The trap here is treating scale-to-zero as a pure win, when it directly conflicts with a requirement to avoid cold starts.

48
MCQhard

A team serves a model with Databricks Model Serving and wants to send production traffic to a newly registered model version while keeping the ability to revert instantly if quality degrades. They prefer not to edit the endpoint configuration to switch versions. Which approach best fits this requirement?

A.Update the endpoint to reference the new model version by number and rely on the previous version remaining in the registry.
B.Deploy a separate endpoint per model version and update the client to choose which endpoint to call.
C.Create a model alias such as Champion in Unity Catalog, point the served entity at that alias, and move the alias between versions to promote or roll back.
D.Use the workspace model registry and transition the model version between Staging and Production stages to switch traffic.
AnswerC

Aliases are mutable pointers to a model version, and a served entity can reference an alias instead of a fixed version. Moving the alias from the old version to the new one shifts traffic without editing the endpoint, and moving it back reverts instantly. This matches the requirement to promote and roll back without configuration changes.

Why this answer

Serving through a Unity Catalog alias decouples the endpoint configuration from the specific model version. Promotion and rollback become operations on the alias, so traffic shifts to the new version or back to the previous one without editing or redeploying the endpoint, which is exactly the instant-revert capability the team needs.

Exam trap

The trap here is believing that stage transitions or version pinning can switch production traffic, when the served entity must reference an alias to allow promotion and rollback without a configuration change.

49
MCQmedium

When deploying a model using Model Serving, how does Databricks ensure that the environment remains consistent between the training workspace and the serving environment?

A.By requiring the user to provide a Dockerfile
B.By capturing the model's dependencies during log_model
C.By strictly enforcing the use of the latest stable libraries
D.By running the model on the same training cluster
AnswerB

MLflow automatically logs the environment dependencies, including Python packages and versions, when log_model is called. The serving infrastructure reads this metadata to rebuild the environment in the container. This ensures that the code runs in an environment identical to the one used during training and testing phases.

Why this answer

Databricks uses Conda or virtual environment dependencies captured during the MLflow logging process. When the model is logged, the environment details, including library versions, are saved. During deployment, the serving infrastructure creates a container that replicates this environment.

This practice is crucial for avoiding 'dependency hell', where models fail in production due to subtle library version mismatches compared to the original training environment.

Exam trap

Candidates frequently think Databricks automatically installs local cluster libraries into serving endpoints, forgetting that dependencies must be explicitly captured during the log_model process.

50
MCQhard

A machine learning engineer has an existing Databricks Model Serving endpoint named churn-endpoint serving version 3 of a model. The team has validated version 5 and wants to direct live traffic to it while keeping the deployment reversible if quality degrades. What is the most appropriate action?

A.Delete the endpoint and create a new one with a different name that serves model version 5
B.Update the served entity in the endpoint configuration to reference model version 5, allowing rollback by re-pointing to version 3
C.Transition model version 5 to Production in the Model Registry and assume the endpoint automatically follows the stage change
D.Keep the endpoint on version 3 and write a scheduled job that calls version 5 and overwrites predictions in the serving table
AnswerB

Updating the served entity to the new model version URI performs a managed rolling deployment while preserving the ability to revert by re-pointing to the earlier version. This keeps the change reversible and uses the endpoint's native configuration, matching both the traffic redirection and rollback requirements without extra infrastructure.

Why this answer

The endpoint serves whichever model version URI its configuration references, so updating the served entity to version 5 redirects live traffic through a managed rollout and keeps rollback simple by re-pointing to version 3. Deleting and recreating, batch overwrites, and registry stage changes do not achieve reversible live traffic redirection.

Exam trap

The trap here is believing that a Model Registry stage transition automatically changes which model version a serving endpoint uses.

51
Multi-Selectmedium

A machine learning engineer is preparing to deploy a scikit-learn model as a Databricks Model Serving endpoint. The model expects a pandas DataFrame with specific column names and types. Which two actions should the engineer take to ensure the endpoint correctly validates and processes inference requests? (Choose two.)

Select 2 answers
A.Convert the model to the `mlflow.pyfunc` flavor using a custom wrapper that hardcodes the expected column order.
B.Set the endpoint's `workload_size` to Large so the endpoint can coerce incoming data types automatically.
C.Provide a representative `input_example` when logging the model so MLflow can infer and store the input schema.
D.Log the model with an explicit `ModelSignature` that captures the training DataFrame's column names and data types.
E.Enable inference tables on the endpoint so that incoming requests are automatically reformatted to match the model's schema.
AnswersC, D

Passing an `input_example` gives MLflow a concrete sample to infer the input schema and store it alongside the model. Even when an explicit signature is used, an input example helps document expected payloads and can be used to validate the schema. Together with the signature, it ensures the endpoint knows the correct column names and types. This is a recommended practice for reliable deployments.

Why this answer

Logging the model with an explicit `ModelSignature` and a representative `input_example` ensures the Serving endpoint knows the exact column names and types to expect. The signature drives request validation, while the input example documents and helps infer the schema. Together they prevent malformed requests from reaching the model and keep training and serving schemas aligned.

Exam trap

The trap here is assuming that endpoint sizing or inference tables can fix schema mismatches, when schema enforcement comes from the logged MLflow signature.

52
MCQhard

A team operates a Databricks Model Serving endpoint with min_instances set to 0 and max_instances set to 4. During a nightly batch job, the endpoint receives a burst of requests and some clients observe elevated latency. The team wants to keep costs low during idle periods while reducing cold-start latency during bursts. Which configuration change best achieves this?

A.Set max_instances to 8 so the endpoint can scale further during the burst.
B.Set min_instances to 4 so the endpoint always matches the maximum capacity.
C.Set min_instances to 1 and leave max_instances at 4 so at least one instance is always warm.
D.Enable scale-to-zero by setting both min_instances and max_instances to 0 and rely on queueing.
AnswerC

Keeping min_instances at 1 preserves a warm replica that can answer requests immediately, eliminating scale-from-zero cold starts while max_instances still caps cost during spikes. This balances the cost-saving goal during idle periods with the latency goal during bursts, because only one instance runs continuously rather than four.

Why this answer

Model Serving endpoints scale between min_instances and max_instances. With min_instances at 0, the endpoint scales to zero when idle and must cold-start a replica when traffic resumes, causing the observed latency. Setting min_instances to 1 keeps a single warm replica available for immediate responses while still allowing scale-out up to max_instances during bursts, achieving both cost and latency goals.

Exam trap

The trap here is confusing the roles of min_instances and max_instances, assuming a higher maximum alone removes cold-start latency.

53
MCQhard

A team's Model Serving endpoint occasionally returns errors when the upstream feature store is slow. They want the endpoint to retry transient failures and reduce cold-start latency for bursty traffic. Which combination of endpoint settings best addresses both concerns?

A.Enable inference tables and route all traffic through a single replica to simplify retry logic.
B.Increase the endpoint's workload size to Large and disable request timeouts on the client.
C.Configure a non-zero min_instances for warm capacity and implement client-side retry with backoff for transient upstream errors.
D.Set min_instances to zero and enable scale-to-zero to lower cost during idle periods.
AnswerC

Keeping at least one instance warm avoids provisioning delay during bursts, directly reducing cold-start latency. Retrying transient failures with exponential backoff in the calling application or model code handles the intermittent feature store slowness. Together these settings address both the latency and reliability concerns described, making this the balanced operational choice.

Why this answer

Warm capacity via a non-zero minimum instance count prevents the provisioning delay that causes cold-start latency during bursts. Because the endpoint itself does not automatically retry upstream feature store calls, retry with backoff belongs in the client or model logic. Combining warm replicas with retry logic addresses latency and transient failures without sacrificing scalability or observability.

Exam trap

The trap here is assuming the serving endpoint automatically retries upstream dependency failures, when that resilience must be implemented by the caller or model code.

54
MCQmedium

A machine learning engineer is deploying an MLflow model to a Databricks Model Serving endpoint. The model was trained on a Spark DataFrame and logged with MLflow using the default signature detection. During testing, the endpoint returns predictions, but the engineer notices the input schema shown in the Serving UI does not match the actual DataFrame column types used during training. Which MLflow logging step should the engineer verify first to resolve this schema mismatch?

A.The model's conda environment file lists incompatible library versions, so the Serving endpoint silently coerces input types to match the environment.
B.The model was logged with `mlflow.pyfunc.log_model` without explicitly passing the `signature` argument, so MLflow inferred it from a sample that may not represent the full schema.
C.The Serving endpoint was configured with `min_instances` greater than zero, which forces the endpoint to rewrite the schema for autoscaling compatibility.
D.The model was registered in the Model Registry under a different name than the one used in the Serving endpoint configuration, causing the endpoint to load a stale schema.
AnswerB

When a model is logged without an explicit signature, MLflow attempts to infer it from the input example or from the model's internal schema, which can be incomplete or incorrect. The Serving endpoint relies on that logged signature to validate and parse incoming requests, so the engineer should check whether the signature was explicitly provided and correct. Supplying a proper `ModelSignature` during logging ensures the endpoint enforces the expected column names and types.

Why this answer

The logged MLflow model signature is the authoritative contract for the Serving endpoint's expected input columns and types. When a signature is missing or inferred from an unrepresentative sample, the endpoint may expose an incorrect schema. Explicitly passing a `ModelSignature` during `log_model` ensures column names and data types align with training data, resolving the mismatch before redeployment.

Exam trap

The trap here is assuming that any logged MLflow model automatically carries a complete and correct signature, when in fact inference can be incomplete.

55
Multi-Selecthard

A team owns a Databricks Model Serving endpoint that receives sporadic bursts of traffic. They want to reduce cold-start latency during bursts while keeping cost predictable, and they also need to capture the request payloads and predictions for later monitoring. Which two configuration choices should they make? (Choose two.)

Select 2 answers
A.Set scale_to_zero_enabled to true on the served entity to lower idle cost.
B.Increase the endpoint's workload size from Small to Large.
C.Set min_instances to a value greater than zero so a warm replica is always available.
D.Attach an MLflow experiment to the endpoint to record prediction requests.
E.Enable inference tables on the endpoint so requests and responses are logged to a Delta table.
AnswersC, E

Setting min_instances above zero keeps at least that many replicas provisioned and ready, so requests arriving after an idle period do not wait for a new container to load the model. This directly addresses cold-start latency during bursts, at the cost of paying for the always-on capacity. It does not by itself record request payloads or predictions.

Why this answer

Cold-start latency during sporadic bursts is best mitigated by keeping at least one replica warm with min_instances greater than zero, avoiding scale-to-zero teardown. Capturing request payloads and predictions for monitoring is done with inference tables, which write each request and response to a Delta table. Together these meet both the latency and the observability goals without conflicting.

Exam trap

The trap here is treating scale_to_zero_enabled, which saves cost by idling replicas, as a latency optimization, when it is actually the primary cause of cold starts on a serving endpoint.

56
MCQhard

Refer to the exhibit. A user attempts to update a model stage in the Model Registry and receives this error. What is the most appropriate action to resolve this?

A.Re-register the model under a new name
B.Request the workspace admin to update the model ACLs
C.Upgrade the model to a higher version
D.Switch to a different Databricks cluster
AnswerB

The correct administrative path is to update the Access Control List (ACL) for the specific model object. A workspace administrator or the current owner has the authority to grant the required 'CAN_MANAGE' permission to the user, allowing them to perform the requested stage transition securely and officially.

Why this answer

The error clearly indicates a lack of permission to modify the model's state. In Databricks, access control lists (ACLs) are strictly enforced on objects in the registry. To perform administrative actions like changing a model's stage, the user must be assigned the appropriate permission level (CAN_MANAGE) by an owner or admin.

This enforces the principle of least privilege, preventing unauthorized users from promoting unvetted models to production.

Exam trap

Candidates often try to modify model code or rewrite endpoint configurations to fix permission errors, failing to recognize that access control lists (ACLs) restrict registry actions.

57
MCQhard

A team has several model versions registered in Unity Catalog. They want to serve a specific version through a Databricks Model Serving endpoint and later promote a newer version without changing the endpoint URL used by applications. Which approach should they use?

A.Create the endpoint pointing to a fixed model version, then delete and recreate the endpoint with the new version when promoting.
B.Create the endpoint pointing to a Unity Catalog alias such as 'champion', and update the alias to reference the new model version when promoting.
C.Create the endpoint pointing to the model name without a version, and rely on the platform to always serve the most recently created version.
D.Create one endpoint per model version and place a load balancer in front that routes to the newest endpoint.
AnswerB

Serving endpoints can target a model alias rather than a fixed version. Pointing the endpoint at an alias like 'champion' lets the team repoint the alias to a new version, and the endpoint picks up the change without altering the URL. This decouples consumers from version numbers and supports controlled promotion.

Why this answer

Databricks Model Serving endpoints can reference a Unity Catalog model alias instead of a fixed version. Applications call a stable endpoint URL, and the team promotes new versions by moving the alias, such as from 'champion' to a newer version. This preserves the URL, avoids endpoint recreation, and keeps promotion governed through alias updates rather than client-side changes.

Exam trap

The trap here is assuming the endpoint must be rebuilt for each new version, when a Unity Catalog alias lets the same endpoint follow promotions.

58
Multi-Selectmedium

A team is preparing to deploy an MLflow model to a Databricks Model Serving endpoint and wants to diagnose why requests are failing before contacting support. Which two actions allow them to inspect the endpoint's behavior and errors? (Choose two.)

Select 2 answers
A.Re-register the model with a new signature and redeploy the endpoint to force the error to surface in the UI.
B.Query the endpoint's event logs and metrics in the Databricks workspace to review build and update events.
C.SSH into the serving container and read the application log files directly from the filesystem.
D.Enable inference tables on the endpoint to capture request and response payloads in a Delta table.
E.Attach a notebook to the serving cluster and run the model's predict method interactively to reproduce the error.
AnswersB, D

Endpoint event logs and metrics expose provisioning, build, and update events along with latency and error rates. These reveal whether the failure is a deployment or configuration problem, such as a failed model build or a scaled-to-zero event, complementing request-level payload inspection from inference tables.

Why this answer

Inference tables capture the actual request and response payloads, exposing malformed inputs and model-side errors, while endpoint event logs and metrics show deployment, scaling, and update events. Together they cover both request-level and infrastructure-level diagnosis. Attaching clusters, SSH access, and redeploying do not provide supported visibility into a managed serving endpoint's runtime.

Exam trap

The trap here is assuming serving endpoints can be inspected like interactive clusters, when diagnostics rely on inference tables and endpoint logs instead.

59
MCQmedium

Which of the following is a primary benefit of using a model serving endpoint versus a batch inference job?

A.Lower cost for infrequent predictions
B.Real-time, low-latency inference
C.Ability to process billions of rows at once
D.Simplified tracking of data lineage
AnswerB

Serving endpoints are designed for low-latency, synchronous request-response cycles. They expose the model via an API, allowing applications to request a prediction and receive it in milliseconds. This is the defining requirement for real-time applications where end-user interaction is waiting for the model result.

Why this answer

Model serving endpoints provide low-latency, real-time responses to individual requests via REST APIs. This is essential for applications like fraud detection or recommendation engines, where an immediate prediction is required for a user interaction. Batch inference jobs, conversely, are designed for high-throughput, asynchronous processing of large datasets at scheduled intervals.

Understanding these use cases is vital for selecting the correct deployment architecture to meet performance and latency requirements.

Exam trap

Candidates conflate batch processing with real-time serving. They assume serving endpoints are for throughput, missing that the primary advantage is low-latency, immediate response for individual requests.

60
MCQeasy

A data scientist registers a model in the Databricks Model Registry and wants to record the model's intended input and output schema so that a serving endpoint can validate incoming requests. Which action accomplishes this when logging the model?

A.Add the schema to the endpoint's tags in the serving configuration
B.Store the schema as a separate Delta table and reference it in the model's README
C.Log the model with an inferred or explicit signature using MLflow
D.Set the model version's stage to Production in the Model Registry
AnswerC

Logging the model with a signature records the expected input and output schema in the MLflow model artifact. Model Serving reads that signature to validate request payloads, so this is the correct way to enable schema validation and prevent malformed inputs from reaching the model.

Why this answer

A model signature is part of the logged MLflow model artifact and is what the serving runtime uses to validate request payloads against expected input and output types. Registry stages, endpoint tags, and external documentation are metadata that the runtime does not interpret as a schema, so only logging the model with a signature achieves request validation.

Exam trap

The trap here is assuming registry stages or endpoint tags carry schema information, when only the logged MLflow model signature drives request validation.

61
MCQeasy

A data scientist has trained a scikit-learn model and wants to expose it for real-time inference through Databricks Model Serving. The model is currently logged as an MLflow run artifact but has not been registered anywhere. What must the data scientist do before creating the serving endpoint?

A.Convert the model to an ONNX file for cross-platform compatibility.
B.Register the model in Unity Catalog and note the model name and version.
C.Export the model to a pickle file and upload it to DBFS.
D.Create a Databricks job that runs the model on a schedule.
AnswerB

Databricks Model Serving creates endpoints from registered model versions, referencing a three-level Unity Catalog name such as catalog.schema.model and a version number. Registering the logged model makes the artifacts, signature, and environment metadata available for the serving layer to build and deploy. Without a registered version, there is no model entity the endpoint can point to, so registration is the required prerequisite.

Why this answer

Databricks Model Serving deploys registered model versions, referencing a Unity Catalog model name and version. Logging a model during training stores artifacts but does not make the model deployable until it is registered. Once registered, the endpoint can be created against that version, and the serving layer uses the stored signature and environment to build the container.

Raw file uploads, format conversions, and scheduled jobs do not satisfy this prerequisite.

Exam trap

The trap here is treating a logged MLflow artifact as directly deployable, when Model Serving actually requires a registered model version to point the endpoint at.

62
MCQmedium

A data scientist registers a scikit-learn model in the Unity Catalog model registry with the name `prod.ml_team.fraud_detector`. They now want to serve it with Databricks Model Serving. Which value should be supplied as the model identifier when creating the serving endpoint?

A.prod.ml_team.fraud_detector
B.models:/prod.ml_team.fraud_detector/Production
C.dbfs:/databricks/mlflow/prod/ml_team/fraud_detector
D.runs:/<run_id>/model
AnswerA

Unity Catalog models are referenced by their fully qualified three-level name, catalog.schema.model_name, so prod.ml_team.fraud_detector is exactly the identifier Model Serving expects. Because the model lives in Unity Catalog rather than the workspace registry, the endpoint resolves the model version directly from that namespace, and permissions on the catalog, schema, and model govern who can serve it.

Why this answer

Because the model was registered in Unity Catalog, the endpoint must be created from the three-level namespace catalog.schema.model_name, optionally pinned to a version. Stage-based URIs and raw storage or run paths do not identify a Unity Catalog registered model, so they cannot be used as the served entity. The fully qualified name is the correct identifier.

Exam trap

The trap here is assuming that MLflow stage-based URIs like models:/name/Production still work for models registered in Unity Catalog, when Unity Catalog models are addressed by their three-level name and use aliases instead of stages.

63
MCQmedium

A data scientist trains a model with a feature engineering pipeline and wants batch scoring to happen nightly on a Delta table using Databricks, producing predictions that downstream dashboards read. The scoring job must scale with data volume and be re-runnable if it fails. Which approach best fits these requirements?

A.Deploy the model to a real-time Model Serving endpoint and call it row by row from a notebook loop.
B.Export the model to a pickle file and have the dashboard application load and score records on demand.
C.Use MLflow's predict API to load the model in a Spark job and score the Delta table with a distributed DataFrame.
D.Convert the model to a SQL UDF and run a single-threaded query against the feature table.
AnswerC

Loading the model with the MLflow predict API inside a Spark job lets the scoring run distributed across the cluster, so throughput scales with the data volume and executor count. The job can be scheduled with a Databricks job and retried on failure, and results are written back to Delta for dashboards. This aligns with a nightly, re-runnable bulk scoring pattern.

Why this answer

Nightly bulk scoring on a Delta table is a distributed batch workload, so the model should be loaded with the MLflow predict API inside a Spark job that reads the feature table and writes predictions back to Delta. That approach scales horizontally with data volume and integrates cleanly with Databricks job scheduling and retries, unlike endpoint-per-row calls, dashboard-embedded pickles or single-threaded SQL.

Exam trap

The trap here is assuming a real-time serving endpoint is the default way to use a model, when scheduled distributed batch scoring is the appropriate pattern for bulk table data.

64
MCQmedium

A data scientist registers a scikit-learn model in Unity Catalog as `prod.ml.churn_model`. They then create a Databricks Model Serving endpoint via the REST API using `served_entities` with `entity_name` set to `prod.ml.churn_model` and `entity_version` set to `"3"`. The endpoint creation fails. What is the most likely cause?

A.Model Serving only supports models registered in the workspace Model Registry, not models registered in Unity Catalog.
B.The model version must be referenced as a numeric integer, but the field only accepts a string, so the deployment cannot resolve the artifact path.
C.The model must first be exported to DBFS as an MLflow artifact tarball before it can be referenced by a serving endpoint.
D.The serving endpoint lacks USE CATALOG, USE SCHEMA, and EXECUTE privileges on the Unity Catalog model, so the service principal cannot read the model artifacts.
AnswerD

Databricks Model Serving authenticates as a platform-managed identity and requires explicit Unity Catalog grants on the model's parent catalog, schema, and the model itself. Without USE CATALOG, USE SCHEMA, and EXECUTE on `prod.ml.churn_model`, the endpoint cannot fetch artifacts and creation fails. Granting these to the serving identity or the requesting user resolves the deployment.

Why this answer

Serving a Unity Catalog model requires the serving identity to hold USE CATALOG, USE SCHEMA, and EXECUTE privileges on the catalog, schema, and model. Without these grants, the endpoint cannot retrieve the model version, and creation fails. Referencing the model by three-level name and version string is otherwise correct, so permissions are the most likely root cause.

Exam trap

The trap here is assuming that a user who can view a Unity Catalog model in the UI automatically has the EXECUTE privilege required for Model Serving to load the artifacts.

65
MCQmedium

Why should you use an inference table in Databricks Model Serving?

A.To increase the speed of model inference
B.To automatically retrain the model on new data
C.To provide visibility into production inputs and predictions
D.To secure the endpoint against malicious attacks
AnswerC

Inference tables provide a permanent, queryable record of every inference call, including the input features and the resulting prediction. This is the foundation for monitoring, drift detection, and debugging. By logging this data to Delta tables, you gain full insight into how the model is being used in production.

Why this answer

Inference tables allow you to capture every request and response processed by a model endpoint. This is vital for monitoring model performance over time, debugging issues by examining specific input-output pairs, and detecting data drift. Without this captured data, you essentially have no visibility into how the model is performing in the real world, which makes it impossible to maintain and improve the system effectively.

Exam trap

Candidates think inference tables are for performance tuning or caching. They fail to realize the primary purpose is observability: tracking inputs and outputs for monitoring and drift detection.

66
MCQmedium

A machine learning engineer has registered a scikit-learn model in Unity Catalog as `prod.ml.risk_model` with version 3. They want to serve it through a Databricks Model Serving endpoint that always uses this exact model version, even after new versions are registered. Which endpoint configuration should they use?

A.Set the served entity to `prod.ml.risk_model` and set the version to `latest` so that the endpoint always uses the most recent registered version.
B.Set the served entity to `prod.ml.risk_model` with version 3 and enable scale-to-zero.
C.Set the served entity to `prod.ml.risk_model` with version 3 and do not configure aliases or the `latest` version.
D.Set the served entity to `prod.ml.risk_model` and rely on the `@champion` alias to always resolve to version 3.
AnswerC

Databricks Model Serving lets you serve a specific registered model version by supplying the model name and version number in the served entity configuration. By specifying version 3 and avoiding aliases or the `latest` marker, the endpoint stays pinned to that immutable version even as newer versions are registered, which directly satisfies the requirement of always serving the same model.

Why this answer

Pinning a Model Serving endpoint to a specific registered model version is done by supplying the model name and version number in the served entity configuration. This produces an immutable reference that does not change when new versions are registered. Relying on aliases or `latest` introduces mutability, which conflicts with the requirement to always serve version 3.

Exam trap

The trap here is assuming that using a model alias such as `@champion` gives the same immutability as a numeric version, when aliases are actually mutable pointers.

67
MCQmedium

A data scientist has registered a scikit-learn model in Unity Catalog as `ml_team.prod.churn_model` and wants it served by a Databricks Model Serving endpoint that performs online inference for a web app. The endpoint must automatically pick up new model versions as they are promoted, without the data scientist editing the endpoint each time. Which approach should the data scientist use to configure the served entity?

A.Serve the model by name with the version set to the alias `champion`, so the endpoint resolves the alias to whatever version it currently points to.
B.Create a Databricks job that runs every five minutes, downloads the latest model version, and calls the Model Serving REST API to overwrite the endpoint configuration with the new numeric version.
C.Serve the model by name with the version pinned to the numeric version that is currently in production, then re-create the endpoint whenever a new version is registered.
D.Export the MLflow model to a local directory, upload the artifacts to DBFS, and configure the endpoint to load the model from that DBFS path.
AnswerA

Serving a model by name plus an alias lets Model Serving resolve the alias at request time, so promoting a new version to `champion` updates the served model without endpoint edits. This matches the requirement of automatic pickup on promotion and works with Unity Catalog-registered models, provided the endpoint's identity has EXECUTE on the model and its versions.

Why this answer

Serving by model name with an alias makes the endpoint follow promotions automatically: when a new version is assigned the `champion` alias, subsequent requests resolve to that version. Pinning a numeric version, copying artifacts to DBFS, or polling with a job all require manual or scheduled intervention and cannot deliver the immediate, configuration-free update the scenario demands.

Exam trap

The trap here is assuming that any serving configuration referencing a registered model will automatically follow the newest version, when in fact only an alias or version reference determines what is actually served.

68
MCQhard

A team has registered a model in Unity Catalog and wants to serve it with Databricks Model Serving. Their security policy requires that the endpoint access the model without embedding long-lived credentials, and they want the endpoint to use a dedicated service principal for accessing downstream feature tables. Which configuration should they apply to meet these requirements?

A.Mount the feature tables as a DBFS path and configure the model to read them using a cluster-scoped credential passthrough token.
B.Enable token-based authentication on the endpoint and pass the token in each inference request header from the calling application.
C.Configure the Serving endpoint with a personal access token stored in the endpoint's environment variables.
D.Grant the endpoint's service principal access to the model and feature tables via Unity Catalog privileges, and configure the endpoint to use that service principal.
AnswerD

Databricks Model Serving integrates with Unity Catalog so that an endpoint can run under a service principal identity. By granting that principal `EXECUTE` on the model and `SELECT` on feature tables, the endpoint accesses resources without embedding credentials. This satisfies both the no-long-lived-credentials policy and the need for a dedicated identity. The service principal's Unity Catalog privileges govern all access at runtime.

Why this answer

Unity Catalog allows Databricks Model Serving endpoints to run under a service principal identity, with access governed by Unity Catalog privileges on models and tables. Granting the service principal `EXECUTE` on the registered model and `SELECT` on feature tables enables secure, credential-free access. This design avoids embedding long-lived tokens and provides auditable, least-privilege access for production inference workloads.

Exam trap

The trap here is confusing inbound authentication to the endpoint with the outbound identity the endpoint uses to reach models and data.

69
MCQeasy

A team has just registered a new version of a fraud detection model in the workspace Model Registry. Before routing production traffic to it, they want to send a small percentage of live requests to the new version and compare its predictions against the current production model. Which Databricks Model Serving feature should they use?

A.Inference tables
B.Provisioned throughput
C.Traffic splitting between served entities
D.Model aliases
AnswerC

Databricks Model Serving supports configuring multiple served entities on one endpoint and assigning each a percentage of traffic. By giving the new model version a small percentage and the current production model the remainder, the team can perform a canary-style rollout and compare live behavior before a full switch. This is exactly the built-in mechanism for gradual, controlled deployment of a new model version.

Why this answer

Traffic splitting on a Databricks Model Serving endpoint allows multiple served entities, each pointing to a different model version, to share incoming requests by percentage. This makes it possible to route a small fraction of live traffic to a new version and compare its behavior against the existing production model before promoting it fully. Other listed features address logging, capacity, or naming, not request routing.

Exam trap

The trap here is conflating logging with routing: inference tables record what happened, but only traffic splitting actually sends a percentage of requests to the new model version.

70
MCQmedium

A team wants to monitor a production Model Serving endpoint for data drift and to capture the exact request payloads and responses for later auditing. They want this logging to happen automatically without adding code to the client application. What should they configure on the endpoint?

A.Enable inference tables on the endpoint so requests and responses are automatically logged to a Delta table for monitoring and auditing.
B.Attach an MLflow experiment to the endpoint so each prediction is recorded as a run with parameters and metrics.
C.Add client-side logging middleware in the application that writes each request and response to a Delta table before and after calling the endpoint.
D.Configure the endpoint to write its container stdout and stderr to a log delivery location in cloud storage.
AnswerA

Inference tables capture the request and response payloads of a serving endpoint into a Delta table automatically, with no client changes. Teams can then join those logs against training data to compute drift metrics and retain an audit trail. Enabling the feature is a configuration step on the endpoint, matching the requirement for automatic, code-free logging.

Why this answer

Inference tables are the native Model Serving feature that logs request and response payloads to a Delta table automatically, enabling drift monitoring and auditing without touching client code. Client middleware, container logs, and MLflow experiments either require application changes, lack payload structure, or are not built for high-volume inference logging.

Exam trap

The trap here is assuming that operational logs or MLflow tracking can substitute for inference tables, when only inference tables capture structured request and response payloads at the endpoint.

71
MCQeasy

A machine learning engineer has a registered model and wants to expose it as a REST endpoint that their application can call for real-time predictions. They need the endpoint to be created and managed natively within Databricks, with the ability to enable scale-to-zero during idle periods. Which Databricks capability should they use?

A.MLflow Model Registry webhooks that notify an external service whenever a model version is registered, letting that service host the model.
B.A Databricks job that runs a Python script loading the model and starting a Flask web server on a driver node, exposed via a public URL.
C.Databricks Model Serving, creating an endpoint that serves the registered model and configuring the served entity with scale-to-zero enabled.
D.A SQL warehouse with a user-defined function that loads the model artifact and returns predictions for incoming queries.
AnswerC

Databricks Model Serving provides managed REST endpoints for registered models and supports scale-to-zero, which shuts down compute when there is no traffic and cold-starts on the next request. This satisfies both the native management and idle-cost requirements. The endpoint exposes a standard REST interface and integrates with Unity Catalog for permissions and lineage.

Why this answer

Databricks Model Serving is the native mechanism for deploying registered models as REST endpoints with managed lifecycle, authentication, and optional scale-to-zero. A Flask process on a job cluster, registry webhooks, or a SQL warehouse cannot provide a stable, secure, autoscaling prediction API, so they fail the scenario's requirements for native management and idle shutdown.

Exam trap

The trap here is assuming that any Databricks compute that can run Python can also host a production-grade model endpoint, when only Model Serving provides managed REST inference.

72
MCQmedium

A machine learning engineer is deploying a model to a Databricks Model Serving endpoint and needs to send feature values that are computed by a separate upstream pipeline. The engineer wants the endpoint to accept a JSON payload describing a single record with named fields. Which approach correctly describes how the client should format the request?

A.Send a JSON body containing a base64-encoded serialized pandas DataFrame along with a pickle content type.
B.Send the record as form-encoded key-value pairs in the request body with a content type of application/x-www-form-urlencoded.
C.Send a JSON body with a dataframe_records or dataframe_split field containing the named columns matching the model signature.
D.Send a JSON body with an inputs field containing a nested array of numeric values in the exact column order of training.
AnswerC

Databricks Model Serving accepts pandas DataFrame inputs in JSON using the dataframe_records or dataframe_split format, where each record maps column names to values. This aligns with the signature logged for a scikit-learn or similar model and lets the endpoint validate and score the record without custom parsing.

Why this answer

For models logged with a pandas DataFrame signature, Databricks Model Serving accepts JSON using dataframe_records or dataframe_split, where each record is an object of column names to values. This preserves the named-column contract and allows signature validation. Positional array formats and non-JSON encodings do not match the DataFrame interface and would either fail validation or misalign features.

Exam trap

The trap here is assuming any JSON structure works, when the endpoint expects the specific dataframe_records or dataframe_split format for DataFrame-signature models.

Ready to test yourself?

Try a timed practice session using only Model Deployment questions.