Courseiva

CCNA Model Deployment Questions

61 questions · Model Deployment · All types, answers revealed

1
MCQhard

Refer to the exhibit. The JSON configuration represents an existing Databricks Model Serving endpoint. You need to update this endpoint to support a traffic split between version 5 and version 6 for A/B testing. Which update strategy is correct?

A.Update the config to replace 'model_version': '5' with 'model_version': '6' in the current served_models list.
B.Create two separate endpoints, one for version 5 and one for version 6, then split traffic using a load balancer.
C.Modify the served_models array to include both versions with specific traffic weights assigned to each.
D.Delete the current endpoint and recreate it with a new configuration that includes only version 6.
AnswerC

Adding both versions to the `served_models` list enables traffic routing control. By specifying the `traffic` field for each, you can define the percentage of requests allocated to each version. This configuration is the standard method for A/B testing in Databricks, providing safe, granular control over model deployment transitions.

Why this answer

To perform A/B testing, you must define multiple served models within the `served_models` array in the endpoint configuration. Each entry requires a `traffic` weight that sums to 100%. This allows Databricks to route incoming requests according to the specified percentages.

This method is crucial for safely rolling out new models, as it allows you to observe performance metrics on a subset of real-world traffic before a full production cutover.

Exam trap

Candidates try to create multiple endpoints for A/B testing rather than modifying the `served_models` array within a single endpoint configuration to split traffic weights.

2
Multi-Selectmedium

A team is deploying a model to Databricks Model Serving and wants to implement a canary release strategy to gradually shift traffic from the current model version to a new version. Which TWO configurations are required to achieve this? (Choose two.)

Select 2 answers
A.Create a serving endpoint with multiple served entities, each referencing a different model version.
B.Deploy two separate endpoints and use an external load balancer to distribute traffic between them.
C.Enable automatic canary deployment by setting a flag in the model's MLflow metadata.
D.Configure the endpoint to use a single served entity and rely on the model's internal logic to route requests based on a random seed.
E.Use the Databricks REST API to update the traffic configuration for the endpoint, specifying the percentage of traffic for each served entity.
AnswersA, E

Databricks Model Serving supports multiple served entities within a single endpoint, each pointing to a different model version. This allows you to route traffic to different versions, which is essential for canary releases. You can then adjust the traffic split between the entities to gradually shift traffic.

Why this answer

Canary releases in Databricks Model Serving are achieved by configuring a single endpoint with multiple served entities, each pointing to a different model version. Traffic is then split between these entities using the endpoint's traffic configuration, which can be updated via the REST API. This allows gradual shifts and monitoring.

Exam trap

The trap here is thinking that canary deployments require multiple endpoints or automatic flags, when actually they are configured within a single endpoint using multiple served entities and explicit traffic weights.

3
MCQmedium

An ML engineer is deploying a model to Databricks Model Serving that requires a custom Python package not available in the default environment. The model was logged with MLflow and includes the package in its conda environment. What must the engineer ensure for the endpoint to successfully load the model?

A.The package must be installed on the driver node of the Databricks cluster used for serving.
B.The package must be installed via an init script that runs when the endpoint starts.
C.The package must be uploaded to DBFS and referenced in the model's signature.
D.The package must be included in the model's conda environment and the endpoint must be configured to use that environment.
AnswerD

When logging an MLflow model, the conda environment specifies the dependencies. Databricks Model Serving reads this environment file and installs the listed packages into the serving container. The engineer must ensure the custom package is correctly listed in the conda environment and that the endpoint is created without overriding the environment. This allows the serving environment to replicate the training environment, making the custom package available.

Why this answer

Databricks Model Serving builds the serving environment from the MLflow model's conda environment. To use a custom package, it must be listed in that environment. The endpoint will then install it automatically.

Cluster-based methods, DBFS uploads, or init scripts do not affect the serverless serving container, so they are not correct.

Exam trap

The trap here is assuming that serving endpoints run on Databricks clusters where you can install packages or run init scripts, when they actually use isolated serverless containers built from the model's conda environment.

4
MCQeasy

A team has deployed a model to a Databricks Model Serving endpoint. They want to monitor the endpoint's performance and detect data drift over time. Which Databricks feature should they use to automatically track inference data and compute drift metrics?

A.Inference tables
B.MLflow tracking server
C.Model serving endpoint logs
D.Databricks SQL dashboards
AnswerA

Inference tables automatically capture the request and response payloads for a serving endpoint and store them in a Delta table. This allows you to monitor model performance, detect data drift, and analyze predictions over time using Databricks SQL or notebooks.

Why this answer

Inference tables are a Databricks feature that automatically logs the input and output of a model serving endpoint to a Delta table. This data can then be used to compute drift metrics, monitor model performance, and trigger alerts. It is the built-in solution for capturing inference data for monitoring purposes.

Exam trap

The trap here is confusing operational logs with inference data capture; endpoint logs do not contain the payloads needed for drift detection.

5
MCQeasy

A data science team has deployed a model to Databricks Model Serving and wants to ensure that the endpoint can handle sudden spikes in traffic without manual intervention. Which feature should they configure?

A.A larger workload type with more memory and CPU.
B.Deploy multiple endpoints and use a round-robin DNS.
C.Enable scale-to-zero to reduce cold starts during spikes.
D.Autoscaling with a defined minimum and maximum replica count.
AnswerD

Autoscaling automatically adjusts the number of replicas based on incoming traffic, ensuring the endpoint can handle spikes without manual scaling. By setting minimum and maximum replicas, you bound the scaling to control cost and capacity. This is the intended feature for handling variable load in Databricks Model Serving.

Why this answer

Autoscaling is the Databricks Model Serving feature that dynamically adjusts the number of replicas based on traffic, allowing the endpoint to handle spikes without manual intervention. Configuring minimum and maximum replicas ensures that scaling is bounded and cost-effective. Other options either provide static capacity or are not designed for dynamic load handling.

Exam trap

The trap here is confusing scale-to-zero with autoscaling; scale-to-zero reduces cost during idle periods but can cause cold starts, while autoscaling adds replicas to handle increased load.

6
MCQmedium

A machine learning engineer needs to deploy a custom scikit-learn model to a Databricks Model Serving endpoint with a strict response time SLA of under 50 milliseconds. The model includes an extensive text-cleaning pipeline that utilizes heavy regex matching. How should the engineer package the model to ensure maximum inference efficiency and meet the low-latency requirement?

A.Store the preprocessing logic in a Delta table and have the serving endpoint query the table asynchronously during the scoring request.
B.Develop a separate Azure or AWS Lambda function to handle the text cleaning before forwarding the request payload to the model endpoint.
C.Implement a custom MLflow PyFunc model where both the text preprocessing and the scikit-learn predictor are encapsulated inside the predict method.
D.Deploy the scikit-learn model natively without custom wrapper code and require client applications to execute the regex cleaning logic locally.
AnswerC

Encapsulating preprocessing within a custom MLflow PyFunc guarantees that raw input strings are cleaned and transformed consistently in the same execution context as the model. This eliminates extra network round trips, optimizes memory usage, and ensures predictable sub-50ms inference performance.

Why this answer

Integrating the preprocessing directly into the MLflow PyFunc wrapper ensures that the input transformation runs within the optimized inference container memory space, avoiding out-of-band network calls and reducing serialization overhead. This architectural pattern prevents latency bottlenecks often introduced by separate preprocessing microservices, satisfying strict production SLAs.

Exam trap

Candidates separate preprocessing logic into an external API or upstream service, creating network latency bottlenecks that violate strict response time SLAs.

7
MCQhard

An ML engineer is deploying a model that includes a custom Python class for preprocessing. During deployment to a Model Serving endpoint, the model fails to load with a 'ModuleNotFoundError'. What is the most likely cause of this error despite having the class in the training notebook?

A.The custom class was not saved as a separate .py file and included in the 'code_paths' parameter during logging.
B.The Model Serving endpoint does not support custom Python classes for security reasons.
C.The data scientist forgot to install the 'databricks-model-serving' library in the training cluster.
D.The custom class must be registered in the Unity Catalog as a separate 'Function' entity.
AnswerA

Notebook-defined classes exist only in the memory of the training session. To make them available to the Model Serving endpoint, the class definition must be in a Python file that is explicitly uploaded to the MLflow artifact store along with the model during the logging process.

Why this answer

When MLflow logs a model, it captures the environment but not necessarily the local code or classes defined in the notebook unless they are part of a package or provided as a code dependency. For custom classes to be available in the serving container, they must be included in the 'code_paths' argument of the log_model function.

Exam trap

Candidates assume that because a custom class is defined and works inside an interactive notebook, MLflow will automatically pickle or capture it without explicit file inclusion.

8
Multi-Selectmedium

A team is deploying a scikit-learn model to Databricks Model Serving and wants to minimize cold-start latency so that the first request after a period of inactivity is still fast. Which TWO actions help achieve this? (Choose two.)

Select 2 answers
A.Set the endpoint's scale-to-zero behavior so that instances remain warm and are not fully shut down during idle periods.
B.Reduce the size of the model artifact and its logged dependencies so that container startup and model loading complete faster.
C.Enable Inference Tables on the endpoint so that request payloads are cached and replayed on subsequent cold starts.
D.Increase the number of served model versions on the endpoint so that requests can be spread across more copies of the model.
E.Lower the endpoint's concurrency setting so that each replica handles fewer simultaneous requests and starts faster.
AnswersA, B

Scale-to-zero controls whether the endpoint's compute is torn down when there is no traffic. Disabling it, or configuring the endpoint to keep instances provisioned, means the model stays loaded and the first request after idle time does not pay the container start and model load cost, which is the dominant contributor to cold-start latency.

Why this answer

Cold-start latency comes from provisioning compute and loading the model environment. Keeping instances warm by avoiding full scale-down removes the provisioning cost, and shrinking the model artifact plus its dependencies shortens the load phase. Neither logging nor routing changes affect how quickly a replica becomes ready.

Exam trap

The trap here is confusing Inference Tables, which log payloads for audit, with a caching or warm-up mechanism that reduces cold-start latency.

9
MCQhard

A team is deploying a model to Databricks Model Serving that requires a specific version of a Python library that conflicts with the version pre-installed in the serving environment. They include the library version in the model's requirements.txt. However, upon deployment, the endpoint fails to start, and logs indicate a dependency conflict. What is the most likely cause of this failure?

A.The library version conflicts with a pre-installed library that is required by the serving infrastructure, causing a dependency resolution failure.
B.The model's requirements.txt is not being parsed correctly due to a syntax error.
C.The serving environment ignores requirements.txt and uses only the pre-installed libraries.
D.The specified library version is incompatible with the Python version used by the serving environment.
AnswerA

Databricks Model Serving has a set of pre-installed libraries that are critical for the serving runtime. If your requirements.txt specifies a version that conflicts with these, pip may fail to resolve dependencies, or the endpoint may crash at runtime. This is a common cause of deployment failures when custom dependencies clash with the base environment.

Why this answer

Model Serving environments come with pre-installed libraries that support the serving infrastructure. If a model's requirements.txt specifies a version that conflicts with these, the dependency resolver may fail, preventing the endpoint from starting. The solution is to align the requested version with the pre-installed one or use a custom container image to isolate dependencies.

Exam trap

The trap here is assuming that any library version can be installed, when in fact the serving environment has fixed dependencies that can cause conflicts.

10
Multi-Selectmedium

When using the Unity Catalog Model Registry, what are the primary advantages of using 'Aliases' over 'Versions' when calling a model from a production application? (Select TWO)

Select 2 answers
A.Aliases allow the application to always point to a stable name like '@prod' instead of a hardcoded version number.
B.Aliases improve performance by caching the model weights on the client side.
C.Aliases allow for easier rollbacks by simply reassigning the alias to a previous version.
D.Aliases are required to enable GPU acceleration for models in Unity Catalog.
E.Only models with an alias can be used in Spark structured streaming jobs.
AnswersA, C

By using a symbolic name, the engineering team can update the model version that the alias points to in the registry. The production application, which calls the alias, will automatically start receiving predictions from the new version without requiring a code redeployment or restart.

Why this answer

Aliases provide a layer of abstraction between the application code and the specific model version. This decoupling is a cornerstone of MLOps, as it allows for seamless updates and rollbacks without modifying the client-side code that consumes the model. It also improves readability by using descriptive names like '@champion'.

Exam trap

Test-takers often recommend hardcoding specific integer version numbers in production applications, leading to brittle codebases that require manual code deployments for every model update.

11
MCQmedium

An ML engineer has deployed a model to Databricks Model Serving and wants to update the endpoint to serve a new model version without changing the endpoint URL or causing downtime. Which approach is correct?

A.Update the existing endpoint's configuration to reference the new model version and apply the update.
B.Use MLflow's `transition_model_version_stage` to move the new version to 'Production', which automatically updates the serving endpoint.
C.Create a new endpoint with the new model version, then delete the old endpoint and update DNS to point to the new endpoint.
D.Modify the model's signature in Unity Catalog to match the new version, and the endpoint will detect the change and update automatically.
AnswerA

Databricks Model Serving allows you to update an existing endpoint's configuration, including the model version, via the REST API or UI. The endpoint URL remains the same, and Databricks performs a rolling update to minimize downtime. This is the standard way to deploy a new model version to an existing endpoint without disrupting service.

Why this answer

To update a serving endpoint to a new model version without downtime, you update the endpoint's configuration to reference the new version and apply the change. Databricks handles the rolling update, ensuring the endpoint URL remains unchanged. Other options would either cause downtime, require client changes, or rely on non-existent automatic updates.

Exam trap

The trap here is believing that model registry stage transitions or signature changes automatically propagate to serving endpoints.

12
MCQeasy

When deploying a model to a Databricks Model Serving endpoint, what is the purpose of the 'Small', 'Medium', and 'Large' workload size settings?

A.They define the maximum number of concurrent requests the endpoint can handle.
B.They determine the geographic region where the model will be hosted.
C.They specify the amount of CPU and memory allocated to each model instance.
D.They select the version of the MLflow library used for deployment.
AnswerC

Each size tier provides a specific amount of RAM and vCPU. A 'Small' instance might be sufficient for a simple linear regression, while a 'Large' instance would be necessary for complex ensembles or models with large memory footprints to ensure they don't run out of memory during execution.

Why this answer

The workload size setting in Databricks Model Serving defines the compute resources (CPU and Memory) allocated to each instance of the model. Choosing the right size is a trade-off between the complexity of the model's computation and the cost of the infrastructure. Larger models or those with heavy preprocessing requirements need more resources to maintain low latency.

Exam trap

Candidates often mistake these settings for 'number of instances' or 'scaling limits'. They are specifically for compute resource allocation (CPU/RAM) per instance to match model complexity.

13
MCQhard

An ML engineer is deploying a model to Databricks Model Serving that requires a custom Python package. The package is not available in the default environment and must be installed from a private PyPI repository. Which method ensures the package is available to the model at serving time?

A.Add the package to the cluster's init script and ensure the cluster is attached to the serving endpoint.
B.Include the package as a wheel file in the model's artifact directory and reference it in the model's conda environment.
C.Upload the package to DBFS and use a %pip install command in a notebook before deploying the model.
D.Specify the package's index URL and credentials in the model's conda environment file, and ensure the serving endpoint has network access to the private repository.
AnswerD

Databricks Model Serving allows you to specify a custom index URL and credentials in the conda environment file (e.g., via pip_requirements or conda_env). The serving environment must have network access to the private repository. This method securely installs the package from the private PyPI repository during environment build.

Why this answer

To use a private PyPI repository, the model's conda environment must include the index URL and credentials. Databricks Model Serving builds the environment based on this specification, provided it has network access to the repository. This ensures the custom package is installed and available during inference.

Exam trap

The trap here is assuming that notebook-level installations or cluster init scripts carry over to the serving environment, which is isolated and built from the model's environment specification.

14
MCQmedium

A data scientist is deploying a model to Databricks Model Serving. The model was trained using a scikit-learn pipeline that includes a custom transformer. The custom transformer is defined in a Python module that is not part of the model artifact. What should the data scientist do to ensure the model can be served successfully?

A.Include the custom transformer code in the model's conda environment as a pip dependency.
B.Log the model with the custom transformer code included in the model's artifacts using MLflow's custom logging.
C.Re-train the model without the custom transformer, as Model Serving does not support custom code.
D.Convert the custom transformer to a built-in scikit-learn transformer before logging.
AnswerB

MLflow allows you to log custom code along with the model by including it in the model's artifacts and specifying it in the model's Python function or by using the `code_path` parameter in `mlflow.pyfunc.log_model`. This ensures the custom transformer is available at inference time. Databricks Model Serving will then load the code from the model artifact, making it the correct and straightforward solution.

Why this answer

To deploy a model with custom code, you must include that code in the model's artifacts when logging with MLflow. Using the `code_path` parameter in `mlflow.pyfunc.log_model` allows you to specify additional code files that will be packaged with the model. Databricks Model Serving then loads this code, ensuring the custom transformer is available during inference.

This is the standard method for handling custom logic in served models.

Exam trap

The trap here is thinking that custom code must be installed as a separate package; instead, it can be bundled directly with the model artifact using MLflow's code_path.

15
MCQmedium

An ML engineer is updating a model serving endpoint to use a new model version. They want to gradually shift traffic from the old version to the new version to monitor performance before full rollout. Which feature of Databricks Model Serving should they use?

A.Configure traffic splitting on the existing endpoint by specifying a percentage of traffic to route to each served model version.
B.Create a new endpoint for the new model version and use a load balancer to distribute traffic.
C.Use the 'Champion' and 'Challenger' aliases in Unity Catalog to automatically route traffic based on model performance.
D.Enable A/B testing by deploying the new model to a separate endpoint and using Databricks SQL to compare metrics.
AnswerA

Databricks Model Serving supports traffic splitting, allowing you to route a percentage of requests to different model versions within the same endpoint. By specifying the traffic percentage for each served entity, you can gradually shift traffic from the old to the new version. This enables safe rollout and performance monitoring without disrupting the endpoint.

Why this answer

Databricks Model Serving allows configuring traffic splitting on a single endpoint by assigning a percentage of traffic to each served model version. This enables gradual rollout, allowing the team to monitor the new version's performance while still serving the old version. It is the native and recommended approach for safe model updates.

Exam trap

The trap here is thinking that aliases or separate endpoints automatically handle gradual traffic shifting, when in fact traffic splitting must be explicitly configured on the endpoint.

16
Multi-Selecthard

A team is deploying a model that requires custom Python libraries not available in the default Databricks Runtime. Which TWO methods can be used to ensure these dependencies are available in the Model Serving environment? (Select TWO)

Select 2 answers
A.Include a 'requirements.txt' file or a 'conda.yaml' when calling mlflow.log_model().
B.Manually install the libraries on the driver node of the cluster used for training.
C.Use the pip_requirements parameter in the mlflow.sklearn.log_model (or similar) function.
D.Add the libraries to the Spark configuration of the Model Serving endpoint.
E.Upload the wheel files to a DBFS location and reference them in the endpoint UI.
AnswersA, C

Providing a requirements file or Conda environment during logging tells MLflow exactly which packages and versions are needed. When the model is deployed to an endpoint, Databricks uses this information to build a container image with all the necessary dependencies pre-installed and ready for execution.

Why this answer

Model Serving environments are reconstructed based on the metadata captured when the model was logged. To include custom libraries, they must be explicitly defined during the logging process. This ensures that the production environment exactly matches the development environment, preventing 'missing module' errors during real-time inference requests.

Exam trap

Test-takers often think installing libraries directly in the serving endpoint settings or cluster configuration is sufficient, forgetting that dependencies must be captured during model logging.

17
Multi-Selectmedium

An ML engineer is transitioning a model from the Workspace Model Registry to the Unity Catalog (UC) Model Registry. Which TWO statements describe benefits or requirements of using Unity Catalog for model management? (Select TWO)

Select 2 answers
A.Unity Catalog supports model sharing across multiple Databricks workspaces.
B.Models must be stored in a legacy DBFS location to be registered in Unity Catalog.
C.Unity Catalog models use a three-level namespace: catalog, schema, and model name.
D.The legacy MLflow Stages (Staging, Production) are the primary way to manage UC models.
E.Only models logged with the SparkML flavor can be registered in the Unity Catalog.
AnswersA, C

By utilizing a centralized Metastore, Unity Catalog allows models to be registered once and accessed by authorized users across any workspace linked to that Metastore. This facilitates better collaboration and standardization across large organizations that use separate environments for development, testing, and production stages.

Why this answer

Unity Catalog centralizes governance for all data and AI assets, providing a unified interface for access control and lineage. Transitioning to UC enables fine-grained permissions and allows models to be shared across different workspaces within the same Metastore. This is a significant improvement over the legacy workspace-specific registry which lacked cross-workspace visibility and unified auditing.

Exam trap

Candidates often confuse Unity Catalog with legacy workspace registry features, failing to recognize that UC's main value is cross-workspace sharing and the three-level namespace structure.

18
Multi-Selecthard

A data science team is preparing to deploy a high-throughput recommendation model using Databricks Model Serving. Which TWO factors must be considered to optimize endpoint latency and resource utilization? (Choose two)

Select 2 answers
A.Configuring the minimum and maximum number of concurrent scaling units based on expected query traffic spikes.
B.Converting all feature lookup tables from Delta Lake format directly into local CSV files inside the serving container.
C.Selecting the appropriate CPU or GPU compute instance type matching the model's computational complexity and memory footprint.
D.Enabling Apache Spark adaptive query execution on the serving endpoint cluster to optimize joins during scoring.
E.Ensuring the MLflow model uses pickle serialization exclusively because it outperforms all other serialization formats.
AnswersA, C

Configuring scaling units correctly allows Databricks Model Serving to automatically scale compute resources up during traffic surges and scale down to zero during idle periods. This balance prevents latency degradation under heavy load while minimizing unnecessary cloud expenditure when request volumes drop.

Why this answer

Selecting appropriate instance types with GPU acceleration and configuring autoscaling bounds based on traffic patterns directly dictate inference latency and operational cost. These considerations are critical in production because underprovisioned endpoints cause timeout errors while overprovisioned endpoints waste cloud computing resources unnecessarily.

Exam trap

Candidates choose data storage options or training parameters instead of focusing on runtime compute configurations and scaling boundaries that directly affect latency and throughput.

19
MCQmedium

A data scientist has deployed a model to Databricks Model Serving and wants to monitor its performance over time. They need to track prediction drift and data quality issues. Which Databricks feature should they use to automatically capture inference logs and compute metrics?

A.Set up a Databricks job to periodically query the endpoint.
B.Use the model's signature to validate incoming data.
C.Configure MLflow tracking to log model predictions.
D.Enable inference tables on the serving endpoint.
AnswerD

Inference tables automatically capture request and response data from a Model Serving endpoint and store them in a Delta table. This enables monitoring of prediction drift, data quality, and model performance by analyzing the logged data. It is the native Databricks feature designed for this purpose, integrating with Lakehouse Monitoring.

Why this answer

Inference tables are a Databricks feature that automatically logs request and response payloads from a Model Serving endpoint into a Delta table. This data can then be used with Lakehouse Monitoring to track prediction drift, data quality, and model performance. Other options do not provide automatic, scalable logging of production inference data.

Exam trap

The trap here is assuming that MLflow tracking or manual queries can replace inference tables for production monitoring, but they lack automatic capture and integration with monitoring tools.

20
Multi-Selecthard

An ML engineer is deploying a model to Databricks Model Serving and needs to ensure the endpoint can handle sudden spikes in traffic without downtime. The model has a large memory footprint and takes several seconds to load. Which TWO configurations should the engineer implement to achieve this? (Choose two.)

Select 2 answers
A.Set a low maximum concurrency per instance to force more instances to be created.
B.Set the minimum provisioned concurrency to a value greater than zero to keep instances warm.
C.Configure the endpoint to use a larger workload size to accommodate the model's memory footprint.
D.Use a smaller model version to reduce load time.
E.Enable scale-to-zero to reduce costs during periods of no traffic.
AnswersB, C

Setting a minimum provisioned concurrency ensures that a specified number of model instances are always running, ready to handle requests. This eliminates cold-start latency during traffic spikes, as instances are pre-loaded with the model. For a model with a large memory footprint and slow load time, this is crucial to avoid downtime and maintain performance.

Why this answer

To handle traffic spikes without downtime for a model with a large memory footprint and slow load time, the engineer should keep instances warm by setting a minimum provisioned concurrency greater than zero, and ensure sufficient resources by selecting a larger workload size. These two configurations work together to provide immediate capacity and prevent cold starts, ensuring the endpoint remains responsive under sudden load.

Exam trap

The trap here is focusing solely on scaling out quickly while overlooking the need to keep warm instances and provide adequate memory, which are essential for models with slow load times and large footprints.

21
MCQmedium

When deploying a model to Databricks Model Serving, you notice that inference latency is higher than expected. Which diagnostic approach is most effective for identifying the bottleneck?

A.Re-train the model with a smaller dataset to see if it improves performance.
B.Check the built-in request metrics and logs for the serving endpoint to analyze duration distribution.
C.Increase the number of instances in the model serving endpoint without checking logs.
D.Restart the Databricks workspace to clear cached inference results.
AnswerB

The built-in monitoring tools provide granular data on request duration and system performance. Analyzing this distribution allows you to identify if the latency is systematic or limited to specific types of requests, enabling you to pinpoint if the bottleneck lies in compute resources, model execution, or network overhead.

Why this answer

Monitoring tools provided by Databricks, such as the built-in request metrics and logs, are essential for identifying latency bottlenecks. By examining request volume, processing duration, and resource utilization, you can determine if the latency is due to model complexity, infrastructure constraints, or external data dependencies. This allows for data-driven optimization, such as choosing a larger workload size, optimizing the model architecture, or implementing caching for frequently accessed data inputs.

Exam trap

Candidates often look to external APM tools or rewrite model code immediately, overlooking the built-in request metrics and logs readily available directly in the Databricks serving interface.

22
MCQhard

An ML engineer is deploying a model to a Databricks Model Serving endpoint. The model's inference function logs predictions to a Delta table for monitoring. During testing, they notice that the logging adds significant latency. They need to reduce the impact on inference latency. Which approach should they take?

A.Enable auto-scaling to handle the additional load from logging.
B.Reduce the frequency of logging by sampling predictions.
C.Move the logging to an asynchronous background process.
D.Increase the workload size of the serving endpoint.
AnswerC

Asynchronous logging decouples the logging operation from the inference path. The model returns predictions immediately, and logging occurs in the background. This significantly reduces latency because the response is not delayed by I/O operations. Databricks Model Serving supports asynchronous logging via custom code or by using the inference table feature, which logs asynchronously.

Why this answer

Asynchronous logging moves the logging operation outside the critical path of inference, allowing predictions to be returned without waiting for the log write to complete. This is the most effective way to reduce latency caused by logging in a serving endpoint. Databricks Model Serving supports asynchronous logging patterns, such as using the inference table or custom async code.

Exam trap

The trap here is assuming that scaling resources or reducing logging frequency will eliminate the latency, when the real issue is synchronous blocking I/O.

23
MCQeasy

A data scientist has trained a model and registered it in Unity Catalog. They now need to deploy it for real-time inference with automatic scaling and a REST API endpoint. Which Databricks feature should they use?

A.Unity Catalog functions
B.Databricks Model Serving
C.Databricks Jobs with a serving cluster
D.MLflow Model Registry webhooks
AnswerB

Databricks Model Serving provides a fully managed, serverless solution to deploy models as REST API endpoints with automatic scaling based on traffic. It integrates with Unity Catalog for model governance and supports real-time inference. This is the standard feature for deploying models for real-time serving in Databricks, offering built-in monitoring and scaling without managing infrastructure.

Why this answer

Databricks Model Serving is the purpose-built feature for deploying models as real-time REST endpoints with automatic scaling. It handles infrastructure, scaling, and monitoring, allowing data scientists to focus on model development. Other options are either for automation, batch processing, or data governance, and do not provide the required real-time serving capabilities.

Exam trap

The trap here is confusing model registry event triggers with actual serving infrastructure, or assuming that batch jobs can handle real-time requests.

24
Multi-Selectmedium

An ML engineer is deploying a model to Databricks Model Serving and needs to ensure that the endpoint can handle traffic spikes while minimizing costs during idle periods. The engineer considers enabling scale-to-zero and configuring autoscaling. Which TWO statements about these features are correct? (Choose two.)

Select 2 answers
A.Scale-to-zero and autoscaling cannot be enabled simultaneously on the same endpoint.
B.Scale-to-zero reduces replicas to zero after a period of inactivity, which can lead to cold-start latency on the next request.
C.Autoscaling scales based on CPU utilization of the serving containers, not on request concurrency.
D.Scale-to-zero is only available for GPU workload types, not for CPU workload types.
E.Autoscaling automatically adjusts the number of replicas based on incoming request load, up to a maximum configured limit.
AnswersB, E

Scale-to-zero is designed to save costs by scaling down to zero replicas when there is no traffic. When a request arrives after idle time, the endpoint must scale up, causing a cold start and increased latency for that request. This is a fundamental trade-off between cost and latency.

Why this answer

Scale-to-zero reduces replicas to zero during inactivity, causing cold-start latency on the next request. Autoscaling adjusts replicas based on load, up to a maximum. These features can be used together, with scale-to-zero effectively setting the minimum replicas to zero.

Autoscaling uses request concurrency as the primary scaling metric, not CPU utilization.

Exam trap

The trap here is thinking that scale-to-zero and autoscaling are mutually exclusive or that autoscaling scales on CPU, when it actually scales on request concurrency.

25
MCQeasy

An ML engineer needs to deploy a model to Databricks Model Serving that requires a specific version of a Python library. The library is available on PyPI. Where should the engineer specify this dependency?

A.In the Databricks cluster's init script.
B.In the model's conda.yaml file.
C.In the endpoint configuration's environment variables.
D.In the model's MLflow signature.
AnswerB

The conda.yaml file is part of the MLflow model artifact and specifies the environment dependencies, including Python packages from PyPI. When deploying to Databricks Model Serving, the serving infrastructure uses this file to build the environment. Therefore, the engineer should add the library and its version to the pip section of conda.yaml.

Why this answer

For a model deployed to Databricks Model Serving, Python dependencies are specified in the model's conda.yaml file. This file is part of the MLflow model artifact and is used by the serving infrastructure to create the environment. Adding the required library and version to the pip section ensures it is installed and available during inference.

Exam trap

The trap here is assuming that dependencies can be set via environment variables or init scripts, when they must be declared in the model's conda.yaml.

26
MCQhard

A company has deployed a model to a Databricks Model Serving endpoint. The model's predictions must be logged to a Delta table for monitoring and auditing. The ML engineer wants to enable inference logging without modifying the model's code. Which approach achieves this with minimal effort?

A.Use a Databricks job to periodically query the endpoint's logs via the REST API and insert them into a Delta table.
B.Wrap the model's predict method to write inputs and outputs to a Delta table using Spark, then redeploy the model.
C.Configure the model to log its predictions to MLflow, then enable MLflow tracking for the serving endpoint.
D.Enable inference logging on the serving endpoint by specifying a Delta table path in the endpoint configuration.
AnswerD

Databricks Model Serving supports inference logging, which automatically captures request and response payloads and writes them to a specified Delta table. This can be enabled in the endpoint configuration without changing the model code. It is the intended feature for auditing and monitoring, and requires only setting the logging destination.

Why this answer

Inference logging is a built-in feature of Databricks Model Serving that automatically logs request and response data to a Delta table. It can be enabled via the endpoint configuration without altering the model code. Other options require code changes, custom jobs, or misuse of MLflow tracking, and do not provide the same seamless auditing capability.

Exam trap

The trap here is assuming that MLflow tracking can be used for inference logging, or that manual code changes are necessary, when Databricks provides a native inference logging feature.

27
MCQeasy

A data scientist has deployed a model to a Databricks Model Serving endpoint. The endpoint is configured with scale-to-zero enabled and a workload size of Small. After a period of inactivity, the endpoint scales down to zero. A client application sends a request to the endpoint after this idle period. What happens to the first request?

A.The request fails with a 503 Service Unavailable error because the endpoint is offline.
B.The request is immediately served by a cold-start replica with no added latency.
C.The request is queued until the endpoint scales up, then processed with increased latency.
D.The request is routed to a fallback model version that is always kept warm.
AnswerC

When scale-to-zero is enabled, the endpoint scales down to zero replicas after inactivity. The first request after idle time triggers a scale-up, causing the request to be queued until a replica is available. This results in higher latency for that initial request. Subsequent requests are served with normal latency.

Why this answer

With scale-to-zero, the endpoint reduces replicas to zero after inactivity. When a new request arrives, the system must provision resources and load the model, causing the request to be queued and served with higher latency. This is expected behavior and not an error.

Exam trap

The trap here is assuming that scale-to-zero causes request failures or that a warm fallback exists, when in fact the request is queued and served after a cold start.

28
MCQeasy

An ML engineer has deployed a model to Databricks Model Serving and wants to monitor the endpoint's performance over time. They need to track the number of requests, latency, and error rates. Which Databricks feature provides these metrics out-of-the-box?

A.Delta Live Tables
B.Databricks Model Serving endpoint metrics in the Databricks UI
C.Unity Catalog audit logs
D.MLflow Tracking
AnswerB

Databricks Model Serving provides built-in metrics such as request count, latency, and error rates, which are displayed in the endpoint's detail page in the Databricks UI. These metrics are available without additional configuration and can be used to monitor the health and performance of the endpoint.

Why this answer

Databricks Model Serving includes built-in metrics that are accessible in the Databricks UI. These metrics cover request counts, latency, and error rates, providing immediate visibility into endpoint performance. Other options like MLflow Tracking or Unity Catalog audit logs serve different purposes and do not offer out-of-the-box endpoint monitoring.

Exam trap

The trap here is assuming MLflow Tracking, which is used during training, also monitors deployed endpoints, when in fact Model Serving has its own metrics dashboard.

29
MCQhard

An ML engineer is updating a production model serving endpoint to use a new model version. The endpoint currently serves version 1 with the 'Champion' alias. The engineer wants to test version 2 with a small percentage of live traffic before full rollout. Which deployment strategy should they use in Databricks Model Serving?

A.Create a new endpoint for version 2 and use a load balancer to split traffic.
B.Configure traffic splitting on the existing endpoint to route a percentage to version 2.
C.Update the 'Champion' alias to point to version 2 and monitor performance.
D.Deploy version 2 to a staging endpoint and compare metrics offline.
AnswerB

Databricks Model Serving supports serving multiple model versions within a single endpoint by specifying a traffic split percentage for each version. This allows canary deployments where a small portion of traffic goes to the new version while the majority remains on the stable version. It is the recommended approach for testing new versions with live traffic.

Why this answer

Databricks Model Serving allows you to serve multiple model versions on a single endpoint and split traffic between them by specifying percentages. This enables a canary deployment where a small fraction of requests go to the new version, allowing real-world testing before a full rollout. Other options either cause a full cutover, require external tools, or do not use live traffic.

Exam trap

The trap here is thinking that updating the alias or creating a separate endpoint is the way to do canary testing, but Databricks provides native traffic splitting within an endpoint.

30
MCQhard

A fraud detection model is deployed to a Databricks Model Serving endpoint. The team wants to test a new model version without affecting existing predictions. They need to send a copy of live traffic to the new version and log its predictions for comparison, while the current version continues to serve all responses. Which feature should they use?

A.Use MLflow's pyfunc flavor to log both models and select the active one at request time.
B.Enable the endpoint's shadow mode by specifying the new model version as a shadow.
C.Deploy the new model version to a separate endpoint and use a load balancer to mirror requests.
D.Configure the endpoint with a traffic split of 50/50 between the current and new model versions.
AnswerB

Databricks Model Serving supports shadow mode, where a copy of the request is sent to a shadow model version while the primary version's response is returned to the client. The shadow model's predictions can be logged for analysis. This exactly matches the requirement to test a new version without affecting existing predictions. The shadow model receives the same input but its output is not served, making it ideal for safe validation.

Why this answer

Shadow mode on Databricks Model Serving allows a new model version to receive a copy of live traffic while the primary version continues to serve all responses. This enables safe evaluation of the new model's predictions without impacting users. Traffic splitting would affect some responses, and external load balancing is not native and adds complexity.

Therefore, enabling shadow mode is the correct approach.

Exam trap

The trap here is confusing shadow deployment with traffic splitting; shadow mode mirrors requests without affecting responses, while traffic splitting changes which model serves the response.

31
MCQmedium

An ML engineer wants to ensure that only models that have passed a specific validation suite can be assigned the 'Champion' alias in Unity Catalog. What is the recommended way to automate this process?

A.Manually checking the validation results and updating the alias in the UI.
B.Using a Databricks Workflow to run validation and then calling the MLflow Client to update the alias.
C.Setting a SQL trigger on the Model Registry table to update the alias automatically.
D.Configuring the Model Serving endpoint to automatically promote the newest version.
AnswerB

A Databricks Workflow can orchestrate the entire process: loading the new model version, running a suite of performance and bias tests, and then using the `set_registered_model_alias` method if the tests succeed. This provides a fully automated, hands-off, and verifiable path to production.

Why this answer

Automation in the model lifecycle is best achieved using Databricks Workflows or CI/CD pipelines. By creating a task that runs a validation notebook, the system can programmatically update model aliases via the MLflow Client API only after all tests pass. This ensures a consistent, governed promotion process that prevents low-quality models from reaching production.

Exam trap

Candidates frequently select manual UI actions or legacy workspace notebooks instead of automated Databricks Workflows combined with the MLflow Client API for governed production promotions.

32
MCQhard

An ML engineer is deploying a model to Databricks Model Serving that uses a custom Python function as a pre-processing step. The function relies on a global variable defined in a separate module. After deployment, the endpoint returns errors indicating the global variable is not defined. The engineer confirmed the module is included in the model's conda environment. What is the most likely cause?

A.Databricks Model Serving runs each inference in a separate process, so global variables are not shared across requests.
B.The model was logged with `mlflow.sklearn.log_model` instead of `mlflow.pyfunc.log_model`, so custom pre-processing code is ignored.
C.The global variable is not serialized with the model because MLflow only saves the model's predict method and its direct dependencies.
D.The conda environment does not include the module because MLflow only captures packages installed via pip, not local modules.
AnswerC

MLflow's default model saving mechanism serializes the model object and its immediate dependencies, but it does not automatically capture global variables or module-level state from custom modules unless they are explicitly referenced within the model's class or function. If the pre-processing function relies on a global variable from another module, that state may not be preserved during serialization, leading to a NameError or undefined variable at serving time.

Why this answer

The error arises because MLflow's serialization does not capture global variables from external modules unless they are explicitly included in the model's code. When the model is loaded for serving, the pre-processing function may reference a global variable that was not saved, resulting in an undefined variable error. To fix this, the engineer should ensure all necessary state is encapsulated within the model or use `code_path` to include the module and initialize variables properly.

Exam trap

The trap here is assuming that any module in the conda environment will have its global state preserved, when in fact MLflow serializes only the model object and its direct code dependencies.

33
MCQmedium

An ML engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model's inference function requires access to an external feature store table for real-time feature lookup. Which approach allows the model to retrieve these features during serving while maintaining low latency and avoiding per-request authentication complexity?

A.Package the feature store data as a static file within the model artifact and load it during initialization.
B.Use the Databricks Feature Store's online table and call the feature lookup client within the model's predict method.
C.Enable the model to query the Delta table directly using Spark during each inference request.
D.Configure the serving endpoint to call an external REST API that returns feature values, using a personal access token stored in the environment.
AnswerB

Databricks Feature Store provides an online table that is optimized for low-latency lookups. The feature lookup client can be embedded in the model's predict method, and the serving endpoint automatically authenticates to the online store. This approach avoids managing credentials per request and ensures the model uses consistent feature values.

Why this answer

The Databricks Feature Store online table is purpose-built for low-latency feature serving. By embedding the feature lookup client in the model's predict method, the endpoint can retrieve features in real time without managing authentication. This integration is a core pattern for deploying models that rely on feature store data, ensuring consistency between training and serving.

Exam trap

The trap here is assuming that any external data source can be queried directly from the model without considering the serving environment's limitations and the need for optimized, authenticated access.

34
MCQhard

A financial services company uses Databricks Model Serving to deploy a real-time fraud detection model. The endpoint is configured with scale-to-zero enabled. During a period of no traffic, the endpoint scales down to zero. When a sudden burst of requests arrives, the first few requests experience high latency. Which mechanism is responsible for this behavior?

A.The requests are being queued because the endpoint's maximum concurrency is set too low.
B.The endpoint is experiencing a cold start as it provisions resources and loads the model into memory.
C.The model's dependencies are being downloaded from PyPI on each request.
D.The model's container image is being rebuilt from scratch for each request.
AnswerB

When scale-to-zero is enabled, the endpoint reduces to zero replicas during idle periods. Upon receiving new requests, it must provision new instances, pull the container image, and load the model. This cold start introduces latency for the initial requests until the endpoint is fully warmed up.

Why this answer

Scale-to-zero reduces cost by scaling down to zero replicas when idle. When traffic resumes, the endpoint must cold start: provision compute, pull the image, and load the model. This causes latency for the first few requests.

The other options describe issues that would persist regardless of scaling, such as concurrency limits or per-request builds.

Exam trap

The trap here is confusing cold start latency with concurrency or image rebuild issues, which are unrelated to scaling from zero.

35
MCQeasy

A data scientist has a model registered in Unity Catalog and wants to let an external application score it over HTTPS without embedding Databricks credentials in the application. The application's identity is already a service principal in the workspace. Which approach should be used to authenticate calls to the Model Serving endpoint?

A.Configure the endpoint to allow unauthenticated access and rely on network isolation to prevent unauthorized callers.
B.Generate a Databricks personal access token for a workspace user and hardcode it in the external application's configuration file.
C.Embed the workspace URL and a Unity Catalog metastore credential in the request body so the endpoint can verify the caller's identity from the payload.
D.Have the external application authenticate as the service principal and obtain an OAuth token from Databricks to present as a bearer credential on each request.
AnswerD

Databricks supports OAuth machine-to-machine authentication, where a service principal exchanges its client ID and secret for a short-lived OAuth token that is sent as a bearer credential. This gives the external application its own non-human identity with rotatable secrets and avoids embedding a user's personal credentials.

Why this answer

Databricks Model Serving endpoints accept OAuth bearer tokens, and service principals can perform machine-to-machine OAuth to obtain them. That gives the external application a distinct, auditable identity with rotatable secrets, avoiding a hardcoded personal token or an unauthenticated endpoint.

Exam trap

The trap here is treating network isolation or a credential placed in the request body as authentication, when the endpoint validates a bearer token from the Authorization header.

36
MCQmedium

A data scientist needs to perform batch inference on a large dataset stored in Delta Lake using a model registered in the Unity Catalog. Which approach is most efficient for leveraging Spark's distributed computing capabilities while using the MLflow model?

A.Loading the model with mlflow.pyfunc.load_model and using a for-loop to iterate over DataFrame rows.
B.Using mlflow.pyfunc.spark_udf to wrap the model and applying it to the Spark DataFrame columns.
C.Converting the Delta table to a Pandas DataFrame and using the standard model.predict() method.
D.Calling the Model Serving REST API for every row in the Delta Lake table using a standard Python request.
AnswerB

The spark_udf function automatically handles the distribution of the model environment and weights to all worker nodes. It allows the model to process data partitions in parallel, significantly reducing the time required for batch inference on massive Delta Lake tables while maintaining a simple, high-level API.

Why this answer

MLflow provides a built-in function to load models as Spark User Defined Functions (UDFs). This allows the model to be distributed across the executor nodes of a Spark cluster, enabling parallel processing of large datasets. This method is preferred over manual iteration as it integrates seamlessly with the Spark DataFrame API and optimizes resource utilization during high-volume batch jobs.

Exam trap

Candidates often try to use standard Python loops or UDFs that do not leverage Spark's parallelization, which is inefficient for large datasets and fails to utilize cluster resources.

37
MCQmedium

A fraud detection team has a model registered in Unity Catalog as main.ml.fraud_model. They need to serve it in real time, but compliance requires that every scoring request automatically generate an audit record in a Delta table, and that the model only be promoted to production after a human reviews the audit logs from a canary period. Which deployment configuration satisfies these requirements?

A.Deploy the model with Databricks Model Serving and configure the endpoint to write request logs to a mounted cloud storage path using a cluster-scoped log4j appender.
B.Deploy the model with Databricks Model Serving and enable Inference Tables, then use the endpoint's request/response logs to drive the review before promoting the model version via a UC alias.
C.Deploy the model behind a Databricks SQL warehouse by registering it as a Python UDF, and rely on the warehouse's query history as the audit trail.
D.Deploy the model with Databricks Model Serving and enable model monitoring, which persists every raw request and response to a Delta table for human review.
AnswerB

Inference Tables capture the payloads and responses of every scored request into a Delta table automatically, which is exactly the audit record compliance needs. Because the endpoint references a UC model version through an alias, the team can point the production alias at the canary version only after reviewing those logs, giving a clean, reversible promotion path.

Why this answer

Inference Tables are the Databricks mechanism that automatically logs the request and response payloads of a served model into a Delta table, which directly produces the compliance audit trail. Pairing that with a Unity Catalog alias lets the team hold production traffic on the prior version and repoint the alias only after reviewing the canary logs.

Exam trap

The trap here is assuming that enabling model monitoring also captures raw request and response payloads, when monitoring only derives metrics from the Inference Table that must be enabled separately.

38
MCQmedium

A machine learning engineer needs to deploy a custom scikit-learn model to a Databricks Model Serving endpoint with maximum throughput and minimum latency. The model requires an external preprocessing Python script during inference. Which deployment approach best leverages MLflow and Databricks architecture?

A.Log the model and preprocessing code using custom pyfunc, register it in Unity Catalog, and configure a serverless Model Serving endpoint.
B.Deploy the raw scikit-learn model artifact to Model Serving and execute the preprocessing logic inside a separate Spark structured streaming job.
C.Write the preprocessing code inside a Databricks notebook and call the model endpoint using a client-side REST API request.
D.Save the model weights to DBFS and build an external Flask application on Amazon EC2 to manage traffic routing and inference.
AnswerA

Custom pyfunc models encapsulate both the estimator and arbitrary preprocessing code into the artifact. Registering this artifact in Unity Catalog enables secure governance, and deploying it to a serverless endpoint provides auto-scaling and low-latency inference capabilities.

Why this answer

Packaging the custom preprocessing logic along with the scikit-learn estimator into an MLflow PyFunc model ensures that all inference data transformations happen inside the serving container natively. This prevents client-side processing bottlenecks and guarantees consistent feature engineering between training and real-time serving environments.

Exam trap

Candidates often suggest performing preprocessing in the Spark application before calling the model. This creates training-serving skew and latency issues compared to native PyFunc model packaging.

39
MCQmedium

Which statement best describes the role of the 'Signature' in an MLflow model when deploying to a Databricks Model Serving endpoint?

A.It is a cryptographic hash used to verify the integrity of the model files.
B.It specifies the data types and names of the input features and the output predictions.
C.It lists the specific users and groups who have permission to call the endpoint.
D.It contains the hyperparameters used during the final training run of the model version.
AnswerB

The signature serves as a contract between the model and the client. It ensures that the serving endpoint can reject malformed requests (e.g., a string where a float is expected) before they cause a crash in the model's prediction code, improving the overall robustness of the service.

Why this answer

An MLflow Model Signature defines the schema of the inputs and outputs for a model. This is crucial for Model Serving as it allows the endpoint to perform input validation before the data reaches the model. It also enables the UI to generate a 'Call' template, making it easier for developers to test the API.

Exam trap

Candidates often confuse MLflow Signatures with model artifacts or performance metrics, believing signatures dictate model accuracy or handle training data transformations automatically during inference, rather than focusing solely on schema definition.

40
MCQeasy

A data scientist has registered a model in Unity Catalog and wants to deploy it to a Databricks Model Serving endpoint. What is the simplest way to create the endpoint?

A.Create a Databricks job that runs the model's predict method on a schedule.
B.Write a Python script using the MLflow library to deploy the model to a local server.
C.Use the Databricks UI to create a new serving endpoint and select the registered model.
D.Export the model as a pickle file and upload it to a Databricks cluster.
AnswerC

The Databricks UI provides a straightforward, guided workflow to create a serving endpoint by selecting a registered model from Unity Catalog. It automatically configures the endpoint with the model's environment and dependencies, making it the simplest method for deployment.

Why this answer

The Databricks UI offers a simple, integrated way to create a Model Serving endpoint by selecting a registered Unity Catalog model. It handles configuration and deployment automatically, making it the easiest method for data scientists.

Exam trap

The trap here is confusing model deployment methods, such as local serving or batch jobs, with the managed Model Serving endpoint creation process.

41
MCQhard

A machine learning engineer has deployed a model to a Databricks Model Serving endpoint. The model requires a custom Python package that is not available in the default environment. The engineer has already logged the model with MLflow and included the package in the conda environment. However, upon deployment, the endpoint fails to start. What is the most likely cause?

A.The model's MLflow flavor is not supported by Model Serving.
B.The custom package is not available in a public PyPI repository, and no additional index URL was specified.
C.The custom package is not installed on the cluster used for serving.
D.The endpoint's workload size is too small to install the package.
AnswerB

Databricks Model Serving installs dependencies from the conda environment specified in the MLflow model. If the custom package is hosted in a private repository or requires a specific index URL, that must be included in the conda environment's pip section. Without it, the installation fails, causing the endpoint to fail to start. This is a common pitfall when using private packages.

Why this answer

When deploying a model with custom dependencies, Databricks Model Serving uses the conda environment logged with the MLflow model. If the package is not on PyPI or requires a private index, you must specify the index URL in the conda environment. Without it, the installation fails, and the endpoint cannot start.

Ensuring the environment specification includes all necessary sources is critical for successful deployment.

Exam trap

The trap here is assuming that Model Serving automatically has access to all Python packages, but it only installs what is specified in the model's conda environment, including any required index URLs.

42
MCQeasy

A data scientist has registered a model in Unity Catalog and wants to deploy it to a Databricks Model Serving endpoint. The model version is 3 and the model name is 'fraud_model'. Which identifier should be used to reference this model version when creating the endpoint?

A.The model's unique ID from the MLflow model registry, such as 'd2f1a3b4c5e6'.
B.The full three-level name and version in the format 'catalog.schema.fraud_model/3'.
C.The full three-level name 'catalog.schema.fraud_model' and the version number '3' specified separately.
D.The model name and version in the format 'fraud_model/3'.
AnswerC

In Unity Catalog, models are referenced by their three-level namespace: catalog.schema.model. The version is a separate integer. When creating a serving endpoint, you provide the model name and the version as distinct fields. This ensures the correct model version is loaded. This option correctly separates the name and version, which is the required format for Databricks Model Serving with Unity Catalog models.

Why this answer

Databricks Model Serving requires the model name and version to be specified separately. For Unity Catalog models, the name is the three-level namespace (catalog.schema.model), and the version is an integer. This allows the serving infrastructure to resolve the exact model version.

Other formats like slash-separated or using the MLflow model version ID are not supported for endpoint creation.

Exam trap

The trap here is assuming that the model version is part of the model name or that a unique ID can be used, when in fact the name and version must be provided as separate fields in the endpoint configuration.

43
MCQmedium

A machine learning engineer needs to deploy a custom PyTorch model to a Databricks Model Serving endpoint. The model requires a custom pre-processing step that is not part of the standard MLflow transformers or pyfunc flavor. Which deployment approach ensures the custom logic executes reliably within the serverless serving container?

A.Register the vanilla PyTorch state dict and apply pre-processing transformations inside the client application before sending HTTP payloads.
B.Define a separate Spark UDF inside the serving endpoint configuration file to intercept incoming JSON batches.
C.Implement a custom MLflow pyfunc PythonModel subclass that encapsulates both the PyTorch model and the custom pre-processing transformations, then log it with MLflow.
D.Store the pre-processing code in a separate volume and reference its absolute file path in the model serving endpoint environment variables.
AnswerC

Subclassing MLflow PythonModel allows developers to bundle custom inference logic, tokenizers, or scalers directly into the logged artifact. Databricks Model Serving natively understands the pyfunc flavor, executing the overridden predict method securely within the managed container environment.

Why this answer

Packaging the custom pre-processing logic directly into the MLflow pyfunc model artifact by overriding the predict context ensures that all required transformations travel with the model weights. This guarantees identical execution behavior between local testing and production serverless endpoints without relying on external pipeline code.

Exam trap

Candidates rely on default MLflow model flavors for custom architectures, forgetting that non-standard pre-processing steps require custom wrapper logic to execute inside serverless endpoints.

44
MCQmedium

An ML engineer needs to deploy a model to Databricks Model Serving that requires a custom Python package. The package is not available in the default Databricks Runtime and must be installed from a private PyPI repository. Which approach should be used to include this package in the model's environment?

A.Install the package on the driver node of the serving cluster using a %pip magic command before deployment.
B.Add the package to the cluster's init script and reference the cluster when creating the serving endpoint.
C.Upload the package as a Databricks notebook and import it at runtime within the model's predict function.
D.Include the package in the model's conda.yaml or requirements.txt file when logging the model with MLflow.
AnswerD

When logging a model with MLflow, you can specify dependencies in a conda.yaml or requirements.txt file. Databricks Model Serving uses these files to build the serving container, installing the listed packages. By including the private PyPI package with its index URL in the requirements, the serving environment will have the necessary dependency, ensuring the model runs correctly.

Why this answer

Databricks Model Serving builds a container for the model based on the environment specified during MLflow model logging. By including the private PyPI package in the conda.yaml or requirements.txt with the appropriate index URL, the package is installed in the serving environment. This is the supported method for custom dependencies.

Exam trap

The trap here is assuming that serving endpoints can use cluster init scripts or interactive pip commands, which are not available in the managed serving environment.

45
MCQmedium

A data science team is using Databricks Model Serving to deploy a model that must process sensitive data. They need to ensure that all inference requests are logged for auditing purposes, including the input data and predictions. Which approach should they take?

A.Use MLflow Tracking to log each inference request as a run in an experiment.
B.Set up a Databricks Job that periodically queries the endpoint and logs the responses.
C.Enable inference logging on the Model Serving endpoint and configure a Delta table to store the logs.
D.Configure the model to write logs to a file in DBFS using a custom Python logger.
AnswerC

Databricks Model Serving supports inference logging, which captures request and response payloads to a Delta table. This feature is designed for auditing and monitoring, allowing you to store input data and predictions. By enabling it and specifying a Delta table, the team can meet compliance requirements and analyze model performance over time.

Why this answer

Inference logging in Databricks Model Serving is the built-in feature for capturing request and response data to a Delta table. It is designed for auditing, monitoring, and compliance, providing a scalable and managed solution. Other methods like custom logging or MLflow Tracking are not suited for production-scale inference logging and may introduce reliability or performance issues.

Exam trap

The trap here is assuming that any logging mechanism (like MLflow or custom code) can serve the purpose, without recognizing that Model Serving has a dedicated inference logging feature that simplifies compliance.

46
Multi-Selectmedium

A data science team is deploying a model to Databricks Model Serving and needs to ensure that the endpoint can handle sudden spikes in traffic without dropping requests. They want to configure auto-scaling appropriately. Which TWO parameters should they adjust to control the scaling behavior? (Choose two.)

Select 2 answers
A.Enable scale-to-zero to reduce costs during idle periods.
B.Adjust the model's batch size in the predict function to process more requests per inference.
C.Set the minimum number of replicas to a value greater than zero.
D.Configure the maximum number of replicas to a high value to allow scaling out.
E.Set the endpoint's workload size to 'Large' to increase per-replica capacity.
AnswersC, D

Setting a minimum replica count ensures that a baseline capacity is always available, reducing cold start impact and allowing the endpoint to handle initial bursts without waiting for scaling. This is crucial for latency-sensitive applications that cannot tolerate cold starts.

Why this answer

Auto-scaling in Databricks Model Serving is governed by the minimum and maximum replica counts. The minimum ensures baseline capacity to handle initial bursts, while the maximum allows the endpoint to scale out to meet high demand. These two parameters directly control the scaling range and are essential for handling traffic spikes without dropping requests.

Exam trap

The trap here is focusing on cost-saving features like scale-to-zero, which actually hinder burst handling, instead of the core scaling parameters.

47
MCQhard

Refer to the exhibit. A machine learning team has updated their model serving endpoint configuration as shown in the JSON. Which deployment strategy is being implemented, and what is the primary risk associated with this specific configuration?

A.Blue/Green deployment; the primary risk is the high cost of running two identical clusters simultaneously.
B.Canary deployment; the primary risk is increased 'cold start' latency for the v2 model due to low traffic volume.
C.A/B testing; the primary risk is that the workload size 'Small' is insufficient for 90% of the traffic.
D.Shadow deployment; the primary risk is that v2 will interfere with the predictions returned by v1.
AnswerB

In a Canary setup with a 90/10 split and scale-to-zero enabled, the v2 instance will likely idle frequently. When the 10% of requests do arrive, the system must provision the instance from scratch, causing significant delays for those specific users compared to the more frequently used v1 version.

Why this answer

The exhibit demonstrates a Canary deployment where a small fraction of traffic (10%) is routed to a new model version (v2) while the majority remains on the stable version (v1). This allows for real-world testing with minimal impact. However, since both versions are set to scale-to-zero, the 10% traffic might not be frequent enough to keep v2 warm, leading to high latency for those users.

Exam trap

Candidates often overlook the 'cold start' penalty in Canary deployments. They assume that if it works for 10% of traffic, it works for everything, forgetting the resource initialization latency.

48
MCQhard

An ML engineer is deploying a model to Databricks Model Serving that uses a custom transformer requiring a GPU. The endpoint must handle high throughput with low latency. Which workload type and configuration should be selected?

A.Use a 'Large' workload type and specify GPU requirements in the model's conda environment.
B.Use a 'Medium' workload type and configure the endpoint to use GPU by setting an environment variable.
C.Use a 'Small' workload type with CPU and enable GPU acceleration via a model parameter.
D.Use a 'GPU_SMALL' or 'GPU_MEDIUM' workload type, depending on the model's memory and compute needs.
AnswerD

Databricks Model Serving offers GPU-enabled workload types such as 'GPU_SMALL' and 'GPU_MEDIUM'. These provide GPU instances suitable for models that require GPU acceleration. Selecting the appropriate size based on memory and compute requirements ensures high throughput and low latency for GPU-dependent models.

Why this answer

Databricks Model Serving provides specific GPU-enabled workload types, such as 'GPU_SMALL' and 'GPU_MEDIUM', which include GPU instances. For a model requiring GPU, selecting one of these workload types is necessary. The choice between small and medium depends on the model's memory and compute demands to achieve high throughput and low latency.

Exam trap

The trap here is assuming that GPU can be enabled through environment variables or conda specifications, when in fact it must be selected as a workload type during endpoint configuration.

49
MCQmedium

A machine learning engineer has a model registered in Unity Catalog as prod.ml.iris_model. They need to deploy it to a real-time serving endpoint that automatically scales based on traffic and provides a REST API for predictions. The model's signature is logged. Which deployment method should they use?

A.Use MLflow's built-in serving command to start a local REST server on a cluster and expose it via a public URL.
B.Use Databricks Model Serving to create a new endpoint and select the model from Unity Catalog, specifying the model version or alias.
C.Create a Databricks job that runs a Python script to load the model and listen for HTTP requests on a driver node.
D.Export the model as a Docker image using MLflow and deploy it to a Kubernetes cluster managed outside Databricks.
AnswerB

Databricks Model Serving is the managed service for real-time inference. It integrates with Unity Catalog, allowing you to select a registered model by name and version or alias. The endpoint provides a REST API, auto-scales based on load, and handles the serving infrastructure, which matches the requirement for a scalable real-time endpoint.

Why this answer

Databricks Model Serving provides a fully managed, scalable solution for deploying models registered in Unity Catalog. It automatically creates a REST endpoint, handles scaling, and integrates with governance features. The other options either rely on manual infrastructure or are intended for development, not production real-time serving.

Exam trap

The trap here is assuming that any method that exposes a REST API, such as MLflow serve or a custom Flask app, is equivalent to a managed serving endpoint.

50
MCQmedium

A machine learning engineer needs to deploy a custom PyTorch model to a Databricks Model Serving endpoint. The model requires custom post-processing logic and loading auxiliary tokenizer files alongside the serialized weights. Which approach provides the correct mechanism to package and serve this custom artifact?

A.Log the model as a standard torch.jit artifact without metadata and point the serving endpoint directly to the raw .pt file in Unity Catalog.
B.Create an external FastAPI application outside Databricks, wrap the model, and route requests via a custom reverse proxy.
C.Subclass mlflow.pyfunc.PythonModel, implement the load_context and predict methods, bundle the tokenizer files in artifacts, and log via mlflow.pyfunc.log_model.
D.Store the tokenizer files in a Delta table and configure the serving endpoint to query the table on every incoming inference request.
AnswerC

Subclassing mlflow.pyfunc.PythonModel satisfies the custom post-processing and auxiliary file constraints: load_context loads the bundled tokenizer artifacts at initialisation, while predict applies the bespoke logic around the PyTorch weights. Logging via mlflow.pyfunc.log_model registers the whole bundle, which Model Serving then deploys as a single custom artifact.

Why this answer

Custom PyTorch models requiring custom code and auxiliary files must be logged using MLflow's pyfunc flavor with custom artifact dependencies. This allows packaging artifacts and arbitrary Python code safely so that the serving infrastructure can instantiate the pyfunc wrapper, execute custom tokenization, and perform required post-processing cleanly during real-time inference.

Exam trap

Candidates mistakenly try to log raw PyTorch weight files directly without a pyfunc wrapper, failing to provide the required custom tokenization and post-processing logic inside the serving container.

51
MCQmedium

An ML engineer has registered a scikit-learn model in the Unity Catalog Model Registry and wants to serve it as a real-time endpoint using Databricks Model Serving. The model's MLflow signature expects a JSON payload with an array of records. The engineer creates a serving endpoint with a workload size of Medium and configures the served entity to use the latest model version. Which additional configuration is required to enable automatic payload logging to a Delta table for monitoring?

A.Enable inference tables on the endpoint by specifying a Unity Catalog table location.
B.Set the environment variable MLFLOW_ENABLE_SYSTEM_METRICS_LOGGING to true in the model's conda environment.
C.Attach a delivery log to the endpoint and specify a Delta table path.
D.Configure the model signature to include a 'log_payloads' parameter set to true.
AnswerA

Inference tables capture request and response payloads for served models. To enable them, you must configure the endpoint with a Unity Catalog table location where logs are written. This allows monitoring and debugging without altering the model artifact. The other options do not provide payload logging.

Why this answer

Automatic payload logging in Databricks Model Serving is achieved through inference tables, which require specifying a Unity Catalog table location when creating or updating the endpoint. This captures request and response data for monitoring and debugging. Other logging mechanisms like delivery logs or MLflow environment variables do not record inference payloads.

Exam trap

The trap here is confusing delivery logs with inference tables; delivery logs track endpoint build events, not request/response payloads.

52
Multi-Selecthard

An ML engineer is deploying a model to Databricks Model Serving and needs to enable automatic scaling based on traffic. The model has variable inference latency and the team wants to optimize cost while maintaining performance. Which TWO configurations are required to achieve this? (Choose two.)

Select 2 answers
A.Deploy the model using a GPU-enabled workload type to reduce latency.
B.Enable auto-scaling by specifying a target concurrency per replica.
C.Configure the minimum and maximum number of replicas for the endpoint.
D.Set the endpoint to use a dedicated cluster with autoscaling enabled.
E.Set the scale-to-zero option to true so the endpoint can shut down when idle.
AnswersB, C

Auto-scaling in Databricks Model Serving is driven by a target concurrency metric. You specify the desired number of concurrent requests per replica, and the system adjusts the replica count to maintain that target. This is essential for scaling based on traffic, as it directly ties replica count to request load.

Why this answer

To enable automatic scaling for a Databricks Model Serving endpoint, you must define the minimum and maximum replica counts and set a target concurrency per replica. These settings allow the serving infrastructure to adjust the number of replicas in response to traffic, ensuring performance while controlling cost. Other options either do not directly enable scaling or are not applicable to Model Serving.

Exam trap

The trap here is confusing cost-saving features like scale-to-zero with autoscaling, which requires replica bounds and a concurrency target.

53
MCQhard

A team has deployed a model to Databricks Model Serving and enabled inference tables. They notice that the inference table contains request and response payloads but no ground truth labels. They want to automatically join ground truth labels for monitoring. What should they do?

A.Use the MLflow Model Registry webhook to trigger a job that updates the inference table with ground truth labels.
B.Set the 'log_inputs' and 'log_outputs' flags to true in the endpoint configuration, which will automatically include ground truth labels.
C.Enable 'auto_capture_config' with 'ground_truth_column' set to the label column name in the inference table.
D.Configure a Databricks SQL query that joins the inference table with a ground truth Delta table and schedule it to refresh a monitoring metric.
AnswerD

Inference tables store request and response data, but ground truth labels must be provided separately. The standard approach is to create a Delta table containing ground truth labels and join it with the inference table using a unique request ID. A scheduled Databricks SQL query or a Lakeflow pipeline can perform this join and compute monitoring metrics.

Why this answer

Databricks Model Serving inference tables capture request and response payloads but do not include ground truth labels. To monitor model quality, you must join the inference table with a separate ground truth table. This is typically done by scheduling a query or pipeline that performs the join and computes metrics, which can then be used for model monitoring.

Exam trap

The trap here is believing that inference tables automatically capture ground truth labels or that a simple configuration flag can enable it, when in reality ground truth must be joined from an external source.

54
MCQhard

Refer to the exhibit. An administrator notices that the cost for this specific endpoint is higher than expected even when there is no traffic. Based on the exhibit, what is the most likely cause of the high idle cost?

A.The 'entity_version' is set to a legacy version that uses more expensive hardware.
B.The 'scale_to_zero_enabled' parameter is set to false, keeping an instance active at all times.
C.The endpoint is using a GPU-accelerated instance by default for all Unity Catalog models.
D.Multiple versions of the model are being served simultaneously, doubling the cost.
AnswerB

When scale-to-zero is disabled, the system maintains the minimum number of provisioned instances (defaulting to 1). This provides the benefit of zero latency for the first request after an idle period but results in constant resource consumption and associated cloud costs.

Why this answer

The configuration shows that 'scale_to_zero_enabled' is set to false. This means that at least one instance of the model is running 24/7, regardless of whether any requests are being made. While this ensures there is never a cold start, it leads to continuous billing for the compute resources even during nights and weekends.

Exam trap

Candidates often assume endpoints automatically scale down to zero when idle, forgetting that scale-to-zero must be explicitly enabled to avoid continuous compute charges.

55
MCQhard

An ML engineer has a Databricks Model Serving endpoint that is currently serving a registered model version. A new model version is registered in Unity Catalog and must be rolled out to the endpoint without any downtime. Which approach should the engineer use to safely transition traffic to the new model version?

A.Modify the model version in Unity Catalog by overwriting the existing version with the new model artifacts.
B.Create a new endpoint with the new model version and manually update the client applications to point to the new endpoint URL.
C.Update the served entities of the endpoint to point to the new model version and rely on Databricks to perform a rolling update.
D.Delete the existing endpoint and recreate it with the new model version.
AnswerC

Databricks Model Serving supports updating the served entities of an existing endpoint to reference a new model version. The platform performs a zero-downtime rolling update, provisioning new capacity with the updated model before shifting traffic, so in-flight requests continue to be served by the old version until the new one is ready.

Why this answer

Updating the served entities of an existing endpoint is the supported method for zero-downtime model updates in Databricks Model Serving. The platform handles the rolling update by provisioning new resources with the updated model version and shifting traffic only when ready, so the endpoint remains available throughout the process.

Exam trap

The trap here is assuming that any change to the served model requires endpoint recreation or a new endpoint, when in fact updating the served entities of the existing endpoint triggers a zero-downtime rolling update.

56
MCQhard

A machine learning engineer is deploying a model to Databricks Model Serving and wants to implement a blue-green deployment strategy. They have registered two model versions in Unity Catalog: version 1 (current production) and version 2 (new candidate). They want to route 10% of traffic to version 2 for testing while keeping 90% on version 1. Which feature should they use to achieve this?

A.Create two separate endpoints and use a load balancer to distribute traffic based on weights.
B.Use the model registry's stage transitions to mark version 2 as 'Staging' and version 1 as 'Production', then enable automatic traffic mirroring.
C.Configure the endpoint with two served entities and use traffic splitting percentages.
D.Deploy version 2 as a separate endpoint and use Unity Catalog aliases to switch traffic instantly.
AnswerC

Databricks Model Serving supports serving multiple model versions within a single endpoint by defining multiple served entities. Each served entity can be assigned a percentage of traffic. By setting version 1 to 90% and version 2 to 10%, the engineer achieves a blue-green or canary deployment. This allows safe testing of the new version.

Why this answer

Databricks Model Serving allows multiple served entities per endpoint, each with a traffic percentage. This enables canary or blue-green deployments by routing a portion of traffic to a new model version while the rest goes to the stable version. Other methods like separate endpoints or aliases do not provide built-in percentage-based splitting.

Exam trap

The trap here is confusing model registry aliases or stages with traffic splitting; aliases switch all traffic, while traffic splitting within an endpoint allows gradual rollout.

57
MCQmedium

An ML engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model expects a single feature vector of 10 float values per request. The endpoint must return predictions in under 100 ms. Which approach should the engineer use to minimize per-request overhead?

A.Log the model with a signature that specifies a tensor input and enable the endpoint's request batching.
B.Enable autoscaling on the endpoint and send requests with a batch size of 1.
C.Use MLflow's pyfunc flavor with a custom predict method that processes one row at a time.
D.Log the model using the native scikit-learn flavor and deploy it with a workload size that provides sufficient CPU.
AnswerD

The native scikit-learn flavor allows the model server to load and invoke the model directly without an extra Python wrapper, reducing per-request overhead. Choosing an appropriate workload size ensures enough CPU resources to handle the model's computation quickly. This combination minimizes framework overhead and provides the necessary compute, making it the best choice to meet the sub-100 ms latency requirement for single-vector requests.

Why this answer

For low-latency single-vector inference, minimizing framework overhead is critical. The native scikit-learn flavor avoids the extra Python layer that pyfunc introduces, and a suitable workload size provides the CPU needed to execute the model quickly. Autoscaling and request batching address throughput rather than per-request latency, so they do not help meet the strict latency target.

Exam trap

The trap here is assuming that autoscaling or request batching reduces the latency of a single small request, when they primarily improve throughput under load.

58
MCQhard

A financial institution deploys a credit scoring model using Databricks Model Serving. The model must log all incoming requests and outgoing responses to a Delta table for auditing. The ML engineer needs to enable this logging with minimal performance impact. Which solution should they implement?

A.Enable inference logging on the serving endpoint, specifying a Delta table location for logs.
B.Implement a custom logging wrapper in the model's predict method that writes each request to a Delta table.
C.Use a Databricks job to periodically query the endpoint's access logs and write them to a Delta table.
D.Configure the model to send logs to an external Kafka topic, then use a Databricks job to ingest from Kafka into Delta Lake.
AnswerA

Databricks Model Serving provides built-in inference logging that captures request and response payloads and writes them to a Delta table. This feature is designed for auditing and monitoring, and it operates asynchronously to minimize impact on inference latency. Configuring the endpoint with a Delta table path enables this logging.

Why this answer

Databricks Model Serving includes a built-in inference logging feature that asynchronously logs request and response data to a Delta table. This is the most efficient and least intrusive method. Custom wrappers or external systems add latency and complexity.

Access logs do not contain payloads. Therefore, enabling inference logging directly is the correct solution.

Exam trap

The trap here is assuming that access logs contain full payloads or that custom logging is necessary, when Databricks provides a native asynchronous logging feature.

59
MCQeasy

An ML engineer needs to deploy a model to Databricks Model Serving. The model was logged with MLflow and registered in Unity Catalog. The engineer wants to ensure that only the latest version of the model is served and that the endpoint can be updated without downtime. Which approach should they use?

A.Delete the existing endpoint and create a new one with the same name and configuration, pointing to the new model version.
B.Use MLflow's transition request to move the model to Production stage, which automatically updates the serving endpoint.
C.Create a new endpoint for each model version and update the client application to point to the new endpoint URL.
D.Update the existing endpoint to serve the new model version using the Databricks UI or REST API, which performs a rolling update.
AnswerD

Databricks Model Serving allows you to update an endpoint to a new model version without downtime. The service performs a rolling update, gradually shifting traffic to the new version while maintaining availability. This ensures that the latest version is served and clients continue to use the same endpoint URL.

Why this answer

Updating an existing endpoint to a new model version via the UI or REST API triggers a rolling update, ensuring zero downtime and keeping the same endpoint URL. This is the standard method for deploying a new version without disrupting clients.

Exam trap

The trap here is assuming that MLflow stages or aliases automatically update serving endpoints, or that recreating an endpoint is necessary for updates.

60
MCQhard

An ML engineer is deploying a model to Databricks Model Serving and wants to implement A/B testing between two model versions. The engineer needs to route a percentage of traffic to each version and collect performance metrics. Which feature of Databricks Model Serving should the engineer use?

A.Deploy two separate endpoints and use an external load balancer to distribute traffic between them.
B.Enable the 'Canary' deployment option in the endpoint configuration to automatically split traffic.
C.Create a single endpoint with multiple model versions and configure traffic splitting between them.
D.Use MLflow's model registry to assign a stage to each model version and route traffic based on the stage.
AnswerC

Databricks Model Serving supports serving multiple model versions on a single endpoint with traffic splitting. You can specify the percentage of traffic routed to each version, enabling A/B testing. This is the built-in feature for such scenarios, allowing you to compare performance and metrics.

Why this answer

To perform A/B testing between model versions, the engineer should create a single Model Serving endpoint that serves multiple model versions with traffic splitting. This allows a specified percentage of requests to be routed to each version, and Databricks provides metrics for each version, facilitating performance comparison. This native feature simplifies A/B testing.

Exam trap

The trap here is thinking that MLflow stages or separate endpoints with external load balancers are needed for A/B testing, when Databricks Model Serving provides built-in traffic splitting on a single endpoint.

61
MCQhard

A team has deployed a model to Databricks Model Serving and wants to enable autoscaling to handle variable traffic. They configure the endpoint with scale_to_zero_enabled set to true and a min_provisioned_concurrency of 0. After deployment, they notice that the endpoint takes several seconds to respond to the first request after a period of inactivity. What is the cause of this latency?

A.The model is being reloaded from Unity Catalog on each request, causing cold-start latency.
B.The endpoint is using a GPU workload, and GPU initialization takes several seconds after each idle period.
C.The model's Python environment is being reinstalled on each request due to missing dependencies in the container image.
D.The endpoint is scaling from zero, which requires provisioning resources and loading the model before serving the request.
AnswerD

When scale_to_zero_enabled is true and min_provisioned_concurrency is 0, the endpoint scales down to zero replicas during inactivity. The first request after idle time triggers a cold start: the system must provision compute resources and load the model into memory. This provisioning and loading time causes the several-second latency observed, which is inherent to scale-to-zero behavior.

Why this answer

Enabling scale_to_zero with zero minimum concurrency allows the endpoint to shut down completely when idle, reducing cost. However, the next request must wait for compute resources to be provisioned and the model to be loaded, resulting in cold-start latency. To avoid this, set a min_provisioned_concurrency greater than zero or disable scale_to_zero, keeping at least one replica warm.

Exam trap

The trap here is attributing cold-start latency to model artifact retrieval or environment setup instead of the scale-from-zero provisioning process.

Ready to test yourself?

Try a timed practice session using only Model Deployment questions.