Which TWO of the following are mandatory requirements for developing an AI application using the Databricks Mosaic AI Model Serving environment?
Trap 1: The model must be stored in a legacy Databricks Workspace folder.
Legacy workspace paths are not compatible with the modern Unity Catalog serving architecture. All models must be registered within Unity Catalog to leverage the unified governance model, lineage tracking, and permission structures that Mosaic AI services require to function effectively within a secure enterprise data environment.
Trap 2: The model must be served using a shared-access mode interactive…
Model serving endpoints use dedicated, managed inference infrastructure, not interactive clusters. Interactive clusters are designed for development and exploratory analysis, whereas inference services require specific, optimized runtime environments that provide scalability, high availability, and auto-scaling capabilities tailored specifically for model serving workloads rather than general-purpose data science tasks.
Trap 3: The application code must perform manual model sharding across…
Databricks Mosaic AI handles model sharding and distribution automatically within the managed serving environment. Developers do not need to manually manage hardware-level parallelism or model sharding, as the infrastructure layer abstracts these complexities to focus on ease of deployment, scalability, and simplified management of inference workloads.
- A
The model must be stored in a legacy Databricks Workspace folder.
Why it fails: Legacy workspace paths are not compatible with the modern Unity Catalog serving architecture. All models must be registered within Unity Catalog to leverage the unified governance model, lineage tracking, and permission structures that Mosaic AI services require to function effectively within a secure enterprise data environment.
- B
The model must be registered as a model version within a Unity Catalog schema.
Unity Catalog acts as the central repository for model artifacts and versions in Databricks. Registering the model here provides the necessary metadata, lineage, and access control required by the serving infrastructure to deploy the model securely and ensure it remains reachable by authorized internal or external applications.
- C
The model must be served using a shared-access mode interactive cluster.
Why it fails: Model serving endpoints use dedicated, managed inference infrastructure, not interactive clusters. Interactive clusters are designed for development and exploratory analysis, whereas inference services require specific, optimized runtime environments that provide scalability, high availability, and auto-scaling capabilities tailored specifically for model serving workloads rather than general-purpose data science tasks.
- D
The serving endpoint must be configured with a defined compute resource.
An inference endpoint needs specific compute resources (like GPU instances) to execute model logic. Defining this configuration is a mandatory step in the deployment process, allowing the platform to provision the necessary infrastructure to meet the required throughput and latency demands of the model being deployed for users.
- E
The application code must perform manual model sharding across nodes.
Why it fails: Databricks Mosaic AI handles model sharding and distribution automatically within the managed serving environment. Developers do not need to manually manage hardware-level parallelism or model sharding, as the infrastructure layer abstracts these complexities to focus on ease of deployment, scalability, and simplified management of inference workloads.