A machine learning engineer needs to deploy a Generative AI application using Databricks Model Serving. The application requires access to a private vector database within the same VPC. Which configuration step is mandatory for this deployment?
Trap 1: Enable public endpoint access on the Model Serving cluster.
Enabling public access exposes the model endpoint to the internet, which violates security policies for private data integration. Model serving requires internal network routing to access private VPC resources securely, so public exposure is counterproductive and insecure for enterprise-grade generative AI applications that handle sensitive internal documentation.
Trap 2: Deploy the model as a library within a standard Databricks notebook.
Deploying a model as a library inside a notebook is not suitable for high-concurrency production serving. Model Serving provides a managed, scalable REST API endpoint that supports automatic scaling and resource isolation, which is necessary for serving generative AI models reliably in a production environment compared to interactive notebooks.
Trap 3: Increase the maximum instance count to ensure network path…
Increasing the instance count only addresses horizontal scaling and does not solve the fundamental network routing problem between the serving endpoint and the private VPC. Network connectivity policies are distinct from cluster scaling configurations and must be explicitly defined to establish paths to restricted internal network resources.
- A
Enable public endpoint access on the Model Serving cluster.
Why it fails: Enabling public access exposes the model endpoint to the internet, which violates security policies for private data integration. Model serving requires internal network routing to access private VPC resources securely, so public exposure is counterproductive and insecure for enterprise-grade generative AI applications that handle sensitive internal documentation.
- B
Configure a custom serving endpoint with a private network connection policy.
Configuring a serving endpoint with proper network connectivity settings allows the model to communicate with internal resources. By leveraging private connectivity, the model serving infrastructure can reach the VPC-hosted vector database while keeping traffic isolated from the public web, fulfilling the architectural requirements for a secure RAG deployment.
- C
Deploy the model as a library within a standard Databricks notebook.
Why it fails: Deploying a model as a library inside a notebook is not suitable for high-concurrency production serving. Model Serving provides a managed, scalable REST API endpoint that supports automatic scaling and resource isolation, which is necessary for serving generative AI models reliably in a production environment compared to interactive notebooks.
- D
Increase the maximum instance count to ensure network path availability.
Why it fails: Increasing the instance count only addresses horizontal scaling and does not solve the fundamental network routing problem between the serving endpoint and the private VPC. Network connectivity policies are distinct from cluster scaling configurations and must be explicitly defined to establish paths to restricted internal network resources.