A generative AI engineer is designing a multi-stage RAG application on Databricks. The application first retrieves documents using Vector Search, then reranks them with a cross-encoder model, and finally calls a foundation model endpoint to generate an answer. The engineer wants to ensure that the entire pipeline is reproducible and that each stage can be independently versioned and deployed. Which design approach best meets these requirements?
MLflow models encapsulate code, environment, and signatures, enabling independent versioning. Composing them into a pipeline allows the entire multi-stage application to be logged as a single model, which can be served with Databricks Model Serving. This provides reproducibility and independent versioning of each stage while offering a unified deployment artifact. It is the recommended pattern for complex generative AI pipelines.
Why this answer
Packaging each stage as an MLflow model with its own signature and dependencies allows independent versioning. Composing them into a single MLflow pipeline enables logging and serving the entire multi-stage RAG application as one model, ensuring reproducibility. This approach is supported by Databricks Model Serving and MLflow, and it simplifies deployment and tracing compared to separate endpoints or scripts.
Exam trap
The trap here is assuming that separate endpoints or a single script provide sufficient versioning and reproducibility, when the key is to encapsulate each stage as an MLflow model and compose them into a unified pipeline.