Databricks-GenAI-Assoc Design Applications Practice Question
A data science team wants to expose a RAG chain as a REST API so that an external web application can send questions and receive answers. The chain is developed with Databricks LangChain integrations and must be deployed with autoscaling and built-in monitoring. Which Databricks capability should they use?
⚠ Common exam trap
The trap here is equating any Databricks compute that can run Python with a production API endpoint, overlooking that Model Serving is the managed serving layer.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Model Serving with a custom MLflow pyfunc model that wraps the chain.
Model Serving is the Databricks capability for hosting models and custom Python logic as REST endpoints with autoscaling and monitoring. Wrapping the LangChain chain in an MLflow pyfunc model makes it deployable through that service, so the web application can call a stable API. Scheduled jobs, SQL warehouses, and manual Flask apps do not provide the same managed serving characteristics.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
A Databricks job that runs the chain on a schedule and writes answers to a Delta table.
Why it's wrong here
A scheduled job is batch-oriented and does not expose a synchronous REST API for interactive web requests. It would introduce latency and require the web application to poll a table, which is not how a chatbot API is consumed. It also lacks the request-level autoscaling and serving metrics that Model Serving provides.
- ✓
Model Serving with a custom MLflow pyfunc model that wraps the chain.
Why this is correct
Databricks Model Serving deploys MLflow models as scalable REST endpoints and supports custom pyfunc models, which can wrap a LangChain chain. It provides autoscaling, request logging, and integration with inference tables for monitoring, matching the requirement to serve the chain as an API with operational visibility.
- ✗
An all-purpose cluster with a Flask app started manually on the driver node.
Why it's wrong here
Running Flask on an all-purpose cluster driver is not a managed serving solution. The endpoint would not autoscale, would lack built-in monitoring, and would be tied to a cluster lifecycle that is not intended for production APIs. It also bypasses Databricks governance and logging features that Model Serving offers.
- ✗
A Databricks SQL warehouse with a query that calls the chain through a UDF.
Why it's wrong here
SQL warehouses are optimized for SQL analytics, not for hosting arbitrary Python RAG chains as low-latency APIs. While UDFs can call external services, they are not designed to serve an interactive chatbot with autoscaling and request monitoring. This approach would be awkward for the web application and would not provide the needed serving features.
About these practice questions
Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.