Databricks-GenAI-Assoc · domain
Design Applications
The Design Applications domain covers architecting RAG and generative AI applications on Databricks: chunking and embedding with Foundation Model APIs, vector search indexes, prompt objects, multi-turn conversation state, secret management via Databricks secret scopes, and evaluating retrieval quality with metrics like context relevance and groundedness. Questions are scenario-based, asking you to pick the right Databricks component or design pattern.
Focused practice
Practice Design Applications questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Design Applications
You must be able to choose the right Databricks building blocks for a RAG or GenAI app: vector search for retrieval, prompt objects for templating, secret scopes for credentials, and history-aware query construction for multi-turn chat. The single most important thing is matching each design need to the correct Databricks feature.
Selecting retrieval evaluation metrics such as context relevance and groundedness for RAG pipelines
Designing multi-turn RAG where follow-up questions need query rewriting or conversation history in retrieval
Storing LLM API keys using Databricks secret scopes and dbutils.secrets rather than hardcoding
Using the Prompt object and MLflow prompt registry to version and serve RAG prompts
Watch out for
Common Design Applications exam traps
- ▸Confusing answer quality metrics like correctness with retrieval metrics such as context relevance when the question asks about retrieved context quality
- ▸Ignoring conversation history in retrieval, so follow-up questions retrieve irrelevant chunks because the query lacks prior-turn context
- ▸Embedding API keys directly in notebooks or code instead of referencing Databricks secret scopes via dbutils.secrets.get
Question index
All Design Applications questions (57)
Click any question to see the full explanation, or start a practice session above.
A team is designing a GenAI application that must call an external LLM provider from a Databricks notebook and from a Model Serving endpoint. The security team requires that the provider API key never appear in notebook source code, Git history, or the model artifact, and that the same governed credential be reused by both the notebook and the serving endpoint. Which design satisfies these requirements?
Easy2When designing a production RAG application, which technique is most effective for preventing the LLM from hallucinating based on outdated information?
Medium3An engineer has a Databricks Vector Search index that stores chunk embeddings for a technical manual. Users report that queries containing exact part numbers return irrelevant chunks because the embedding model blurs numeric tokens. The engineer wants retrieval to consider both dense vector similarity and exact keyword matching on the same Delta table without building a separate search system. Which Databricks Vector Search feature should the engineer enable?
Hard4When designing a RAG application for a regulated industry, what is the most important Unity Catalog feature for ensuring data compliance during retrieval?
Medium5A generative AI engineer is building a RAG chain on Databricks using LangChain. The chain must call an external LLM provider via a Databricks Model Serving endpoint that proxies the provider, and the engineer wants to avoid hardcoding credentials in notebooks. Which approach should the engineer use to authenticate calls from the chain to the serving endpoint?
Medium6A GenAI engineer is building a customer-support assistant on Databricks. The assistant must invoke an external LLM endpoint that requires an API token, and the token must not appear in notebook source or logs. The engineer wants to store the token once and let notebooks and Databricks Jobs reference it by name. Which Databricks capability should the engineer use to satisfy this requirement?
Medium7In a RAG application, which architectural component ensures that the system handles changes in the source data effectively?
Hard8Which TWO factors should be considered when choosing an embedding model for a RAG application?
Medium9An engineer is designing a GenAI application where a Databricks workflow must call an external LLM provider that enforces a strict rate limit and occasionally returns transient errors. The engineer wants the application to degrade gracefully instead of failing the whole job. Which design choice best addresses this requirement?
Medium10A team is deploying a Databricks RAG chatbot that must serve interactive traffic with low latency while also allowing the data science team to test prompt variations safely. The engineer wants the production endpoint to keep serving stable traffic even while a new prompt template is being evaluated. Which approach best meets these requirements?
Medium11Which THREE strategies improve the quality of retrieval in a Databricks Vector Search-based RAG application?
Medium12Which Databricks component is recommended for orchestrating the end-to-end RAG pipeline, including data ingestion, transformation, and vector indexing?
Easy13A team is designing a Databricks GenAI application that must return grounded answers with citations to source documents. The application uses Databricks Vector Search for retrieval and a Foundation Model API for generation. Which TWO design choices are required to return accurate citations alongside each answer? (Choose two.)
Hard14An organization needs to build a RAG application on Databricks that minimizes data egress and maximizes security by keeping all data within the workspace perimeter. Which architectural pattern best satisfies this requirement?
Medium15Which TWO factors should be prioritized when selecting an embedding model for a domain-specific RAG application on Databricks?
Medium16A team is deploying a RAG chatbot that answers questions from internal policy documents. The documents change frequently, and the team wants retrieval to reflect updates within minutes without re-running a full embedding job. The source Delta table already has change data feed enabled. Which approach should the team use to keep the Databricks Vector Search index current?
Easy17An engineer is designing a Databricks RAG application where the same retrieved context must be reused across several prompt variants during evaluation. They want to reduce token cost and improve consistency between variants. Which design choice best supports this?
Hard18A generative AI engineer is designing a multi-stage RAG application on Databricks. The application first retrieves documents using Vector Search, then reranks them with a cross-encoder model, and finally calls a foundation model endpoint to generate an answer. The engineer wants to ensure that the entire pipeline is reproducible and that each stage can be independently versioned and deployed. Which design approach best meets these requirements?
Hard19A GenAI engineer is designing a multi-turn chat application on Databricks. Users report that the assistant forgets details from earlier in long conversations and sometimes answers using only the most recent message. The team wants the model to reliably use facts stated several turns earlier without exceeding the model's context window. Which design should the engineer implement?
Medium20A GenAI engineer is designing a Databricks application that must ground answers in a large corpus of internal policy documents. The corpus is updated by a nightly Delta job, and the application must cite the source document for every answer. The engineer is deciding how to structure the retrieval and generation stages. Which TWO design choices best satisfy the grounding and citation requirements? (Choose two.)
Hard21A GenAI engineer is building a RAG chatbot backed by a Databricks Vector Search index that is synced from a Delta table with Change Data Feed enabled. Source documents are frequently updated, and some are deleted. The engineer wants the index to reflect deletions and updates automatically with minimal operational overhead. Which configuration should be used?
Medium22A team is deploying a RAG application using Databricks Model Serving with a foundation model endpoint and a Databricks Vector Search index. During load testing, they observe that p95 latency spikes when the retriever returns many chunks, and the LLM occasionally truncates context. They want to reduce latency while preserving answer quality. Which change is most appropriate?
Hard23What is the primary benefit of using Unity Catalog when designing generative AI applications in Databricks?
Easy24An engineer is building a GenAI application that must return structured JSON output conforming to a specific schema so downstream systems can parse it reliably. The team wants to enforce the schema at generation time rather than post-processing free-form text. Which approach should they use?
Easy25A financial services company is designing a RAG assistant that must never return answers containing personally identifiable information from its knowledge base. The team plans to filter retrieved chunks before they are placed into the prompt. Which Databricks design element should they use to enforce this filtering consistently?
Medium26Which design pattern is best for protecting the LLM from prompt injection attacks when building a customer-facing chatbot on Databricks?
Medium27A team is designing a Databricks RAG application that must return answers with citations to source documents. They plan to use Databricks Vector Search and want the LLM to reference specific chunks. Which two design choices are required to produce reliable citations? (Choose two.)
Hard28Refer to the exhibit. The Databricks Vector Search index is currently in an 'ONLINE' status with incremental sync. If a developer performs a DELETE operation on the source Delta table, what happens to the vector index?
Medium29An engineer is designing a GenAI application that must call an external LLM provider's API. The provider key must not appear in notebook code, job logs, or Git. The team already uses Databricks and wants the key to be injected into the serving endpoint at runtime. Which Databricks capability should the engineer use?
Easy30Which Databricks feature is specifically designed to monitor model quality and drift in production?
Medium31A data science team wants to expose a RAG chain as a REST API so that an external web application can send questions and receive answers. The chain is developed with Databricks LangChain integrations and must be deployed with autoscaling and built-in monitoring. Which Databricks capability should they use?
Easy32A generative AI engineer is designing a RAG application on Databricks that uses a foundation model served via Databricks Model Serving. The application must handle peak loads gracefully and provide consistent response times. The engineer is evaluating design patterns for scaling and reliability. Which TWO design choices should the engineer implement? (Choose two.)
Hard33Refer to the exhibit. What is the correct way to log a custom RAG chain so it can be loaded using the provided code?
Medium34A generative AI engineer is designing a Databricks RAG application that ingests a Delta table containing 400 million support articles into a Databricks Vector Search index. The team wants the lowest-latency online serving with the smallest possible index while preserving retrieval quality for the most common queries. Which design decision best meets these requirements?
Medium35A GenAI engineer is building a Databricks RAG application where the retrieval step returns the top-5 chunks for each user question. The engineer wants to add a second LLM call that evaluates whether each retrieved chunk contains enough information to answer the question, and then filters out chunks that fail this evaluation before passing the remaining chunks to the final answer-generation prompt. Which design pattern is the engineer implementing?
Hard36A GenAI application uses a Databricks Model Serving endpoint hosting a foundation model, and the team wants structured JSON output that conforms to a fixed schema for downstream parsing. Responses sometimes include prose or markdown fences that break the parser. Which design change is most likely to produce reliably parseable output?
Hard37An engineer needs to ensure that only authorized users can access the RAG chatbot. Which Databricks security feature should be used to enforce this access control?
Medium38When designing an application that requires fine-tuning a small model (like Llama-3-8B) on Databricks, which THREE factors must be considered to ensure a successful training job?
Hard39A GenAI engineer is building a Databricks RAG application that answers questions over a Delta table containing 40 million support tickets. Users report that simple keyword lookups return irrelevant results because the tickets use inconsistent terminology. The engineer needs semantic retrieval that stays synchronized as new tickets stream in every few minutes. Which Databricks component should be used to serve this retrieval layer?
Medium40A generative AI engineer is designing an agent on Databricks that uses a LangChain agent with tool-calling capabilities. The agent must call a Databricks SQL warehouse to run queries and a Vector Search index for retrieval. Which design consideration is most important for controlling agent behavior in production?
Hard41When evaluating the performance of a RAG application, which metric is most useful for measuring the quality of the retrieved context?
Medium42An engineer is designing a Databricks RAG application that must support multi-turn conversations where follow-up questions refer to earlier turns. They want the retrieval step to remain accurate as the conversation progresses. Which TWO design elements should they include? (Choose two.)
Hard43A GenAI engineer is designing an agentic application on Databricks that uses a foundation model to decide which external tools to call. The team wants the agent to be able to invoke a Databricks SQL warehouse query and a Python function registered as a Unity Catalog function, and they need the model to select tools based on natural language requests. Which design element is required for the model to select and invoke these tools correctly?
Hard44Which approach is recommended for managing secrets, such as API keys for external LLM providers, in Databricks?
Medium45A GenAI team at a retail company has built a RAG chatbot on Databricks that answers customer questions from a product catalog stored in Delta Lake. The catalog is updated nightly, and the team wants the chatbot to reflect those updates without manual intervention. They are deciding how to keep the Vector Search index synchronized. Which approach should they use?
Medium46What is the primary function of the 'Prompt' object in a Databricks RAG architecture?
Easy47An engineer is designing a RAG application that uses Databricks Vector Search to retrieve documents and a foundation model endpoint to generate answers. The team wants to log all user queries, retrieved documents, and generated responses for auditing and continuous improvement. They also need to monitor for drift in retrieval quality over time. Which Databricks capability should they integrate into the application design?
Medium48You are building an application that uses Model Serving to host a fine-tuned LLM. Which configuration is required to optimize for high-concurrency request throughput?
Medium49A GenAI engineer is designing a RAG application on Databricks that must support multi-turn conversations where users refer to earlier messages, and the application must keep responses grounded in retrieved documents. Which TWO design elements are required to meet these requirements? (Choose two.)
Medium50An application requires low-latency retrieval of RAG metadata stored in Databricks. Which storage approach balances performance and cost while ensuring seamless integration with Unity Catalog?
Medium51An engineer is building a multi-turn GenAI assistant on Databricks. The assistant must answer follow-up questions that reference earlier turns, such as 'what about its warranty?', while keeping each request within the model's context limit. Which design should the engineer implement?
Medium52Refer to the exhibit. An engineer is testing a model endpoint. The outputs are too brief and often stop mid-sentence. What is the most likely cause?
Hard53Refer to the exhibit. What is the cause of this error when logging a RAG chain to MLflow?
Medium54A GenAI engineer is designing a retrieval-augmented generation application whose source documents are long PDFs. Early testing shows that answers are vague because retrieved chunks contain several unrelated topics, and the language model frequently cites content that does not support its claims. The engineer wants to improve chunk quality before indexing. Which TWO changes should the engineer make to the ingestion pipeline? (Choose two.)
Hard55A generative AI team is building a customer-support assistant on Databricks. The assistant must answer questions using the company's private knowledge base, and the team wants to minimize latency while ensuring that the LLM only uses retrieved documents. They plan to use Databricks Vector Search with a Delta table as the source. Which design choice best balances low latency and grounded responses?
Medium56An engineer is designing a RAG application where the LLM must answer questions using only the retrieved context and must refuse to answer when the context is insufficient. Which prompt design approach best enforces this behavior?
Medium57An engineer is designing a Databricks GenAI application that summarizes customer meeting notes. The notes contain personally identifiable information, and company policy requires that the summarization model never sends note text to an external provider. The team already has a fine-tuned open model registered in Unity Catalog. Which deployment choice satisfies the policy?
EasyOther domains
All Databricks-GenAI-Assoc exam domains
Frequently asked questions
- What does the Design Applications domain cover on the Databricks-GenAI-Assoc exam?
- You must be able to choose the right Databricks building blocks for a RAG or GenAI app: vector search for retrieval, prompt objects for templating, secret scopes for credentials, and history-aware query construction for multi-turn chat. The single most important thing is matching each design need to the correct Databricks feature.
- How many questions are in this domain?
- This page lists all 57 Design Applications questions in the Databricks-GenAI-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Design Applications questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.