You must be able to choose the right Databricks building blocks for a RAG or GenAI app: vector search for retrieval, prompt objects for templating, secret scopes for credentials, and history-aware query construction for multi-turn chat. The single most important thing is matching each design need to the correct Databricks feature.
Start practicing
Design Applications — choose a session length
Free · No account required
Domain overview
The Design Applications domain covers architecting RAG and generative AI applications on Databricks: chunking and embedding with Foundation Model APIs, vector search indexes, prompt objects, multi-turn conversation state, secret management via Databricks secret scopes, and evaluating retrieval quality with metrics like context relevance and groundedness. Questions are scenario-based, asking you to pick the right Databricks component or design pattern.
Exam objectives
Selecting retrieval evaluation metrics such as context relevance and groundedness for RAG pipelines
Designing multi-turn RAG where follow-up questions need query rewriting or conversation history in retrieval
Storing LLM API keys using Databricks secret scopes and dbutils.secrets rather than hardcoding
Using the Prompt object and MLflow prompt registry to version and serve RAG prompts
Confusing answer quality metrics like correctness with retrieval metrics such as context relevance when the question asks about retrieved context quality
Ignoring conversation history in retrieval, so follow-up questions retrieve irrelevant chunks because the query lacks prior-turn context
Embedding API keys directly in notebooks or code instead of referencing Databricks secret scopes via dbutils.secrets.get
Click any question to see the full explanation and answer options, or start a focused practice session above.
An organization needs to build a RAG application on Databricks that minimizes data egress and maximizes security by keeping all data within the workspace perimeter. Which architectural pattern best satisfies this requirement?
2Which TWO factors should be prioritized when selecting an embedding model for a domain-specific RAG application on Databricks?
3Which THREE strategies improve the quality of retrieval in a Databricks Vector Search-based RAG application?
4What is the primary benefit of using Unity Catalog when designing generative AI applications in Databricks?
5When designing a production RAG application, which technique is most effective for preventing the LLM from hallucinating based on outdated information?
6Which Databricks feature is specifically designed to monitor model quality and drift in production?
7When designing an application that requires fine-tuning a small model (like Llama-3-8B) on Databricks, which THREE factors must be considered to ensure a successful training job?
8Which design pattern is best for protecting the LLM from prompt injection attacks when building a customer-facing chatbot on Databricks?
9Refer to the exhibit. An engineer is testing a model endpoint. The outputs are too brief and often stop mid-sentence. What is the most likely cause?
10When evaluating the performance of a RAG application, which metric is most useful for measuring the quality of the retrieved context?
11Which Databricks component is recommended for orchestrating the end-to-end RAG pipeline, including data ingestion, transformation, and vector indexing?
12An engineer needs to ensure that only authorized users can access the RAG chatbot. Which Databricks security feature should be used to enforce this access control?
13An application requires low-latency retrieval of RAG metadata stored in Databricks. Which storage approach balances performance and cost while ensuring seamless integration with Unity Catalog?
14Refer to the exhibit. The Databricks Vector Search index is currently in an 'ONLINE' status with incremental sync. If a developer performs a DELETE operation on the source Delta table, what happens to the vector index?
15You are building an application that uses Model Serving to host a fine-tuned LLM. Which configuration is required to optimize for high-concurrency request throughput?
16Refer to the exhibit. What is the correct way to log a custom RAG chain so it can be loaded using the provided code?
17When designing a RAG application for a regulated industry, what is the most important Unity Catalog feature for ensuring data compliance during retrieval?
18Which approach is recommended for managing secrets, such as API keys for external LLM providers, in Databricks?
19What is the primary function of the 'Prompt' object in a Databricks RAG architecture?
20Which TWO factors should be considered when choosing an embedding model for a RAG application?
21Refer to the exhibit. What is the cause of this error when logging a RAG chain to MLflow?
22In a RAG application, which architectural component ensures that the system handles changes in the source data effectively?
23A GenAI engineer is building a Databricks RAG application where the retrieval step returns the top-5 chunks for each user question. The engineer wants to add a second LLM call that evaluates whether each retrieved chunk contains enough information to answer the question, and then filters out chunks that fail this evaluation before passing the remaining chunks to the final answer-generation prompt. Which design pattern is the engineer implementing?
24A GenAI engineer is building a Databricks RAG application that answers questions over a Delta table containing 40 million support tickets. Users report that simple keyword lookups return irrelevant results because the tickets use inconsistent terminology. The engineer needs semantic retrieval that stays synchronized as new tickets stream in every few minutes. Which Databricks component should be used to serve this retrieval layer?
25A generative AI engineer is building a RAG chain on Databricks using LangChain. The chain must call an external LLM provider via a Databricks Model Serving endpoint that proxies the provider, and the engineer wants to avoid hardcoding credentials in notebooks. Which approach should the engineer use to authenticate calls from the chain to the serving endpoint?
26A GenAI engineer is building a customer-support assistant on Databricks. The assistant must invoke an external LLM endpoint that requires an API token, and the token must not appear in notebook source or logs. The engineer wants to store the token once and let notebooks and Databricks Jobs reference it by name. Which Databricks capability should the engineer use to satisfy this requirement?
27A team is designing a Databricks RAG application that must return answers with citations to source documents. They plan to use Databricks Vector Search and want the LLM to reference specific chunks. Which two design choices are required to produce reliable citations? (Choose two.)
28An engineer has a Databricks Vector Search index that stores chunk embeddings for a technical manual. Users report that queries containing exact part numbers return irrelevant chunks because the embedding model blurs numeric tokens. The engineer wants retrieval to consider both dense vector similarity and exact keyword matching on the same Delta table without building a separate search system. Which Databricks Vector Search feature should the engineer enable?
29A GenAI engineer is building a RAG chatbot backed by a Databricks Vector Search index that is synced from a Delta table with Change Data Feed enabled. Source documents are frequently updated, and some are deleted. The engineer wants the index to reflect deletions and updates automatically with minimal operational overhead. Which configuration should be used?
30An engineer is designing a RAG application where the LLM must answer questions using only the retrieved context and must refuse to answer when the context is insufficient. Which prompt design approach best enforces this behavior?
31A GenAI team at a retail company has built a RAG chatbot on Databricks that answers customer questions from a product catalog stored in Delta Lake. The catalog is updated nightly, and the team wants the chatbot to reflect those updates without manual intervention. They are deciding how to keep the Vector Search index synchronized. Which approach should they use?
32A team is deploying a RAG chatbot that answers questions from internal policy documents. The documents change frequently, and the team wants retrieval to reflect updates within minutes without re-running a full embedding job. The source Delta table already has change data feed enabled. Which approach should the team use to keep the Databricks Vector Search index current?
33A team is designing a GenAI application that must call an external LLM provider from a Databricks notebook and from a Model Serving endpoint. The security team requires that the provider API key never appear in notebook source code, Git history, or the model artifact, and that the same governed credential be reused by both the notebook and the serving endpoint. Which design satisfies these requirements?
34A generative AI engineer is designing an agent on Databricks that uses a LangChain agent with tool-calling capabilities. The agent must call a Databricks SQL warehouse to run queries and a Vector Search index for retrieval. Which design consideration is most important for controlling agent behavior in production?
35An engineer is designing a Databricks RAG application where the same retrieved context must be reused across several prompt variants during evaluation. They want to reduce token cost and improve consistency between variants. Which design choice best supports this?
36An engineer is designing a GenAI application where a Databricks workflow must call an external LLM provider that enforces a strict rate limit and occasionally returns transient errors. The engineer wants the application to degrade gracefully instead of failing the whole job. Which design choice best addresses this requirement?
37A team is deploying a Databricks RAG chatbot that must serve interactive traffic with low latency while also allowing the data science team to test prompt variations safely. The engineer wants the production endpoint to keep serving stable traffic even while a new prompt template is being evaluated. Which approach best meets these requirements?
38A data science team wants to expose a RAG chain as a REST API so that an external web application can send questions and receive answers. The chain is developed with Databricks LangChain integrations and must be deployed with autoscaling and built-in monitoring. Which Databricks capability should they use?
39A GenAI engineer is designing a retrieval-augmented generation application whose source documents are long PDFs. Early testing shows that answers are vague because retrieved chunks contain several unrelated topics, and the language model frequently cites content that does not support its claims. The engineer wants to improve chunk quality before indexing. Which TWO changes should the engineer make to the ingestion pipeline? (Choose two.)
40A GenAI engineer is designing a Databricks application that must ground answers in a large corpus of internal policy documents. The corpus is updated by a nightly Delta job, and the application must cite the source document for every answer. The engineer is deciding how to structure the retrieval and generation stages. Which TWO design choices best satisfy the grounding and citation requirements? (Choose two.)
41A financial services company is designing a RAG assistant that must never return answers containing personally identifiable information from its knowledge base. The team plans to filter retrieved chunks before they are placed into the prompt. Which Databricks design element should they use to enforce this filtering consistently?
42A generative AI team is building a customer-support assistant on Databricks. The assistant must answer questions using the company's private knowledge base, and the team wants to minimize latency while ensuring that the LLM only uses retrieved documents. They plan to use Databricks Vector Search with a Delta table as the source. Which design choice best balances low latency and grounded responses?
43An engineer is designing a Databricks GenAI application that summarizes customer meeting notes. The notes contain personally identifiable information, and company policy requires that the summarization model never sends note text to an external provider. The team already has a fine-tuned open model registered in Unity Catalog. Which deployment choice satisfies the policy?
44An engineer is designing a Databricks RAG application that must support multi-turn conversations where follow-up questions refer to earlier turns. They want the retrieval step to remain accurate as the conversation progresses. Which TWO design elements should they include? (Choose two.)
45An engineer is building a multi-turn GenAI assistant on Databricks. The assistant must answer follow-up questions that reference earlier turns, such as 'what about its warranty?', while keeping each request within the model's context limit. Which design should the engineer implement?
46A generative AI engineer is designing a RAG application on Databricks that uses a foundation model served via Databricks Model Serving. The application must handle peak loads gracefully and provide consistent response times. The engineer is evaluating design patterns for scaling and reliability. Which TWO design choices should the engineer implement? (Choose two.)
47An engineer is designing a RAG application that uses Databricks Vector Search to retrieve documents and a foundation model endpoint to generate answers. The team wants to log all user queries, retrieved documents, and generated responses for auditing and continuous improvement. They also need to monitor for drift in retrieval quality over time. Which Databricks capability should they integrate into the application design?
48A GenAI application uses a Databricks Model Serving endpoint hosting a foundation model, and the team wants structured JSON output that conforms to a fixed schema for downstream parsing. Responses sometimes include prose or markdown fences that break the parser. Which design change is most likely to produce reliably parseable output?
49A generative AI engineer is designing a multi-stage RAG application on Databricks. The application first retrieves documents using Vector Search, then reranks them with a cross-encoder model, and finally calls a foundation model endpoint to generate an answer. The engineer wants to ensure that the entire pipeline is reproducible and that each stage can be independently versioned and deployed. Which design approach best meets these requirements?
50A team is deploying a RAG application using Databricks Model Serving with a foundation model endpoint and a Databricks Vector Search index. During load testing, they observe that p95 latency spikes when the retriever returns many chunks, and the LLM occasionally truncates context. They want to reduce latency while preserving answer quality. Which change is most appropriate?
51An engineer is building a GenAI application that must return structured JSON output conforming to a specific schema so downstream systems can parse it reliably. The team wants to enforce the schema at generation time rather than post-processing free-form text. Which approach should they use?
52A GenAI engineer is designing an agentic application on Databricks that uses a foundation model to decide which external tools to call. The team wants the agent to be able to invoke a Databricks SQL warehouse query and a Python function registered as a Unity Catalog function, and they need the model to select tools based on natural language requests. Which design element is required for the model to select and invoke these tools correctly?
53An engineer is designing a GenAI application that must call an external LLM provider's API. The provider key must not appear in notebook code, job logs, or Git. The team already uses Databricks and wants the key to be injected into the serving endpoint at runtime. Which Databricks capability should the engineer use?
54A GenAI engineer is designing a RAG application on Databricks that must support multi-turn conversations where users refer to earlier messages, and the application must keep responses grounded in retrieved documents. Which TWO design elements are required to meet these requirements? (Choose two.)
55A GenAI engineer is designing a multi-turn chat application on Databricks. Users report that the assistant forgets details from earlier in long conversations and sometimes answers using only the most recent message. The team wants the model to reliably use facts stated several turns earlier without exceeding the model's context window. Which design should the engineer implement?
56A team is designing a Databricks GenAI application that must return grounded answers with citations to source documents. The application uses Databricks Vector Search for retrieval and a Foundation Model API for generation. Which TWO design choices are required to return accurate citations alongside each answer? (Choose two.)
57A generative AI engineer is designing a Databricks RAG application that ingests a Delta table containing 400 million support articles into a Databricks Vector Search index. The team wants the lowest-latency online serving with the smallest possible index while preserving retrieval quality for the most common queries. Which design decision best meets these requirements?
You must be able to choose the right Databricks building blocks for a RAG or GenAI app: vector search for retrieval, prompt objects for templating, secret scopes for credentials, and history-aware query construction for multi-turn chat. The single most important thing is matching each design need to the correct Databricks feature.
The Courseiva Databricks-GenAI-Assoc question bank contains 57 questions in the Design Applications domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Design Applications domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included