Databricks-GenAI-Assoc · domain
scenario questions
Practise Databricks Certified Generative AI Engineer Associate scenario questions practice questions — original exam-style scenarios with answer choices, explanations, and analysis of common mistakes.
Focused practice
Practice scenario questions questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about scenario questions
scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
How the topic appears in realistic exam-style scenarios.
Which detail in the question changes the correct answer.
How to eliminate plausible but wrong options.
How to connect the question back to the wider exam objective.
Watch out for
Common scenario questions exam traps
- ▸Answering from memory before reading the full scenario.
- ▸Missing a constraint such as cost, availability, security, scope or command context.
- ▸Choosing a broad answer when the question asks for the most specific fix.
- ▸Ignoring why the wrong options are tempting.
Question index
All scenario questions questions (330)
Click any question to see the full explanation, or start a practice session above.
A data scientist is building a GenAI application that uses the OpenAI API through Databricks external model endpoints. The application must use credentials stored securely in Databricks and must not expose the API key in code or logs. Which Databricks feature should the data scientist use to store and reference the OpenAI API key?
Medium2A team is designing a GenAI application that must call an external LLM provider from a Databricks notebook and from a Model Serving endpoint. The security team requires that the provider API key never appear in notebook source code, Git history, or the model artifact, and that the same governed credential be reused by both the notebook and the serving endpoint. Which design satisfies these requirements?
Easy3When designing a production RAG application, which technique is most effective for preventing the LLM from hallucinating based on outdated information?
Medium4An engineer has a Databricks Vector Search index that stores chunk embeddings for a technical manual. Users report that queries containing exact part numbers return irrelevant chunks because the embedding model blurs numeric tokens. The engineer wants retrieval to consider both dense vector similarity and exact keyword matching on the same Delta table without building a separate search system. Which Databricks Vector Search feature should the engineer enable?
Hard5A GenAI engineer is using MLflow to track experiments for a fine-tuned language model. The engineer wants to ensure that model artifacts and parameters are governed by Unity Catalog, so that access can be controlled and audited. Which Unity Catalog object should the engineer use to register the model?
Easy6Which component of a Databricks Asset Bundle (DAB) allows you to define different configurations (e.g., instance sizes, variables) for development versus production environments?
Medium7A generative AI team stores prompt/response logs in a Unity Catalog Delta table named `main.genai.interaction_logs`. Compliance requires that the raw `prompt_text` column be readable only by members of the `ai_compliance` group, while all other users querying the table must see the literal string `REDACTED`. Which Unity Catalog feature should be implemented directly on the table to satisfy this requirement without creating a separate view?
Medium8A RAG assistant built with Mosaic AI Agent Framework is in production. Product managers complain that answers are sometimes plausible but unsupported by the retrieved documents, and separately that the assistant sometimes ignores documents that clearly contain the answer. The team wants automated, scheduled checks in MLflow LLM Evaluation that separate these two failure modes so each can be triaged independently. Which TWO evaluation metrics should they add to the evaluation suite? (Choose two.)
Medium9When designing a RAG application for a regulated industry, what is the most important Unity Catalog feature for ensuring data compliance during retrieval?
Medium10A GenAI engineer has a Mosaic AI Agent application packaged as a Databricks Asset Bundle with a serving endpoint defined in the bundle's resources. A teammate recently updated the agent's prompt template in the source files, and the engineer now needs to push that change to the existing production endpoint without recreating it. Which Databricks CLI command should the engineer run from the bundle root?
Medium11A data engineering team is using Databricks to prepare data for a RAG application. They want to ensure that document chunks are of consistent size and quality. Which tool should they use within the Databricks notebook environment to achieve this?
Medium12A GenAI engineer registers a vector search index in Unity Catalog that points to a Delta table containing customer support transcripts. The security team requires that when an end-user queries the index through a Databricks notebook, the underlying table's row filter and column mask policies are enforced. Which Unity Catalog feature should the engineer configure?
Medium13A team deploys a customer-support assistant on a Mosaic AI Model Serving endpoint and enables inference table logging to Unity Catalog. Compliance requires that every production response be traceable back to the exact request, the retrieved context, and the model version that produced it, and that reviewers can query this history with SQL months later. Which capability satisfies this requirement?
Easy14A GenAI engineer monitors a customer-facing RAG assistant hosted on Databricks. After a routine re-indexing job, groundedness scores from the MLflow LLM Evaluation job drop sharply while answer relevance stays flat. The application prompt and the LLM serving endpoint were untouched. Which conclusion is best supported by these signals?
Hard15A data engineer is preparing a large corpus of support tickets stored in a Unity Catalog volume for fine-tuning a Llama model on Databricks. They must remove personally identifiable information (PII) before the data reaches the training cluster. Which TWO approaches are appropriate for detecting and redacting PII at scale in this pipeline? (Choose two.)
Medium16An engineer deployed a RAG agent to a Model Serving endpoint and enabled inference tables for payload logging. After a week, reviewers notice that logged requests contain customer email addresses and order identifiers. Compliance requires that raw prompts and responses not be stored in plain text. Which change should the engineer make?
Hard17An AI engineer is developing a real-time customer service chatbot application using Databricks Model Serving and needs to securely store API keys and database credentials without hardcoding them into the application code or notebook. Which approach should the engineer use?
Medium18An engineer built an agent using Mosaic AI Agent Framework and wants the agent to call a Unity Catalog function that returns customer order history. The function must be invoked by the LLM at runtime without exposing raw SQL to the model. Which approach should the engineer use?
Hard19When deploying a Python-based application that interacts with Unity Catalog, which step is essential to ensure the code can authenticate securely to external services without hardcoding tokens?
Easy20A generative AI engineer is building a RAG chain on Databricks using LangChain. The chain must call an external LLM provider via a Databricks Model Serving endpoint that proxies the provider, and the engineer wants to avoid hardcoding credentials in notebooks. Which approach should the engineer use to authenticate calls from the chain to the serving endpoint?
Medium21A media company's RAG assistant answers questions about its streaming catalog. Users report that the assistant often returns answers that ignore the retrieved documents and instead rely on outdated information from the base model. The team wants an MLflow LLM Evaluation metric that measures how much of the final answer is actually derived from the retrieved context rather than the model's prior knowledge. Which metric best fits this need?
Medium22You are building a data preparation pipeline for a generative AI application. The raw data is stored in a Unity Catalog volume as JSON files with nested fields. You need to flatten the nested structure and extract specific fields into a Delta table for downstream embedding. The JSON schema may evolve, with new fields added occasionally. Which approach provides the most robust and maintainable solution?
Hard23A company is deploying a GenAI chatbot that uses a Databricks Model Serving endpoint hosting a fine-tuned Llama 2 model. The endpoint is registered in Unity Catalog. The security team wants to restrict which groups can invoke the endpoint and also log all inference requests for auditing. Which combination of Unity Catalog and Databricks features should the engineer use to meet these requirements?
Hard24A media company uses an LLM-as-a-judge evaluation pipeline on Databricks to score a summarization assistant nightly. The judge is the same model family as the assistant and was given a rubric that rewards stylistic fluency. Over two months, nightly scores climb steadily while editor spot-checks find summaries increasingly omit key facts. Which corrective action best restores the evaluation's ability to detect this regression?
Hard25A team uses an LLM judge to score their RAG agent on Databricks and sees scores that fluctuate by several points between identical evaluation runs. They need more stable, reproducible quality signals for release gating. Which action best addresses the root cause?
Hard26A Generative AI engineer is preparing a Delta table of support tickets for embedding generation. The pipeline computes embeddings with ai_query and stores them in a separate embeddings table keyed by ticket_id. Tickets are frequently updated by agents, and the engineer wants the embeddings table to reflect the latest ticket text without recomputing embeddings for tickets whose text has not changed. Which design best achieves this?
Hard27A GenAI engineer is building a customer-support assistant on Databricks. The assistant must invoke an external LLM endpoint that requires an API token, and the token must not appear in notebook source or logs. The engineer wants to store the token once and let notebooks and Databricks Jobs reference it by name. Which Databricks capability should the engineer use to satisfy this requirement?
Medium28A GenAI engineer is preparing a large text corpus stored in a Unity Catalog volume for fine-tuning a chat model. The raw files are JSON Lines, each containing a 'conversation' array with alternating 'user' and 'assistant' turns, but many records have malformed turns or missing roles. The engineer needs to convert this into a Delta table with a schema of (conversation_id STRING, messages ARRAY<STRUCT<role:STRING, content:STRING>>) while filtering out records where any turn has a null role or empty content. Which approach uses the appropriate Databricks-native capability for this transformation?
Medium29Refer to the exhibit. A developer is registering a model. Why is the model signature, as shown in the exhibit, considered a best practice for model registration?
Hard30In a RAG application, which architectural component ensures that the system handles changes in the source data effectively?
Hard31A platform team is serving a retrieval-augmented generation application. The vector index is built from a Delta table in Unity Catalog, and the team wants every similarity search executed by an end user to be automatically restricted to documents that user is allowed to see, without maintaining a separate index per department. Which approach best enforces this at query time?
Hard32A team is deploying a fine-tuned LLM using Mosaic AI Model Serving. To reduce the cost of serving the model while maintaining acceptable performance, which strategy should be prioritized?
Medium33Which TWO factors should be considered when choosing an embedding model for a RAG application?
Medium34A Generative AI engineer is using MLflow Tracing to monitor a RAG application deployed on Mosaic AI Model Serving. They want to capture the retrieved documents, the final prompt, and the model's response for each request to debug a quality issue. Which approach should they use to ensure all three are logged in a single trace?
Hard35A developer is deploying a custom LangChain agent application as a custom Python model to Databricks Model Serving. The application depends on a specific third-party library that is not included in the standard Databricks runtime environment. How should the developer ensure the dependency is installed when the model is loaded into the serving container?
Medium36A financial services company has registered a Mosaic AI Agent application in Unity Catalog and wants to serve it through a Databricks Model Serving endpoint that queries a Delta table containing sensitive customer records. The security team requires that the endpoint access the table using a dedicated service principal with least privilege, and that the agent's LLM calls go through a governed gateway that logs usage. Which TWO configurations should the team apply to meet these requirements? (Choose two.)
Hard37A GenAI engineer is developing an agent using Databricks Mosaic AI Agent Framework. The agent must call an external API to fetch real-time stock prices. The engineer wants to ensure the agent can be evaluated and deployed with minimal changes between development and production. Which approach should the engineer take to define the tool-calling logic?
Hard38Which approach is most effective for managing the dependencies of a custom ML model when deploying it to Mosaic AI Model Serving?
Medium39A data science team has registered a new version of a retrieval-augmented generation (RAG) agent in Unity Catalog as a model named main.genai.support_agent. They need to make this version available as a low-latency REST endpoint with automatic scaling and no cluster management. Which Databricks capability should they use?
Easy40A team has deployed a RAG chatbot on Databricks and enabled inference table logging. They want to set up automated monitoring to detect when the average response length increases significantly compared to the baseline. Which Databricks feature should they use to create this monitor?
Medium41An engineer is designing a GenAI application where a Databricks workflow must call an external LLM provider that enforces a strict rate limit and occasionally returns transient errors. The engineer wants the application to degrade gracefully instead of failing the whole job. Which design choice best addresses this requirement?
Medium42An AI engineer is building a RAG application and wants to log the retrieval context and generated response for each request to MLflow for evaluation. They are using the `mlflow.langchain` flavor. Which method should they use to log the model so that MLflow automatically captures the necessary artifacts?
Medium43A data engineer wants to track who accessed a specific table in Unity Catalog. Which tool should they use?
Medium44A developer is configuring a model serving endpoint as shown in the exhibit. They observe that the endpoint fails to respond quickly to the first request after a period of inactivity. What is the cause of this behavior?
Medium45A team has developed a GenAI chatbot using the Databricks Mosaic AI Agent Framework and logged it with MLflow. They now want to expose it as a REST API that their customer support portal can call, with autoscaling and built-in monitoring. Which Databricks capability should they use to host the agent?
Easy46A team is developing a generative AI application that uses an LLM to answer questions based on internal documents. They want to ensure the application is robust and provides accurate responses. Which TWO practices should they implement to improve the reliability of the application? (Choose two.)
Medium47A team is deploying a Databricks RAG chatbot that must serve interactive traffic with low latency while also allowing the data science team to test prompt variations safely. The engineer wants the production endpoint to keep serving stable traffic even while a new prompt template is being evaluated. Which approach best meets these requirements?
Medium48When monitoring a RAG application, you notice a high discrepancy between the retrieved context and the generated answer. Which metric would specifically help identify if the model is ignoring the provided context?
Medium49A Databricks team is building a retrieval corpus from mixed-format documents stored in a Unity Catalog volume. They need a preparation pipeline that preserves document structure for later chunking and that records which source file each chunk came from so retrieval results can cite evidence. Which two design choices best meet these requirements? (Choose two.)
Hard50You are preparing a large text corpus for fine-tuning a generative AI model. The corpus is stored in a Delta table with columns: doc_id, raw_text, and metadata. You need to create a cleaned dataset that removes personally identifiable information (PII) and normalizes whitespace, while preserving document boundaries for training. Which two actions should you perform to achieve this in a scalable and maintainable way? (Choose two.)
Hard51Which THREE strategies improve the quality of retrieval in a Databricks Vector Search-based RAG application?
Medium52An organization wants to implement 'Guardrails' on their model outputs. Which deployment strategy best facilitates this?
Hard53A fintech company runs a customer-support RAG assistant on Databricks. Before promoting a new prompt template from staging to production, the ML team must demonstrate that the change does not regress answer quality. Which TWO evaluation practices should they apply in Mosaic AI Agent Evaluation to make this promotion decision defensible? (Choose two.)
Medium54You are preparing a large corpus of customer support emails stored as Parquet files in a Unity Catalog volume. The emails must be cleaned by removing boilerplate signatures and disclaimers before tokenization for a fine-tuning dataset. Your team wants to enforce this transformation as a declarative, testable pipeline stage that fails fast if any cleaned record still contains a known boilerplate marker. Which Databricks capability should you use to implement this cleaning step with built-in data quality expectations?
Medium55A team is preparing a pretraining corpus and must filter out near-duplicate documents before tokenization. The corpus contains 500 million short text records in a Delta table. They want a scalable, deterministic deduplication signal that can be computed per record and compared across the dataset. Which approach best fits this requirement?
Hard56A company is scaling its RAG applications. Which THREE of the following are benefits of using Databricks Vector Search over managing a standalone vector database outside of the platform?
Hard57A GenAI engineer has built a Retrieval Augmented Generation (RAG) application using Databricks Vector Search. The index is created on a Delta table that contains sensitive customer support transcripts. Company policy requires that end users can only retrieve chunks from documents they are authorized to view. The engineer wants to enforce this at query time without duplicating the index. Which approach should the engineer use?
Medium58Which Databricks component is recommended for orchestrating the end-to-end RAG pipeline, including data ingestion, transformation, and vector indexing?
Easy59A team is designing a Databricks GenAI application that must return grounded answers with citations to source documents. The application uses Databricks Vector Search for retrieval and a Foundation Model API for generation. Which TWO design choices are required to return accurate citations alongside each answer? (Choose two.)
Hard60Which of the following is a primary benefit of using Unity Catalog to manage models for a RAG application?
Medium61A team deploys a RAG agent to a Mosaic AI Model Serving endpoint. During load testing, requests intermittently return HTTP 429 responses even though the endpoint shows healthy replicas. The agent calls an external vector search index and a foundation model endpoint. Which action most directly addresses the 429 responses?
Hard62A data engineer is preparing a dataset of product descriptions for embedding generation. The source table in Unity Catalog contains a column 'description' with mixed languages, and the team wants to filter to English-only text before vectorization. They need a scalable, built-in Databricks function that can detect the language of each description without external API calls. Which function should they use?
Easy63An organization needs to build a RAG application on Databricks that minimizes data egress and maximizes security by keeping all data within the workspace perimeter. Which architectural pattern best satisfies this requirement?
Medium64A team runs a RAG application on Mosaic AI Model Serving and logs all requests to an inference table. Reviewers report that answers are sometimes fluent but contradict the retrieved documents. The team wants a recurring, automated check that quantifies this contradiction on production traffic and alerts when it exceeds a threshold. Which approach best fits?
Hard65Which TWO factors should be prioritized when selecting an embedding model for a domain-specific RAG application on Databricks?
Medium66A GenAI engineer is developing an agent on Databricks that must call a custom Python function to query an inventory database. The agent must decide when to invoke the function and must receive the result back into its reasoning loop. Which TWO actions are required to expose the function as a tool the agent can call? (Choose two.)
Hard67A team wants every MLflow evaluation run of their RAG chain to be reproducible months later, including the exact prompt template, model parameters, and retrieved context used. Which practice best ensures this reproducibility?
Easy68A team is deploying a RAG chatbot that answers questions from internal policy documents. The documents change frequently, and the team wants retrieval to reflect updates within minutes without re-running a full embedding job. The source Delta table already has change data feed enabled. Which approach should the team use to keep the Databricks Vector Search index current?
Easy69An organization is deploying a GenAI application using Databricks Model Serving. Which TWO steps are required to ensure the deployment environment handles model governance and observability effectively?
Hard70Refer to the exhibit. A developer wants to update this serving endpoint configuration to ensure it handles high-concurrency requests with consistent latency. Which change should be applied to the configuration?
Hard71A team deploys a Mosaic AI Agent application through Databricks Apps and must let the app call a Model Serving endpoint that enforces Unity Catalog permissions on the underlying model. Which TWO configurations are required for the app to authenticate and be authorized to query that endpoint? (Choose two.)
Hard72An engineer is designing a Databricks RAG application where the same retrieved context must be reused across several prompt variants during evaluation. They want to reduce token cost and improve consistency between variants. Which design choice best supports this?
Hard73Which THREE practices are recommended when using MLflow for managing LLM experiments in Databricks?
Medium74Refer to the exhibit. What is the most likely cause for this error in a production RAG application?
Hard75Refer to the exhibit. A developer wants to enable monitoring for their deployed LLM endpoint. Given the current configuration, what must the developer change to ensure that request/response logs are captured for analysis?
Medium76A generative AI engineer is designing a multi-stage RAG application on Databricks. The application first retrieves documents using Vector Search, then reranks them with a cross-encoder model, and finally calls a foundation model endpoint to generate an answer. The engineer wants to ensure that the entire pipeline is reproducible and that each stage can be independently versioned and deployed. Which design approach best meets these requirements?
Hard77A GenAI engineer is designing a multi-turn chat application on Databricks. Users report that the assistant forgets details from earlier in long conversations and sometimes answers using only the most recent message. The team wants the model to reliably use facts stated several turns earlier without exceeding the model's context window. Which design should the engineer implement?
Medium78A GenAI engineer is designing a Databricks application that must ground answers in a large corpus of internal policy documents. The corpus is updated by a nightly Delta job, and the application must cite the source document for every answer. The engineer is deciding how to structure the retrieval and generation stages. Which TWO design choices best satisfy the grounding and citation requirements? (Choose two.)
Hard79A developer is building an agentic workflow using Databricks and LangChain. The agent needs to decide whether to answer a user's query directly or call an external tool to retrieve additional information. The developer wants to ensure the agent's decisions are logged for debugging and auditing. Which approach should they take to achieve this?
Hard80A GenAI engineer is building a chatbot using Databricks Foundation Model APIs. They need the chatbot to maintain conversation context across multiple user turns and also allow the use of custom tools like a weather API. Which approach should they take?
Medium81A generative AI engineer trains a model on a Delta table that contains customer support transcripts. Before registering the model, security requires that the training data be classified so that policies can be applied consistently across the lakehouse. The engineer wants to attach a governed label to the transcript column indicating it contains sensitive personal data. Which Unity Catalog feature should be used?
Hard82A Generative AI engineer at a retail company has deployed a RAG chatbot on a Databricks Mosaic AI Model Serving endpoint. The team wants to capture per-request evaluation data, including the user's question, the retrieved context chunks, and the model's response, so they can analyze response quality over time. Which Databricks feature should they use to collect this data for downstream evaluation?
Easy83A team is preparing a Databricks Asset Bundle that deploys a GenAI application consisting of a job that builds a vector index and a Model Serving endpoint that hosts the agent. Before merging, they want the pipeline to validate and deploy the bundle to a staging workspace automatically. Which TWO bundle capabilities should the pipeline rely on? (Choose two.)
Medium84A team runs a RAG chatbot whose MLflow evaluation with the built-in groundedness judge previously scored well. After they swap the retriever for a new embedding model, groundedness scores drop sharply even though the generator model and prompt are unchanged. They confirm the judge model itself is unchanged. Which action should they take FIRST to diagnose the regression?
Medium85A team has deployed a model that is experiencing high latency. How should they identify if the bottleneck is the model inference or the preprocessing code?
Medium86A GenAI engineer is building a RAG chatbot backed by a Databricks Vector Search index that is synced from a Delta table with Change Data Feed enabled. Source documents are frequently updated, and some are deleted. The engineer wants the index to reflect deletions and updates automatically with minimal operational overhead. Which configuration should be used?
Medium87Which Databricks feature is best suited for maintaining the lineage of data used during the preparation of training sets for Generative AI?
Easy88A GenAI engineer is deploying a chain that calls an external LLM API. The chain must not block the serving thread while waiting on the remote API, and the endpoint must handle many concurrent requests. Which implementation approach should the engineer choose when logging the model?
Medium89A GenAI engineer is building a retrieval-augmented generation (RAG) application using Databricks Vector Search. During testing, they observe that for some queries, the retrieval step returns document chunks that are semantically similar but not actually relevant to the user's question, leading to poor answer quality. They want to improve retrieval precision without changing the embedding model. Which of the following approaches is most appropriate?
Medium90An engineer is troubleshooting a Vector Search index that fails to update. The index relies on a Delta table that is frequently updated. What is the most likely cause for the index failing to reflect new data?
Hard91You need to store embeddings generated by an LLM in a Delta table. Which data type is most efficient for storing these high-dimensional vector arrays in Databricks?
Hard92When evaluating an LLM for deployment, what is a crucial 'non-functional' requirement that must be addressed?
Medium93A GenAI engineer is preparing a corpus of HTML product pages stored in a Unity Catalog volume for a RAG application. The pages contain navigation bars, script tags, and boilerplate footers that add noise to embeddings. Which Databricks-native approach best removes this noise while preserving the main article text before chunking?
Medium94An organization is deploying an LLM application using Databricks Asset Bundles (DABs). Which TWO of the following are primary benefits of using DABs for the deployment process?
Medium95Refer to the exhibit. What is the impact of min_instances: 0 on this deployment?
Medium96A team is building a RAG application using Databricks Vector Search with a Delta table as the source. They need the index to automatically reflect new and updated chunks as the source table changes, without rebuilding the entire index each time. Which configuration should they use?
Hard97A team uses MLflow LLM Evaluation with an LLM judge to score a summarization agent. On a 400-row golden dataset the judge marks 96% of summaries as relevant. Spot-checking reveals the judge approves nearly every summary whenever the summary is fluent, even when key facts are missing. The team wants a defensible quality signal before approving a release. Which action best addresses this judge weakness?
Hard98You are deploying a RAG application. You need to ensure the model uses the most recent vector data without redeploying the model. What should you use?
Medium99Which of the following metrics is most effective for evaluating a RAG-based chatbot's ability to retrieve relevant context from a vector database during production monitoring?
Medium100Which security configuration is essential when deploying a model that accesses sensitive data stored in Unity Catalog?
Medium101A GenAI engineer is building a retrieval-augmented chatbot whose answers must cite the exact source document and page. The team wants the chatbot's responses to include structured citations that downstream UIs can render. Which Databricks feature should the engineer use to return these structured citations from the model endpoint?
Hard102A team needs to provide fine-grained access control to an LLM application that uses a vector index stored in a Delta table. Which Unity Catalog feature should they use to restrict access to specific rows based on user identity?
Medium103A company wants to share a GenAI application's prompt templates and evaluation datasets with a partner organization. The partner uses its own Databricks account and must not access any other company data. The company also wants to revoke access easily. Which Unity Catalog feature should they use?
Hard104A data engineer wants to share a specific subset of sensitive PII data with an external department using Unity Catalog. The engineering team must ensure the data is anonymized dynamically based on the user's role without creating physical copies. Which feature is most appropriate?
Medium105A Generative AI engineer is setting up MLflow LLM Evaluation for a RAG pipeline on Databricks that answers questions about product manuals. The team wants to evaluate both the retrieval stage and the generation stage in a single evaluation run. Which TWO metric groups should they include in the evaluation configuration? (Choose two.)
Hard106A GenAI engineer prepares a Delta table of product descriptions that will feed a chunking and embedding pipeline. The descriptions are written in mixed languages, and the embedding model supports only English. The engineer needs to ensure that non-English rows are detected and routed for translation before embedding. Which approach is most appropriate?
Hard107A team is building a retrieval-augmented generation application on Databricks and wants to reduce hallucination by improving the quality of retrieved context before it reaches the LLM. Which TWO techniques should they apply? (Choose two.)
Medium108A team runs an LLM-as-a-judge evaluation on Databricks using the built-in `mlflow.evaluate()` with `model_type="databricks-agent"`. They notice that the judge model, a serving endpoint, is producing scores that are consistently inflated compared to human review. Which configuration change should the team make first to improve the reliability of the evaluation?
Medium109A GenAI engineer is building a chatbot using Databricks Model Serving and needs to ensure that the endpoint can only be invoked by users who have been granted the 'genai_users' group. The endpoint is registered in Unity Catalog. Which configuration is required to enforce this access control?
Hard110A team is deploying a RAG application using Databricks Model Serving with a foundation model endpoint and a Databricks Vector Search index. During load testing, they observe that p95 latency spikes when the retriever returns many chunks, and the LLM occasionally truncates context. They want to reduce latency while preserving answer quality. Which change is most appropriate?
Hard111A developer is using MLflow to track experiments for a generative AI application. They want to log a prompt template and its associated parameters so that they can reproduce the exact input to the model later. Which MLflow function should they use to log the prompt template as an artifact?
Easy112When deploying a Generative AI application, why is it recommended to use a dedicated Serving Endpoint rather than a shared interactive cluster?
Medium113What is the primary benefit of using Unity Catalog when designing generative AI applications in Databricks?
Easy114A logistics company monitors a Databricks-hosted RAG assistant that answers questions about shipping regulations. Over one week, the retrieval index was rebuilt after a document refresh, and the team observes that the rate of answers flagged as unsupported by retrieved context rose sharply while retrieval relevance scores stayed flat. Which metric should they inspect first to determine whether the regression originates in the retrieval stage or the generation stage?
Easy115Which Databricks feature is primary for managing the lifecycle, versioning, and deployment readiness of custom Generative AI models?
Easy116A developer is building a RAG application. Which step is essential to prevent the model from hallucinating or providing outdated information?
Medium117An enterprise LLM application on Databricks is experiencing high latency and inconsistent responses. Which approach best enables observability to identify the root cause of these performance bottlenecks within the LLM pipeline?
Medium118A data scientist has developed a GenAI application locally and now needs to share it with stakeholders as an interactive web app inside the Databricks workspace. The app must use the workspace's existing authentication and compute. Which Databricks capability should be used?
Easy119When developing a Generative AI application in Databricks, which tool provides a collaborative environment for engineers to write code, visualize data, and document their experiments using Markdown?
Easy120An engineer is building a GenAI application that must return structured JSON output conforming to a specific schema so downstream systems can parse it reliably. The team wants to enforce the schema at generation time rather than post-processing free-form text. Which approach should they use?
Easy121Which action must an administrator perform to allow a user to use Databricks SQL to query a table that is stored in an external storage location?
Medium122A team wants its Databricks App to read a secret that stores an external API key used by an agent tool. The secret is managed in a Databricks secret scope. Which approach correctly exposes the secret to the running app without putting the value in source control?
Medium123A financial services team runs a RAG assistant on Databricks that answers questions about internal policy documents. During a review, the team finds that for many questions the answer is factually correct but cites a document that does not actually contain the supporting statement. They want an MLflow LLM Evaluation metric that specifically detects when the response is not supported by the retrieved context. Which metric should they add to their evaluation run?
Medium124A data team is developing a Mosaic AI Agent application. Which TWO of the following are mandatory for deploying this application using the Databricks Model Serving infrastructure?
Medium125Which of the following describes the 'Gold' layer in a Medallion architecture, and why is it important for GenAI data preparation?
Easy126A machine learning engineer needs to deploy a custom Mosaic AI Model Serving endpoint that requires access to a private internal database. Which mechanism should be used to securely manage the database credentials?
Medium127A platform team is preparing to deploy a GenAI chat application built on Mosaic AI Model Serving. They want the deployment to enforce least-privilege access for the application and to keep the serving environment reproducible. Which two practices should they implement? (Choose two.)
Medium128A GenAI team is using MLflow to track experiments for a large language model. They want to ensure that only team members can view the experiment results and that the experiment artifacts are stored in a governed location. Which Unity Catalog integration should they use to manage MLflow experiments?
Easy129A GenAI engineer is developing a conversational agent using Databricks. The agent must maintain context across multiple turns and retrieve relevant information from a knowledge base. They want to ensure the agent can handle follow-up questions that refer to previous exchanges. Which TWO techniques should they implement to manage conversation state and retrieval effectively? (Choose two.)
Medium130A Mosaic AI Agent application deployed as a Databricks App intermittently returns stale answers after the team updates the underlying vector index. The app caches a client to the serving endpoint and an index handle at module import time. Which change best resolves the staleness while keeping latency low?
Hard131A Generative AI engineer is configuring a Mosaic AI Model Serving endpoint for a production-grade LLM. Which TWO of the following tasks are necessary to ensure effective monitoring and evaluation of the endpoint? (Choose two)
Hard132A data engineering team is deploying a RAG application using Mosaic AI Model Serving. They need to monitor the quality of the model's responses in production. Which Databricks feature should they use to capture and analyze inference data, such as requests, responses, and latency metrics?
Medium133A team is deploying a GenAI application to a Databricks Model Serving endpoint and wants to ensure the deployment can be rolled back quickly and that traffic is shifted safely during updates. Which TWO practices should the team follow? (Choose two.)
Hard134A financial services company is designing a RAG assistant that must never return answers containing personally identifiable information from its knowledge base. The team plans to filter retrieved chunks before they are placed into the prompt. Which Databricks design element should they use to enforce this filtering consistently?
Medium135An organization requires that all GenAI models deployed in Databricks be tracked and managed with a unified registry for compliance. Which feature should the developer use?
Medium136A GenAI engineer must build a fine-tuning dataset from 40 TB of raw JSONL conversation logs stored in cloud object storage. The logs are immutable and only ever read once, in full, during preprocessing. Which storage and access configuration should be used to minimize cost while keeping the data readable by Spark on Databricks?
Medium137A team is designing an LLM application that requires strict data privacy. Which TWO approaches ensure that sensitive data is not leaked during the model inference process or training?
Hard138When integrating an external LLM via a Databricks Model Serving endpoint, how should the API credentials be managed to ensure they are not exposed in the application code?
Medium139Refer to the exhibit. An engineer is automating the deployment of a model to an endpoint using the Databricks CLI. Based on the error log provided, what is the most appropriate action to resolve this deployment failure?
Hard140A developer is creating a Databricks notebook to orchestrate a GenAI pipeline that includes data ingestion, vector index refresh, and model inference. They want to ensure that the pipeline can be easily tested and deployed across different environments. Which Databricks feature should they use to define the pipeline as code and manage deployments?
Medium141A developer is using Databricks AI Functions to extract structured information from a large set of customer reviews stored in a Delta table. They want to apply a prompt to each review and store the results in a new column. Which function should they use?
Medium142When designing an agentic workflow in Databricks, which TWO practices are essential to ensure the application remains observable and maintainable?
Hard143A team's RAG evaluation shows high context recall but low answer correctness. Retrieved documents contain the needed facts, yet the generated answers frequently contradict them. Which single metric should they examine next to pinpoint whether the generator is ignoring or misusing the retrieved context?
Medium144A team deploys a Mosaic AI Agent application to a Databricks Model Serving endpoint. During load testing they observe that the first request after an idle period takes several seconds, while subsequent requests are fast. They want to eliminate this cold-start penalty for a latency-sensitive customer-facing application while keeping costs reasonable during off-peak hours. Which configuration should they apply?
Hard145A generative AI engineer is packaging a retrieval-augmented generation (RAG) application so it can be deployed as a Databricks App. The app reads the vector index name and the serving endpoint name from environment configuration so the same code can run in dev and prod. Which approach correctly supplies these values at deploy time using Databricks Asset Bundles?
Medium146Which TWO of the following are primary components of the Unity Catalog identity model?
Hard147Which workflow best describes the recommended CI/CD process for updating a Databricks Asset Bundle?
Medium148A GenAI platform team is preparing a Unity Catalog schema to host a retrieval-augmented generation pipeline. They will store prompt templates and evaluation datasets as Delta tables, and they need to expose the pipeline to a group of application developers. The security team asks the platform team to describe how Unity Catalog privileges must be granted for the developers to query the tables. Which TWO statements correctly describe the required privilege model? (Choose two.)
Medium149Refer to the exhibit. An engineer wants to perform a canary deployment by routing 10% of traffic to a new version (version 6). How should the JSON traffic configuration be modified?
Hard150A GenAI engineering team has deployed a customer-support RAG chain on Databricks and registered it in Unity Catalog. They now want MLflow 3 to automatically score every production request for groundedness and relevance without writing custom scoring code, and to persist those assessments against the logged traces. Which approach should they use?
Medium151Which THREE conditions must be met for a user to successfully create a new table in a Unity Catalog schema?
Hard152An engineer is packaging a custom PyFunc model that wraps an open-source LLM and must deploy it to a Databricks Model Serving endpoint. The endpoint must load the model from Unity Catalog and expose it through a REST API. Which two actions are required to make the deployment succeed? (Choose two.)
Medium153A Generative AI engineer is packaging a retrieval-augmented generation application for deployment with Databricks Asset Bundles. The bundle must provision a Databricks job that periodically refreshes a Delta table used as the vector index, and the job requires a specific cluster node type that differs between the development and production workspaces. Which approach correctly handles the node type difference while keeping a single bundle definition?
Medium154A Databricks Generative AI engineer has deployed a RAG application and is now setting up production monitoring. They want to automatically detect when the distribution of incoming user questions diverges from the distribution seen during development, so they can trigger retraining or prompt adjustments. Which Databricks capability should they configure?
Medium155An AI engineer is deploying a RAG application using Databricks Model Serving. They need to ensure the endpoint can handle high traffic with low latency and automatically scale based on demand. Which configuration should they use?
Hard156Which Unity Catalog object should be used to manage access to a folder of JSON files that are not part of a formal Delta table?
Medium157A team is deploying a GenAI agent built with Mosaic AI Agent Framework. The agent must be exposed as a REST API that automatically scales with traffic and records inference logs to Unity Catalog. Which deployment mechanism meets these requirements with the least operational overhead?
Medium158A Generative AI engineer is designing an evaluation harness for a customer-support RAG agent on Databricks. They need metrics that specifically assess the RETRIEVAL stage rather than the generation stage. (Choose two.)
Hard159Which tool in the Databricks ecosystem is best suited for developers to experiment with prompt engineering and tool-calling logic iteratively?
Easy160A healthcare analytics team has a RAG application that must not reveal protected health information from other patients. They want to continuously monitor production traffic on their Mosaic AI Model Serving endpoint and alert when responses contain unsafe content. Which Databricks capability should they configure to evaluate each logged request and response against safety criteria and route flagged records for review?
Hard161A team is using `databricks.agents.deploy()` to publish a Mosaic AI Agent to a serving endpoint. They must expose an environment-specific Vector Search index name and the endpoint name to the deployment without hardcoding values in the notebook, and the same notebook must run in dev and prod. Which approach should the engineer use?
Hard162Which of the following describes the 'drift' phenomenon in the context of LLM monitoring?
Medium163A GenAI engineer has registered a fine-tuned Llama 3 model in Unity Catalog as `prod.ml.models.support_llm` and now needs to expose it as a REST endpoint for a customer-facing chatbot. The team wants Databricks to manage the serving infrastructure, GPU autoscaling, and version upgrades with minimal operational overhead. Which action should the engineer take?
Easy164A team maintains a Mosaic AI Agent application whose endpoint must call an external LLM provider through a secret stored in a Databricks secret scope. During a deployment pipeline run, the endpoint build step fails while resolving the credential, even though the secret scope exists and the notebook test works. Which configuration should the engineer verify first?
Hard165Which design pattern is best for protecting the LLM from prompt injection attacks when building a customer-facing chatbot on Databricks?
Medium166Which component of the Databricks platform allows you to bundle together notebooks, model serving configurations, and pipeline definitions for repeatable deployment?
Easy167A healthcare analytics team deployed a RAG assistant on Databricks whose answers cite clinical policy documents. Compliance requires that every production answer be attributable to a specific retrieved chunk. During evaluation, the team notices the groundedness judge scores are high, but manual review finds answers that blend two policies into a statement neither document supports. Which change to their Mosaic AI Agent Evaluation configuration best detects this failure mode?
Hard168A GenAI team at a retail bank runs a RAG assistant on a Databricks Mosaic AI Model Serving endpoint. During a pilot, they captured end-user thumbs-up/down feedback in a Delta table but never joined it to the trace payloads. Six weeks later, hallucination complaints spike, yet the aggregate thumbs-down rate is unchanged. Which approach best resolves this discrepancy using Databricks-native tooling?
Medium169An AI engineer is using MLflow to track experiments for a generative AI application. They want to log parameters, metrics, and artifacts for each run, and later compare runs to select the best model. Which MLflow component should they use to organize runs into a named group for a specific project?
Easy170A GenAI engineer is preparing a large text corpus for fine-tuning an LLM. The corpus contains many near-duplicate documents and documents in multiple languages. They need to reduce redundancy and ensure language consistency. Which two steps should be performed during data preparation? (Choose two.)
Hard171Which THREE of the following are common challenges when monitoring LLM applications in production that differ significantly from traditional ML monitoring? (Choose three)
Medium172Which Databricks feature allows developers to track experiment parameters, model artifacts, and evaluation metrics in a structured way?
Easy173A GenAI engineer is developing a RAG application that uses a Vector Search index. The index is created from a Delta table that contains sensitive information. The security team requires that the engineer can audit which users have queried the index and what data they retrieved. Which Unity Catalog feature should the engineer enable to capture this information?
Medium174A generative AI engineer registers a model in Unity Catalog and wants a downstream application to call it for inference without granting the application broad workspace access. The engineer is told to grant a specific Unity Catalog privilege on the registered model so the application can invoke it. Which privilege should be granted?
Easy175A GenAI engineer is instrumenting a production RAG application on Databricks to detect quality degradation before users complain. The team wants signals that reveal problems in the retrieval stage specifically, rather than in the generation stage. Which TWO signals should they track? (Choose two.)
Medium176A machine learning engineer needs to deploy a real-time Model Serving endpoint. Which Databricks construct is required to manage the model's environment, dependencies, and artifacts while ensuring version control for the deployment?
Medium177Why should you integrate Mosaic AI Model Evaluation with Unity Catalog?
Easy178A team is using MLflow to track their model evaluation runs. They want to ensure that every evaluation run is reproducible. What is the best practice to achieve this in Databricks?
Hard179Which Databricks asset is best suited for scheduling and orchestrating a multi-step GenAI pipeline that includes data ingestion, vector index updating, and model evaluation?
Easy180A team is designing a Databricks RAG application that must return answers with citations to source documents. They plan to use Databricks Vector Search and want the LLM to reference specific chunks. Which two design choices are required to produce reliable citations? (Choose two.)
Hard181A team wants to review what a Databricks Asset Bundle deployment would change in the production workspace before it actually creates or modifies any resources. Which command should the engineer run first?
Easy182Refer to the exhibit. The Databricks Vector Search index is currently in an 'ONLINE' status with incremental sync. If a developer performs a DELETE operation on the source Delta table, what happens to the vector index?
Medium183When migrating a Databricks Asset Bundle (DAB) project from 'development' to 'production', which TWO actions should an engineer perform?
Medium184Which TWO of the following are benefits of using Delta Lake for data preparation over standard Parquet files on cloud storage?
Medium185A developer is creating a custom model serving endpoint that requires an external API call for data enrichment. What is the recommended way to handle sensitive API keys within the Databricks environment?
Medium186An engineer is designing a GenAI application that must call an external LLM provider's API. The provider key must not appear in notebook code, job logs, or Git. The team already uses Databricks and wants the key to be injected into the serving endpoint at runtime. Which Databricks capability should the engineer use?
Easy187An engineer is using Mosaic AI Agent Evaluation to score a conversational agent that calls tools. The agent sometimes answers correctly but with fabricated citations. Which evaluation approach best surfaces this specific failure mode?
Hard188When building an application that retrieves context from Databricks Vector Search, what is the recommended data format for storing the document chunks?
Easy189When deploying a model endpoint, which Databricks feature provides the capability to review and approve the model before it is promoted to production?
Medium190A developer needs to monitor the performance of an LLM application in production. They want to track the latency of their model endpoint. Where can they find this metric in the Databricks workspace?
Medium191When building a RAG application, a developer wants to ensure that the retrieved context is strictly limited to documents the user has access to. Where should this security logic be enforced?
Medium192An AI engineer is designing a scalable customer support application on Databricks that integrates custom vector search indexes with a fine-tuned LLM. Which TWO architectural components are essential for enabling efficient similarity search and low-latency retrieval within the Databricks ecosystem? (Choose TWO)
Medium193Which Databricks feature is specifically designed to monitor model quality and drift in production?
Medium194A GenAI engineer has built a RAG application that queries a Delta table named `prod_ai.knowledge_base.documents` through a Databricks SQL warehouse. The application uses a service principal `sp-rag-prod` to authenticate. The table contains confidential internal documents. The security team wants to ensure that the service principal can only read the table and cannot modify or delete it. Which Unity Catalog privilege should be granted to `sp-rag-prod` on the table?
Medium195A data scientist wants to use a Unity Catalog registered model in a GenAI pipeline. The model was trained on sensitive data, and the governance team requires that the model's lineage back to the training dataset be traceable. Which Unity Catalog capability provides this traceability?
Easy196What is the purpose of the 'metastores.list' command in the Unity Catalog CLI?
Hard197A team is preparing data for a RAG system and needs to remove duplicates from a large collection of PDF text extracts. What is the most efficient way to perform de-duplication in Databricks?
Medium198A GenAI engineer is preparing a fine-tuning dataset from a Delta table in Unity Catalog that contains raw user feedback. The feedback text includes irregular capitalization, HTML tags, and excessive punctuation. The engineer needs to normalize the text using Spark NLP within a Databricks notebook, ensuring the pipeline is reproducible and scalable. Which approach should the engineer use to apply this transformation?
Medium199A data science team wants to expose a RAG chain as a REST API so that an external web application can send questions and receive answers. The chain is developed with Databricks LangChain integrations and must be deployed with autoscaling and built-in monitoring. Which Databricks capability should they use?
Easy200A data engineer is preparing a dataset for fine-tuning a chat model. The dataset contains conversations with alternating user and assistant messages. They need to format each conversation into a single string with special tokens indicating roles. Which approach is most appropriate in Databricks?
Medium201You need to ingest data from an external JSON source into a Delta table. The source schema is inconsistent. Which strategy is most effective for preparing this data?
Medium202A team is deploying a GenAI application using Mosaic AI Model Serving. They want to ensure that the endpoint can handle sudden spikes in traffic without dropping requests. Which feature should they configure?
Medium203A generative AI engineer is designing a RAG application on Databricks that uses a foundation model served via Databricks Model Serving. The application must handle peak loads gracefully and provide consistent response times. The engineer is evaluating design patterns for scaling and reliability. Which TWO design choices should the engineer implement? (Choose two.)
Hard204A data engineer needs to persist a prepared instruction-tuning dataset so that downstream fine-tuning jobs can read it with ACID guarantees, time travel, and schema enforcement, and so that Unity Catalog can track column-level lineage. Which storage format and registration should be used?
Easy205A developer is packaging a GenAI chat application as an MLflow model that will be deployed to a Mosaic AI Model Serving endpoint. The application needs to load a retrieval index and a prompt template at startup so the first request is not slowed by initialization. Which MLflow logging pattern should the developer use?
Easy206Refer to the exhibit. What is the correct way to log a custom RAG chain so it can be loaded using the provided code?
Medium207When deploying a model to Databricks Model Serving, which THREE of the following are best practices to ensure production reliability?
Medium208Which Unity Catalog object is used to link a specific cloud storage path to a catalog, schema, or table, allowing users to create tables without managing individual storage credentials?
Easy209A data science team wants to expose a fine-tuned LLM as a REST API for an internal web application. They have registered the model in Unity Catalog and want Databricks to manage the serving infrastructure, autoscaling, and request routing. Which Databricks capability should they use?
Easy210A generative AI engineer is designing a Databricks RAG application that ingests a Delta table containing 400 million support articles into a Databricks Vector Search index. The team wants the lowest-latency online serving with the smallest possible index while preserving retrieval quality for the most common queries. Which design decision best meets these requirements?
Medium211A team is deploying a Retrieval-Augmented Generation application as a Mosaic AI Agent using `databricks.agents.deploy()`. The agent must query a Vector Search index and a Delta table in Unity Catalog. During testing, the endpoint returns permission errors when accessing those resources, even though the deploying user has access. Which configuration should the engineer apply to resolve this?
Medium212A team is evaluating their RAG application using Mosaic AI Model Evaluation. Which TWO metrics are most relevant for assessing the quality of the generated responses?
Medium213A GenAI engineer notices that a RAG agent's retrieval stage is returning relevant chunks, but the generated answers frequently omit key facts present in those chunks. The team wants a single evaluation metric that isolates whether the generator is using the provided context. Which metric should they focus on?
Easy214A GenAI team's Mosaic AI Agent application is deployed via a Databricks Asset Bundle and served through a Model Serving endpoint. They need the endpoint to call a Unity Catalog function as a tool during inference, and the function reads from a table the endpoint's service principal cannot currently access. What should the engineer do?
Hard215What is the primary function of a Storage Credential in Unity Catalog?
Easy216A GenAI engineer is building a Databricks RAG application where the retrieval step returns the top-5 chunks for each user question. The engineer wants to add a second LLM call that evaluates whether each retrieved chunk contains enough information to answer the question, and then filters out chunks that fail this evaluation before passing the remaining chunks to the final answer-generation prompt. Which design pattern is the engineer implementing?
Hard217A GenAI application uses a Databricks Model Serving endpoint hosting a foundation model, and the team wants structured JSON output that conforms to a fixed schema for downstream parsing. Responses sometimes include prose or markdown fences that break the parser. Which design change is most likely to produce reliably parseable output?
Hard218When monitoring a production LLM, you detect a drift where the model's responses are becoming increasingly verbose and less helpful compared to the baseline. Which strategy is most effective for detecting this quality decay?
Medium219A GenAI engineer is building a retrieval-augmented generation application on Databricks. They want to store document embeddings and perform fast approximate nearest-neighbor search without managing a separate vector database. They have already created a source Delta table with columns: id (string), text (string), and embedding (array<float>). Which Databricks feature should they use to create a Vector Search index that automatically syncs with the Delta table?
Medium220Refer to the exhibit. An AI engineer is configuring a Databricks Asset Bundle (DAB) to deploy a generative AI application. When executing 'databricks bundle deploy --target prod', which workspace host will the bundle resources be deployed to, and why?
Hard221Which TWO evaluation approaches are most effective for measuring the quality of a RAG pipeline's retrieval stage?
Medium222An engineer needs to ensure that only authorized users can access the RAG chatbot. Which Databricks security feature should be used to enforce this access control?
Medium223A data scientist is developing a GenAI application that uses a foundation model served via Databricks Model Serving. The model endpoint is configured to log inference tables for monitoring. The data science team wants to ensure that the inference logs, which may contain sensitive user prompts, are protected according to Unity Catalog policies. Which Unity Catalog object should be used to store and govern the inference tables?
Easy224A healthcare company uses Databricks to build a GenAI chatbot that answers questions from patient records stored in a Delta table. The records contain PHI, and the company must ensure that the chatbot never returns PHI to unauthorized users. The security team wants to enforce policies at the data layer so that even if the LLM is manipulated, it cannot access PHI. Which Unity Catalog feature should be used to dynamically redact PHI columns based on the user's group membership?
Hard225When designing an application that requires fine-tuning a small model (like Llama-3-8B) on Databricks, which THREE factors must be considered to ensure a successful training job?
Hard226Refer to the exhibit. An engineer is configuring a canary deployment for a churn prediction model. Based on the provided traffic configuration, what is the expected behavior of the endpoint?
Hard227Which command is used within the Databricks CLI to deploy a project defined by a Databricks Asset Bundle?
Medium228A retail company runs a customer-facing RAG assistant on a Mosaic AI Model Serving endpoint. The team wants every production request, response, and retrieved context to be captured automatically into a Unity Catalog Delta table so they can monitor quality and latency trends over time. Which action should they take?
Hard229Your organization requires an auditable record of all LLM evaluation results for compliance. Which Databricks feature provides the best centralized storage for these evaluation runs?
Medium230A team is preparing a fine-tuning dataset from customer reviews stored in a Delta table. They need to filter out reviews shorter than 20 tokens and reviews flagged as spam by a classifier, then write the result to a Unity Catalog table for training. Which approach best fits Databricks best practices?
Medium231A developer needs to store prompt templates, model parameters, and evaluation results for a GenAI application so that each iteration can be compared and reproduced later. Which Databricks capability should they use?
Easy232A Generative AI engineer is preparing a Delta table of product reviews for a retrieval-augmented generation application. The reviews contain HTML tags, inconsistent casing, and occasional very long paragraphs that exceed the embedding model's context window. The engineer wants to clean and normalize the text before chunking and embedding. Which two actions should the engineer take to directly address the stated quality issues? (Choose two.)
Medium233A GenAI engineer is building a Databricks RAG application that answers questions over a Delta table containing 40 million support tickets. Users report that simple keyword lookups return irrelevant results because the tickets use inconsistent terminology. The engineer needs semantic retrieval that stays synchronized as new tickets stream in every few minutes. Which Databricks component should be used to serve this retrieval layer?
Medium234When preparing data for a fine-tuning task, you realize the dataset is severely imbalanced. Which Databricks technique should you use to create a more balanced dataset?
Medium235You are preparing a large dataset for fine-tuning a model using Databricks Delta Live Tables (DLT). Which configuration is best for ensuring data quality and lineage in this pipeline?
Medium236A GenAI engineer is building a RAG application on Databricks. They have registered a foundation model endpoint in Unity Catalog as a model. The application needs to query the endpoint, and the engineer wants to ensure that only members of the group 'genai_team' can invoke it. Which Unity Catalog privilege must be granted on the model object to allow invocation?
Medium237A generative AI engineer is designing an agent on Databricks that uses a LangChain agent with tool-calling capabilities. The agent must call a Databricks SQL warehouse to run queries and a Vector Search index for retrieval. Which design consideration is most important for controlling agent behavior in production?
Hard238A Generative AI engineer is building a Delta Live Tables pipeline that ingests raw JSON event logs into a bronze table, then uses ai_query to classify each event's free-text field. The classification call is expensive, so the engineer wants to avoid re-running it on events that have already been processed in previous pipeline updates. The source table is append-only and new events arrive continuously. Which Delta Live Tables feature should the engineer configure on the bronze table to prevent reprocessing of previously ingested rows?
Medium239What is the primary benefit of using MLflow Model Evaluation for generative AI applications compared to manual evaluation methods?
Easy240Which component in the Databricks GenAI stack is responsible for orchestrating the flow between data retrieval, prompt construction, and model invocation?
Easy241A developer is building a RAG application using Mosaic AI Model Serving. They need to ensure that the model endpoint logs inference requests and responses for audit purposes. Which configuration parameter should they enable?
Medium242Which Unity Catalog feature is best suited for sharing data assets with users outside of your Databricks account?
Medium243When evaluating LLM outputs, which TWO metrics are most appropriate for measuring the 'quality' of a response in a RAG system? (Choose two)
Medium244A developer needs to deploy a custom Python model that requires non-standard library dependencies. Which MLflow feature should the developer use to specify these environment requirements during model logging?
Medium245A data scientist is using Databricks to fine-tune a large language model. The training data is stored in a Unity Catalog table that contains sensitive customer information. The data scientist needs to read the table but should not be able to see the raw values of certain columns. Which Unity Catalog feature should be used to dynamically mask the sensitive columns based on the user's group membership?
Medium246A data engineer is preparing a large Delta table of conversation logs for embedding generation. The table has frequent small appends, and the engineer needs to reduce file fragmentation and improve read throughput before the embedding job runs. Which two actions should the engineer take? (Choose two.)
Medium247When evaluating the performance of a RAG application, which metric is most useful for measuring the quality of the retrieved context?
Medium248A Data Engineer needs to ensure that PII data in a Delta table is masked before serving it to non-privileged users. Which Databricks feature provides the most efficient, centralized control for this requirement?
Medium249A GenAI engineer has a RAG application whose retrieval step uses Databricks Vector Search. Users report that answers are sometimes irrelevant because the retriever pulls chunks from documents the user is not authorized to see. The engineer must enforce per-user document ACLs at query time without re-indexing the corpus. Which approach should the engineer take?
Medium250Which CI/CD approach for model deployment best minimizes downtime during a model update?
Hard251When evaluating a generative AI model using the Mosaic AI Model Evaluation tool, what is the primary purpose of providing a 'baseline' dataset?
Easy252An engineer is designing a Databricks RAG application that must support multi-turn conversations where follow-up questions refer to earlier turns. They want the retrieval step to remain accurate as the conversation progresses. Which TWO design elements should they include? (Choose two.)
Hard253Which TWO factors are most important when selecting a chunking strategy for text data prior to vectorization?
Hard254A data engineer needs to ensure that sensitive PII columns are masked for specific groups while remaining visible to analysts. Which Unity Catalog feature should be used to implement this requirement?
Medium255A GenAI engineer is designing an agentic application on Databricks that uses a foundation model to decide which external tools to call. The team wants the agent to be able to invoke a Databricks SQL warehouse query and a Python function registered as a Unity Catalog function, and they need the model to select tools based on natural language requests. Which design element is required for the model to select and invoke these tools correctly?
Hard256An engineer is deploying a RAG application whose retrieval step calls an external vector database over the public internet. The serving endpoint must reach that database, but corporate policy forbids exposing credentials in the model artifact and forbids public egress from the serving environment. Which configuration satisfies both constraints?
Hard257Which approach is recommended for managing secrets, such as API keys for external LLM providers, in Databricks?
Medium258Refer to the exhibit. A developer is deploying a model using the provided JSON configuration. What is the primary benefit of setting 'scale_to_zero_enabled' to true in this production RAG application?
Medium259When deploying an application using Databricks Asset Bundles (DABs), which file is the primary entry point to define the project structure and configuration?
Medium260An engineer registers a GenAI agent to Unity Catalog and enables inference tables on its Mosaic AI Model Serving endpoint. They want to automatically detect when response quality degrades in production without waiting for human review. Which capability should they configure to achieve continuous automated quality monitoring?
Hard261A GenAI team at a retail company has built a RAG chatbot on Databricks that answers customer questions from a product catalog stored in Delta Lake. The catalog is updated nightly, and the team wants the chatbot to reflect those updates without manual intervention. They are deciding how to keep the Vector Search index synchronized. Which approach should they use?
Medium262A GenAI engineer is building an agent with LangChain on Databricks. The agent must call a Unity Catalog function `catalog.schema.get_weather` to fetch current weather. They want the LLM to decide when to invoke this function. Which LangChain component should they use to expose the Unity Catalog function to the LLM?
Medium263An engineer is preparing to deploy a Mosaic AI Agent application with Databricks Asset Bundles. The bundle defines the agent, the serving endpoint, and a job that refreshes the vector index. The engineer wants the deployment to target a staging workspace and a production workspace with different endpoint names and different Unity Catalog catalog names, without editing files between deployments. Which approach should the engineer use?
Medium264A GenAI engineer is building a multi-step agent that uses Databricks Foundation Model APIs. The agent must decide when to call a weather tool and when to answer directly. The engineer wants to ensure the agent's decision-making is reliable and that failures in tool calls are handled gracefully. Which design approach should the engineer use?
Hard265Refer to the exhibit. Why did the analyst group lose access after the table was recreated?
Hard266Refer to the exhibit. The user is a member of the 'finance_team'. Why might the user encounter an access error when executing this join query?
Hard267A GenAI engineer has registered a RAG chain in Unity Catalog as a model and now needs to deploy it for real-time inference with per-request token usage and latency captured automatically. Which Databricks capability should they enable on the serving endpoint?
Medium268Which TWO of the following are mandatory requirements for developing an AI application using the Databricks Mosaic AI Model Serving environment?
Hard269An engineer is packaging a GenAI agent application with Databricks Asset Bundles so that the same bundle deploys to a development and a production workspace. The agent's serving endpoint name must differ per target, and the production endpoint needs more concurrent capacity. Which mechanism in the bundle configuration should the engineer use?
Medium270Which TWO of the following are benefits of using Databricks Asset Bundles for deploying AI applications?
Medium271Which of the following is a primary reason to prefer Databricks Asset Bundles (DABs) over manual workspace deployment?
Easy272A developer is creating a Databricks notebook to prototype a GenAI application. They need to install the `databricks-langchain` library to use LangChain integrations with Databricks. Which command should they use in the notebook?
Easy273When designing a production-ready Databricks notebook for model inference, which TWO practices improve maintainability and performance?
Medium274A team stores raw documents in a Unity Catalog volume and needs to build a training set for fine-tuning a large language model. The raw files are in mixed formats including PDF, DOCX, and plain text. The team wants a single Delta table where every row is one document with its extracted text and source path, and wants the extraction to run in parallel across the cluster. Which approach should the team use?
Easy275What is the primary function of the 'Prompt' object in a Databricks RAG architecture?
Easy276A developer is using MLflow to track experiments for a RAG application. They want to log the retrieval step's parameters, such as the number of documents retrieved (k) and the embedding model used. Which MLflow API should they use?
Easy277A developer deploys a new model version as shown in the exhibit. What is the purpose of this configuration?
Hard278A company is preparing a generative AI application for production and must demonstrate that model inputs and outputs are traceable and that access to sensitive prompt data is controlled. Which TWO Unity Catalog capabilities should the team rely on to meet these governance objectives? (Choose two.)
Medium279A team has deployed a RAG application on Databricks and wants to monitor the quality of responses in production. They have enabled inference table logging. Which built-in Databricks capability allows them to periodically evaluate the logged requests and responses for quality metrics like groundedness?
Easy280Refer to the exhibit. What is the current permission state for the 'analyst_group' after the execution of the REVOKE statement?
Hard281When developing a feature engineering pipeline using Feature Store, which practice ensures maximum code reusability across training and inference?
Medium282Which TWO actions are necessary to ensure that a model serving endpoint in Databricks remains available and performant during peak traffic hours?
Medium283A team is preparing a Delta table of product descriptions for a RAG application. The table receives continuous upserts from a streaming pipeline, and the embedding job reads the table every hour. Engineers notice the embedding job reprocesses every row on each run even though only a few rows change. Which change should be made to the source table to let the embedding job process only new or updated rows?
Hard284Which THREE components are critical to include in a comprehensive evaluation strategy for a RAG-based Generative AI application?
Hard285A GenAI engineer has written a notebook that calls the Databricks Foundation Model APIs to summarize documents. Before the notebook can run in production, the security team wants to confirm exactly which workspace users and service principals are permitted to invoke the pay-per-token foundation model endpoints. Where should the engineer point them to review and manage those permissions?
Easy286Which Databricks feature is specifically designed to facilitate the rapid development and deployment of LLM applications by providing a managed environment for hosting and testing prompts?
Easy287A Generative AI engineer is deploying a new version of a RAG chain to a Mosaic AI Model Serving endpoint. Before promoting it to production, they want to run an evaluation that checks whether the generated answers are faithful to the retrieved documents. Which Mosaic AI Agent Evaluation metric should they examine?
Easy288An engineer is designing a RAG application that uses Databricks Vector Search to retrieve documents and a foundation model endpoint to generate answers. The team wants to log all user queries, retrieved documents, and generated responses for auditing and continuous improvement. They also need to monitor for drift in retrieval quality over time. Which Databricks capability should they integrate into the application design?
Medium289An engineer uses MLflow LLM Evaluation with mlflow.evaluate() to score a RAG application. The judge model is configured with a temperature of 0.9 and no explicit metric thresholds are set. Reruns of the identical evaluation dataset produce relevance scores that swing by up to 20 percentage points, and the team cannot tell whether a prompt change helped. Which change most directly improves the reliability of the evaluation comparison?
Hard290You are building an application that uses Model Serving to host a fine-tuned LLM. Which configuration is required to optimize for high-concurrency request throughput?
Medium291Refer to the exhibit. An engineer observes an unexpected drop in "relevance" for the latest deployment. What is the most likely cause related to the evaluation process itself?
Medium292A healthcare company uses Databricks to build a RAG application over clinical notes stored in a Unity Catalog table. An auditor requires proof that only authorized personnel can view raw note text, while the RAG application must use embeddings generated from those notes. The team wants to avoid copying data outside Unity Catalog. Which approach best satisfies the auditor while preserving RAG functionality?
Hard293An engineer is cleaning a Delta table of customer support transcripts stored in Unity Catalog. The table has a nested array column named messages, where each element has role and content fields. They need to keep only rows where at least one message has role equal to 'user' and content longer than 20 characters. Which transformation correctly expresses this filter?
Medium294When preparing a dataset for fine-tuning an LLM, you need to ensure the data is representative of the target domain. What is the most effective approach to detect and mitigate sampling bias in your training set using Databricks?
Medium295A developer is building a RAG application using Mosaic AI Model Serving. They need to ensure that the embedding model endpoint is strictly accessed only by specific service principals within the workspace. Which feature should the developer configure to enforce this security requirement?
Medium296A developer is building a retrieval-augmented generation (RAG) application on Databricks. They need to ensure that embeddings are updated automatically when the underlying Delta table changes. Which approach is the most efficient and scalable?
Medium297A team is using MLflow LLM Evaluation with the built-in answer_correctness metric to compare two prompt templates for a question-answering application. They notice that answer_correctness scores are nearly identical, but manual review shows one template produces answers that are factually correct yet omit key supporting details. Which additional built-in metric should they add to their evaluation to surface this difference?
Hard298A team is preparing a large text corpus for embedding generation with a foundation model endpoint on Databricks. They must reduce token cost and improve retrieval quality before vectorization. Which two preprocessing steps should be applied to the raw text? (Choose two.)
Hard299An AI engineer is developing a custom Databricks App using Mosaic AI Agent Framework and needs to deploy the application workspace securely. Which deployment artifact and configuration mechanism should the engineer use to define the app dependencies and entry point?
Medium300A developer is building a RAG application and notices that the retrieval step often returns irrelevant context. Which step in the pipeline should be improved to address this?
Medium301A developer is configuring a RAG application and needs to ensure that the LLM response is based on specific, trusted document snippets. Which technique, when implemented correctly, helps mitigate hallucination by grounding the response in provided context?
Medium302A GenAI engineer is designing a RAG application on Databricks that must support multi-turn conversations where users refer to earlier messages, and the application must keep responses grounded in retrieved documents. Which TWO design elements are required to meet these requirements? (Choose two.)
Medium303A GenAI engineer registers a fine-tuned model in Unity Catalog and wants a downstream application to call it through the Databricks Model Serving endpoint without embedding a long-lived personal access token in the application. The application runs on Azure Databricks and must authenticate as its own identity, and the security team requires that credentials be short-lived and automatically rotated. Which authentication approach should the engineer implement?
Hard304A data engineer needs to prepare a Delta table of customer reviews for embedding generation. The reviews contain HTML tags, inconsistent whitespace, and mixed casing that hurt embedding quality. Which preparation step should be applied before generating embeddings?
Easy305An application requires low-latency retrieval of RAG metadata stored in Databricks. Which storage approach balances performance and cost while ensuring seamless integration with Unity Catalog?
Medium306A Generative AI engineer is evaluating a RAG pipeline using MLflow LLM Evaluation on Databricks. They want to assess both the retrieval quality and the generation quality. Which TWO built-in evaluation metrics should they use? (Choose two.)
Hard307A data engineer is assembling a fine-tuning dataset from a Delta table of conversation transcripts. Each transcript contains a list of message objects with a role and a content field, and the training job requires one row per conversation with the messages serialized into the expected format. Which transformation should the engineer apply?
Medium308An engineer is building a multi-turn GenAI assistant on Databricks. The assistant must answer follow-up questions that reference earlier turns, such as 'what about its warranty?', while keeping each request within the model's context limit. Which design should the engineer implement?
Medium309Refer to the exhibit. An engineer is testing a model endpoint. The outputs are too brief and often stop mid-sentence. What is the most likely cause?
Hard310Refer to the exhibit. What is the cause of this error when logging a RAG chain to MLflow?
Medium311A GenAI engineer is designing a retrieval-augmented generation application whose source documents are long PDFs. Early testing shows that answers are vague because retrieved chunks contain several unrelated topics, and the language model frequently cites content that does not support its claims. The engineer wants to improve chunk quality before indexing. Which TWO changes should the engineer make to the ingestion pipeline? (Choose two.)
Hard312A generative AI team runs a RAG chatbot whose Mosaic AI Model Serving endpoint is monitored in Unity Catalog inference tables. Over two weeks, the percentage of user questions that receive a refusal answer ('I don't have enough information') climbs from 4% to 31%, while retrieval latency and token counts stay flat. The team wants the earliest actionable signal that the retrieval corpus has gone stale rather than the prompt or model. Which monitoring signal should they inspect first?
Medium313A generative AI team is building a customer-support assistant on Databricks. The assistant must answer questions using the company's private knowledge base, and the team wants to minimize latency while ensuring that the LLM only uses retrieved documents. They plan to use Databricks Vector Search with a Delta table as the source. Which design choice best balances low latency and grounded responses?
Medium314An engineer is preparing to deploy a Mosaic AI Agent to a Model Serving endpoint. The agent depends on a custom Python library that is not available on PyPI. Which approach ensures the library is available at serving time?
Medium315Refer to the exhibit. A developer encounters this error when trying to call a Model Serving endpoint from a job. Which action should the developer take to resolve this authorization failure?
Hard316A team runs an offline evaluation of a RAG chatbot using MLflow LLM Evaluation with an LLM judge. The same evaluation dataset produces noticeably different scores when re-run on different days, even though the application code and retrieved documents are unchanged. Which action best addresses this score instability?
Hard317An engineer is designing a RAG application where the LLM must answer questions using only the retrieved context and must refuse to answer when the context is insufficient. Which prompt design approach best enforces this behavior?
Medium318An engineer maintains a GenAI application that uses a Databricks Asset Bundle to deploy a Mosaic AI Agent serving endpoint. A new model version has been logged and validated, and the team wants to roll it out to production with the ability to revert quickly if quality regressions appear. Which deployment approach best satisfies this?
Hard319An engineer has finished building a RAG chatbot and wants to expose it as a Databricks App so business users can reach it through a browser. The app needs a Python web server and a command that starts it. Which artifact in the app's project layout defines the runtime command and dependencies used when the app is deployed?
Easy320A GenAI engineer has built a retrieval-augmented generation (RAG) application using Databricks Vector Search and a Databricks-hosted LLM served via Mosaic AI Model Serving. Users report that responses are sometimes irrelevant or cite incorrect document passages. The engineer wants to systematically improve answer quality by identifying which retrieved chunks are actually being used by the LLM. Which approach should the engineer take to capture the relationship between retrieved context and the generated response for later evaluation?
Medium321A data engineer is preparing a large corpus of customer support transcripts stored as Parquet in a Unity Catalog volume. Before generating embeddings, each transcript must be tokenized and truncated to a maximum token length. The engineer wants to use a Databricks-native approach that runs distributed across the cluster and avoids pulling the full corpus into a single node. Which approach best satisfies these requirements?
Hard322A GenAI engineer is deploying a RAG application that uses Databricks Vector Search and a Foundation Model API. The solution must comply with governance policies that require all data access and model invocations to be auditable and access-controlled at a fine-grained level. Which two Unity Catalog features should the engineer leverage to meet these requirements? (Choose two.)
Medium323A team is deploying a LLM-based application using Databricks Model Serving. They want to implement robust observability and monitoring for their endpoint. Which TWO features should they utilize to track performance and quality metrics? (Select TWO)
Hard324A support team operates a Databricks-hosted RAG assistant and wants end-user feedback to feed their monitoring dashboards. They plan to add a thumbs-up and thumbs-down control to the chat UI and log each vote alongside the request ID. What is the primary value of collecting this feedback for the evaluation and monitoring workflow?
Easy325An engineer is designing a Databricks GenAI application that summarizes customer meeting notes. The notes contain personally identifiable information, and company policy requires that the summarization model never sends note text to an external provider. The team already has a fine-tuned open model registered in Unity Catalog. Which deployment choice satisfies the policy?
Easy326Which Databricks feature is specifically designed to allow developers to programmatically manage and version their entire data and AI infrastructure as code?
Easy327Which of the following describes the purpose of using a 'Feature Store' when preparing data for Generative AI applications?
Medium328Refer to the exhibit. The developer is attempting to log a custom model to the Unity Catalog. Which configuration is missing to ensure the model is registered correctly under the specified Unity Catalog location?
Hard329What is the primary purpose of the 'bundle validate' command in the Databricks Asset Bundles CLI?
Easy330A GenAI engineer is building a retrieval-augmented generation (RAG) application using Databricks Vector Search. They notice that the retriever sometimes returns irrelevant chunks that hurt answer quality. They want to add a reranking step to improve the relevance of retrieved documents before passing them to the LLM. Which component should they add to their RAG pipeline?
MediumOther domains
All Databricks-GenAI-Assoc exam domains
Frequently asked questions
- What does the scenario questions domain cover on the Databricks-GenAI-Assoc exam?
- scenario questions questions test whether you can apply the concept in context, not just recognise a definition.
- How many questions are in this domain?
- This page lists all 330 scenario questions questions in the Databricks-GenAI-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only scenario questions questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.