Be able to instrument a RAG or agent pipeline with MLflow Tracing, run Mosaic AI Agent Evaluation with appropriate judges and metrics, and interpret baseline comparisons. The key is matching each metric to what it actually measures: retrieval quality versus response quality versus safety.
Start practicing
Evaluation and Monitoring — choose a session length
Free · No account required
Domain overview
This domain covers observability, quality measurement, and evaluation tooling for generative AI apps on Databricks. Questions test MLflow tracing and evaluation, Mosaic AI Agent Evaluation, judge-based scoring, and RAG quality metrics. Expect scenario items asking you to pick the right approach for diagnosing latency, comparing evaluation methods, or interpreting baseline and metric choices.
Exam objectives
MLflow Tracing to capture spans across retrieval, LLM calls, and tool steps for latency debugging
Mosaic AI Agent Evaluation with LLM judges scoring correctness, groundedness, relevance, and safety
MLflow Model Evaluation APIs that compute metrics programmatically over evaluation datasets at scale
Baseline datasets in Mosaic AI Model Evaluation used to compare candidate model outputs against reference behavior
Treating manual spot-checks or generic logging as sufficient observability instead of instrumenting the pipeline with MLflow Tracing spans.
Confusing retrieval metrics like context relevance with answer-quality metrics like groundedness and correctness when selecting RAG evaluation metrics.
Assuming a baseline dataset is training data or a fine-tuning input rather than a reference for comparing evaluation results.
Click any question to see the full explanation and answer options, or start a focused practice session above.
A data engineering team is deploying a RAG application using Mosaic AI Model Serving. They need to monitor the quality of the model's responses in production. Which Databricks feature should they use to capture and analyze inference data, such as requests, responses, and latency metrics?
2Which of the following metrics is most effective for evaluating a RAG-based chatbot's ability to retrieve relevant context from a vector database during production monitoring?
3A Generative AI engineer is configuring a Mosaic AI Model Serving endpoint for a production-grade LLM. Which TWO of the following tasks are necessary to ensure effective monitoring and evaluation of the endpoint? (Choose two)
4When evaluating a generative AI model using the Mosaic AI Model Evaluation tool, what is the primary purpose of providing a 'baseline' dataset?
5Which THREE of the following are common challenges when monitoring LLM applications in production that differ significantly from traditional ML monitoring? (Choose three)
6When monitoring a RAG application, you notice a high discrepancy between the retrieved context and the generated answer. Which metric would specifically help identify if the model is ignoring the provided context?
7A team is using MLflow to track their model evaluation runs. They want to ensure that every evaluation run is reproducible. What is the best practice to achieve this in Databricks?
8When evaluating LLM outputs, which TWO metrics are most appropriate for measuring the 'quality' of a response in a RAG system? (Choose two)
9Which of the following describes the 'drift' phenomenon in the context of LLM monitoring?
10Why should you integrate Mosaic AI Model Evaluation with Unity Catalog?
11An enterprise LLM application on Databricks is experiencing high latency and inconsistent responses. Which approach best enables observability to identify the root cause of these performance bottlenecks within the LLM pipeline?
12Which TWO evaluation approaches are most effective for measuring the quality of a RAG pipeline's retrieval stage?
13What is the primary benefit of using MLflow Model Evaluation for generative AI applications compared to manual evaluation methods?
14When monitoring a production LLM, you detect a drift where the model's responses are becoming increasingly verbose and less helpful compared to the baseline. Which strategy is most effective for detecting this quality decay?
15Which THREE components are critical to include in a comprehensive evaluation strategy for a RAG-based Generative AI application?
16Your organization requires an auditable record of all LLM evaluation results for compliance. Which Databricks feature provides the best centralized storage for these evaluation runs?
17Refer to the exhibit. An engineer observes an unexpected drop in "relevance" for the latest deployment. What is the most likely cause related to the evaluation process itself?
18A Databricks Generative AI engineer has deployed a RAG application and is now setting up production monitoring. They want to automatically detect when the distribution of incoming user questions diverges from the distribution seen during development, so they can trigger retraining or prompt adjustments. Which Databricks capability should they configure?
19A team is using MLflow LLM Evaluation with the built-in answer_correctness metric to compare two prompt templates for a question-answering application. They notice that answer_correctness scores are nearly identical, but manual review shows one template produces answers that are factually correct yet omit key supporting details. Which additional built-in metric should they add to their evaluation to surface this difference?
20A Generative AI engineer is deploying a new version of a RAG chain to a Mosaic AI Model Serving endpoint. Before promoting it to production, they want to run an evaluation that checks whether the generated answers are faithful to the retrieved documents. Which Mosaic AI Agent Evaluation metric should they examine?
21A GenAI engineering team has deployed a customer-support RAG chain on Databricks and registered it in Unity Catalog. They now want MLflow 3 to automatically score every production request for groundedness and relevance without writing custom scoring code, and to persist those assessments against the logged traces. Which approach should they use?
22A team runs an LLM-as-a-judge evaluation on Databricks using the built-in `mlflow.evaluate()` with `model_type="databricks-agent"`. They notice that the judge model, a serving endpoint, is producing scores that are consistently inflated compared to human review. Which configuration change should the team make first to improve the reliability of the evaluation?
23A team runs a RAG application on Mosaic AI Model Serving and logs all requests to an inference table. Reviewers report that answers are sometimes fluent but contradict the retrieved documents. The team wants a recurring, automated check that quantifies this contradiction on production traffic and alerts when it exceeds a threshold. Which approach best fits?
24A Generative AI engineer at a retail company has deployed a RAG chatbot on a Databricks Mosaic AI Model Serving endpoint. The team wants to capture per-request evaluation data, including the user's question, the retrieved context chunks, and the model's response, so they can analyze response quality over time. Which Databricks feature should they use to collect this data for downstream evaluation?
25A team runs a RAG chatbot whose MLflow evaluation with the built-in groundedness judge previously scored well. After they swap the retriever for a new embedding model, groundedness scores drop sharply even though the generator model and prompt are unchanged. They confirm the judge model itself is unchanged. Which action should they take FIRST to diagnose the regression?
26A Generative AI engineer is using MLflow Tracing to monitor a RAG application deployed on Mosaic AI Model Serving. They want to capture the retrieved documents, the final prompt, and the model's response for each request to debug a quality issue. Which approach should they use to ensure all three are logged in a single trace?
27A financial services team runs a RAG assistant on Databricks that answers questions about internal policy documents. During a review, the team finds that for many questions the answer is factually correct but cites a document that does not actually contain the supporting statement. They want an MLflow LLM Evaluation metric that specifically detects when the response is not supported by the retrieved context. Which metric should they add to their evaluation run?
28An engineer registers a GenAI agent to Unity Catalog and enables inference tables on its Mosaic AI Model Serving endpoint. They want to automatically detect when response quality degrades in production without waiting for human review. Which capability should they configure to achieve continuous automated quality monitoring?
29A team has deployed a RAG chatbot on Databricks and enabled inference table logging. They want to set up automated monitoring to detect when the average response length increases significantly compared to the baseline. Which Databricks feature should they use to create this monitor?
30A healthcare analytics team has a RAG application that must not reveal protected health information from other patients. They want to continuously monitor production traffic on their Mosaic AI Model Serving endpoint and alert when responses contain unsafe content. Which Databricks capability should they configure to evaluate each logged request and response against safety criteria and route flagged records for review?
31A team's RAG evaluation shows high context recall but low answer correctness. Retrieved documents contain the needed facts, yet the generated answers frequently contradict them. Which single metric should they examine next to pinpoint whether the generator is ignoring or misusing the retrieved context?
32A Generative AI engineer is evaluating a RAG pipeline using MLflow LLM Evaluation on Databricks. They want to assess both the retrieval quality and the generation quality. Which TWO built-in evaluation metrics should they use? (Choose two.)
33A media company's RAG assistant answers questions about its streaming catalog. Users report that the assistant often returns answers that ignore the retrieved documents and instead rely on outdated information from the base model. The team wants an MLflow LLM Evaluation metric that measures how much of the final answer is actually derived from the retrieved context rather than the model's prior knowledge. Which metric best fits this need?
34A Generative AI engineer is designing an evaluation harness for a customer-support RAG agent on Databricks. They need metrics that specifically assess the RETRIEVAL stage rather than the generation stage. (Choose two.)
35A GenAI engineer notices that a RAG agent's retrieval stage is returning relevant chunks, but the generated answers frequently omit key facts present in those chunks. The team wants a single evaluation metric that isolates whether the generator is using the provided context. Which metric should they focus on?
36A Generative AI engineer is setting up MLflow LLM Evaluation for a RAG pipeline on Databricks that answers questions about product manuals. The team wants to evaluate both the retrieval stage and the generation stage in a single evaluation run. Which TWO metric groups should they include in the evaluation configuration? (Choose two.)
37A team has deployed a RAG application on Databricks and wants to monitor the quality of responses in production. They have enabled inference table logging. Which built-in Databricks capability allows them to periodically evaluate the logged requests and responses for quality metrics like groundedness?
38A team wants every MLflow evaluation run of their RAG chain to be reproducible months later, including the exact prompt template, model parameters, and retrieved context used. Which practice best ensures this reproducibility?
39A team runs an offline evaluation of a RAG chatbot using MLflow LLM Evaluation with an LLM judge. The same evaluation dataset produces noticeably different scores when re-run on different days, even though the application code and retrieved documents are unchanged. Which action best addresses this score instability?
40A team uses an LLM judge to score their RAG agent on Databricks and sees scores that fluctuate by several points between identical evaluation runs. They need more stable, reproducible quality signals for release gating. Which action best addresses the root cause?
41A GenAI team at a retail bank runs a RAG assistant on a Databricks Mosaic AI Model Serving endpoint. During a pilot, they captured end-user thumbs-up/down feedback in a Delta table but never joined it to the trace payloads. Six weeks later, hallucination complaints spike, yet the aggregate thumbs-down rate is unchanged. Which approach best resolves this discrepancy using Databricks-native tooling?
42A GenAI engineer is instrumenting a production RAG application on Databricks to detect quality degradation before users complain. The team wants signals that reveal problems in the retrieval stage specifically, rather than in the generation stage. Which TWO signals should they track? (Choose two.)
43A healthcare analytics team deployed a RAG assistant on Databricks whose answers cite clinical policy documents. Compliance requires that every production answer be attributable to a specific retrieved chunk. During evaluation, the team notices the groundedness judge scores are high, but manual review finds answers that blend two policies into a statement neither document supports. Which change to their Mosaic AI Agent Evaluation configuration best detects this failure mode?
44A support team operates a Databricks-hosted RAG assistant and wants end-user feedback to feed their monitoring dashboards. They plan to add a thumbs-up and thumbs-down control to the chat UI and log each vote alongside the request ID. What is the primary value of collecting this feedback for the evaluation and monitoring workflow?
45A generative AI team runs a RAG chatbot whose Mosaic AI Model Serving endpoint is monitored in Unity Catalog inference tables. Over two weeks, the percentage of user questions that receive a refusal answer ('I don't have enough information') climbs from 4% to 31%, while retrieval latency and token counts stay flat. The team wants the earliest actionable signal that the retrieval corpus has gone stale rather than the prompt or model. Which monitoring signal should they inspect first?
46A fintech company runs a customer-support RAG assistant on Databricks. Before promoting a new prompt template from staging to production, the ML team must demonstrate that the change does not regress answer quality. Which TWO evaluation practices should they apply in Mosaic AI Agent Evaluation to make this promotion decision defensible? (Choose two.)
47A GenAI engineer monitors a customer-facing RAG assistant hosted on Databricks. After a routine re-indexing job, groundedness scores from the MLflow LLM Evaluation job drop sharply while answer relevance stays flat. The application prompt and the LLM serving endpoint were untouched. Which conclusion is best supported by these signals?
48An engineer uses MLflow LLM Evaluation with mlflow.evaluate() to score a RAG application. The judge model is configured with a temperature of 0.9 and no explicit metric thresholds are set. Reruns of the identical evaluation dataset produce relevance scores that swing by up to 20 percentage points, and the team cannot tell whether a prompt change helped. Which change most directly improves the reliability of the evaluation comparison?
49A logistics company monitors a Databricks-hosted RAG assistant that answers questions about shipping regulations. Over one week, the retrieval index was rebuilt after a document refresh, and the team observes that the rate of answers flagged as unsupported by retrieved context rose sharply while retrieval relevance scores stayed flat. Which metric should they inspect first to determine whether the regression originates in the retrieval stage or the generation stage?
50A team deploys a customer-support assistant on a Mosaic AI Model Serving endpoint and enables inference table logging to Unity Catalog. Compliance requires that every production response be traceable back to the exact request, the retrieved context, and the model version that produced it, and that reviewers can query this history with SQL months later. Which capability satisfies this requirement?
51A RAG assistant built with Mosaic AI Agent Framework is in production. Product managers complain that answers are sometimes plausible but unsupported by the retrieved documents, and separately that the assistant sometimes ignores documents that clearly contain the answer. The team wants automated, scheduled checks in MLflow LLM Evaluation that separate these two failure modes so each can be triaged independently. Which TWO evaluation metrics should they add to the evaluation suite? (Choose two.)
52A media company uses an LLM-as-a-judge evaluation pipeline on Databricks to score a summarization assistant nightly. The judge is the same model family as the assistant and was given a rubric that rewards stylistic fluency. Over two months, nightly scores climb steadily while editor spot-checks find summaries increasingly omit key facts. Which corrective action best restores the evaluation's ability to detect this regression?
53A team uses MLflow LLM Evaluation with an LLM judge to score a summarization agent. On a 400-row golden dataset the judge marks 96% of summaries as relevant. Spot-checking reveals the judge approves nearly every summary whenever the summary is fluent, even when key facts are missing. The team wants a defensible quality signal before approving a release. Which action best addresses this judge weakness?
54A retail company runs a customer-facing RAG assistant on a Mosaic AI Model Serving endpoint. The team wants every production request, response, and retrieved context to be captured automatically into a Unity Catalog Delta table so they can monitor quality and latency trends over time. Which action should they take?
Be able to instrument a RAG or agent pipeline with MLflow Tracing, run Mosaic AI Agent Evaluation with appropriate judges and metrics, and interpret baseline comparisons. The key is matching each metric to what it actually measures: retrieval quality versus response quality versus safety.
The Courseiva Databricks-GenAI-Assoc question bank contains 54 questions in the Evaluation and Monitoring domain. Click any question to see the full explanation and answer breakdown.
Start with a 10-question focused session to identify your baseline accuracy in this domain. Read every explanation — even for questions you answer correctly — to understand the reasoning. Once you score consistently above 80%, move to a 20–30 question session to confirm depth before moving to the next domain.
Yes — the session launcher on this page draws questions exclusively from the Evaluation and Monitoring domain. Choose 10, 20, 30, or 50 questions for a focused session, or click individual questions to review them one by one.
Save your results, see per-domain analytics, and get readiness scores — free, for every certification.
Sign Up FreeFree forever · Every certification included