Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
A financial services team runs a RAG assistant on Databricks that answers questions about internal policy documents. During a review, the team finds that for many questions the answer is factually correct but cites a document that does not actually contain the supporting statement. They want an MLflow LLM Evaluation metric that specifically detects when the response is not supported by the retrieved context. Which metric should they add to their evaluation run?
⚠ Common exam trap
It's easy for candidates to confuse factual correctness with grounding, assuming that an accurate answer must also be supported by the retrieved context.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
groundedness
Groundedness evaluates whether the response's claims can be traced back to the retrieved context, which is precisely the gap identified in this review. Adding it to the MLflow LLM Evaluation run gives the team a metric that flags answers which sound correct but are not actually supported by the cited policy documents.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
answer_similarity
Why it's wrong here
answer_similarity compares the generated response to a provided ground-truth answer using semantic similarity. It can flag that an answer differs from the reference, but it does not inspect the retrieved context, so it cannot detect that a response is unsupported by the documents the assistant cited in this policy-assistant scenario.
- ✓
groundedness
Why this is correct
The groundedness metric in MLflow LLM Evaluation judges whether the claims in the response are supported by the retrieved context. In this scenario the answers are correct but not backed by the cited documents, which is exactly the failure mode groundedness is designed to surface, making it the right metric to add to the evaluation run.
- ✗
retrieval_precision
Why it's wrong here
retrieval_precision measures how many of the retrieved chunks are relevant to the question, focusing on the retrieval stage rather than the generation stage. Here the retrieved documents may be relevant but the answer still cites content they do not contain, so retrieval precision would not expose the lack of contextual support.
- ✗
toxicity
Why it's wrong here
The toxicity metric checks whether the response contains harmful, offensive, or inappropriate language. The problem described is a factual grounding problem, not a safety problem, so adding toxicity would not reveal that correct-sounding answers are unsupported by the retrieved policy documents the team relies on.
About these practice questions
One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.