Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
A Generative AI engineer is deploying a new version of a RAG chain to a Mosaic AI Model Serving endpoint. Before promoting it to production, they want to run an evaluation that checks whether the generated answers are faithful to the retrieved documents. Which Mosaic AI Agent Evaluation metric should they examine?
⚠ Common exam trap
A common mix-up: candidates confuse retrieval-stage metrics like chunk_relevance with answer-stage metrics, when faithfulness of the final answer is measured by groundedness.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
groundedness
To verify that generated answers are faithful to the retrieved documents, the evaluation must compare the answer against the context. groundedness is the Agent Evaluation metric that performs this check, flagging responses that include unsupported claims. It is the appropriate pre-promotion quality gate for a RAG chain.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
chunk_relevance
Why it's wrong here
chunk_relevance assesses whether the retrieved chunks are relevant to the user's question, which is a retrieval-stage metric. It does not evaluate the final answer's faithfulness to those chunks. Since the engineer wants to verify that the generated answer is grounded in the retrieved documents, this metric addresses a different stage of the pipeline.
- ✗
latency
Why it's wrong here
latency measures how long the endpoint takes to return a response. While important for user experience, it says nothing about whether the answer is faithful to the retrieved documents. The engineer's stated goal is to check answer faithfulness, so latency is not the relevant metric for this evaluation.
- ✗
request_count
Why it's wrong here
request_count is an operational metric tracking how many requests the endpoint has served. It provides no insight into answer quality or faithfulness. Using it to judge whether answers are grounded in retrieved documents would be meaningless, as it only reflects traffic volume.
- ✓
groundedness
Why this is correct
groundedness measures whether the generated answer is supported by the retrieved context, directly assessing faithfulness to the documents. This is exactly the quality dimension the engineer needs to verify before promotion. It helps catch hallucinations where the model adds information not present in the retrieved chunks.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.