Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
A Generative AI engineer is evaluating a RAG pipeline using MLflow LLM Evaluation on Databricks. They want to assess both the retrieval quality and the generation quality. Which TWO built-in evaluation metrics should they use? (Choose two.)
⚠ Common exam trap
The trap here is selecting common NLP metrics like ROUGE or perplexity, which are not part of MLflow's built-in RAG evaluation suite and do not measure retrieval or groundedness.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
`relevance`
For a RAG pipeline, retrieval quality is assessed by relevance, which measures how well retrieved documents match the query. Generation quality is assessed by groundedness, which checks if the answer is supported by the retrieved context. Both are built-in metrics in MLflow LLM Evaluation and are specifically designed for RAG evaluation on Databricks.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
`perplexity`
Why it's wrong here
Perplexity measures how well a language model predicts a sample of text, but it is not a built-in metric in MLflow LLM Evaluation for RAG. It does not evaluate retrieval quality or groundedness. While it can indicate language model fluency, it is not suitable for assessing whether the answer is supported by context. It is not the right choice here.
- ✗
`toxicity`
Why it's wrong here
Toxicity evaluates whether the generated content is harmful or offensive. While important for safety, it does not measure retrieval quality or whether the answer is grounded in the retrieved context. It is a separate dimension of evaluation. The team specifically wants to assess retrieval and generation quality, so toxicity is not the appropriate metric here.
- ✗
`ROUGE`
Why it's wrong here
ROUGE is an n-gram overlap metric for summarization and translation, not a built-in RAG evaluation metric in MLflow. It compares generated text to reference summaries and does not assess retrieval relevance or groundedness. Using it would require reference answers and would not capture semantic quality. It is not designed for this purpose.
- ✓
`relevance`
Why this is correct
Relevance assesses how well the retrieved documents match the user's query, directly evaluating retrieval quality. In MLflow LLM Evaluation, the relevance metric can be computed for the retrieval step. It helps identify whether the retriever is fetching appropriate context. This is essential for diagnosing retrieval failures in a RAG pipeline.
- ✓
`groundedness`
Why this is correct
Groundedness measures whether the generated answer is supported by the retrieved context. It is a key metric for generation quality in RAG, as it detects hallucinations. MLflow provides a built-in groundedness judge that evaluates the answer against the context. Using it helps ensure the model relies on retrieved evidence rather than fabricating information.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.