Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
A Generative AI engineer is designing an evaluation harness for a customer-support RAG agent on Databricks. They need metrics that specifically assess the RETRIEVAL stage rather than the generation stage. (Choose two.)
⚠ Common exam trap
The trap here is treating groundedness as a retrieval metric because it references context, when it actually compares generated claims to context and thus spans both stages.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Context recall, measuring whether the retrieved passages collectively contain the information present in the ground-truth answer.
Context recall and context precision are computed from the retrieved passages and ground-truth relevance, so they quantify what the retriever returned and in what order. Answer relevance, groundedness, and toxicity all depend on the generated response, which means they blend generation behavior into the score. Selecting the two context-based metrics keeps the harness focused on retrieval.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Context recall, measuring whether the retrieved passages collectively contain the information present in the ground-truth answer.
Why this is correct
Context recall compares retrieved passages against the reference answer to determine whether all needed information was retrieved. It is computed purely from retrieval outputs and ground truth, so it isolates the retriever's ability to fetch relevant content independent of how the generator phrases its response.
- ✗
Answer relevance, measuring how well the final generated response addresses the user's original question.
Why it's wrong here
Answer relevance scores the generated response against the query, so it evaluates the generation stage and can be affected by prompt design and decoding. It does not measure what the retriever fetched, making it unsuitable for isolating retrieval performance in this harness.
- ✗
Toxicity, measuring whether the generated response contains harmful or offensive content.
Why it's wrong here
Toxicity is a safety metric over the generated output. It is independent of retrieval quality and is typically used for guardrail monitoring rather than evaluating whether the right documents were fetched, so it does not belong in a retrieval-focused metric set.
- ✗
Groundedness, measuring whether the answer's claims are supported by the retrieved context.
Why it's wrong here
Groundedness compares the generated answer to retrieved context, so it spans both stages and is influenced by the generator. A retriever change can affect it, but it is not a pure retrieval metric and therefore does not satisfy the requirement to isolate the retrieval stage.
- ✓
Context precision, measuring the proportion of retrieved chunks that are actually relevant and how highly they are ranked.
Why this is correct
Context precision evaluates the retrieved chunk list against ground-truth relevance, rewarding relevant chunks ranked near the top. Like context recall, it depends only on retriever output and ground truth, so it directly characterizes retrieval quality without involving the generator's wording.
About these practice questions
One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.