Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
A healthcare analytics team deployed a RAG assistant on Databricks whose answers cite clinical policy documents. Compliance requires that every production answer be attributable to a specific retrieved chunk. During evaluation, the team notices the groundedness judge scores are high, but manual review finds answers that blend two policies into a statement neither document supports. Which change to their Mosaic AI Agent Evaluation configuration best detects this failure mode?
⚠ Common exam trap
The trap here is assuming a high groundedness score means every claim is properly attributed, when response-level groundedness against merged context can pass blended statements whose parts each appear somewhere in the retrieved pool.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable the chunk-level citation assessment so each claim in the response is checked against the specific retrieved chunk it references, rather than scoring the response against the concatenated context.
The compliance requirement is per-answer attribution to a specific chunk, so the evaluation must verify each claim against the chunk it cites. Scoring the whole response against concatenated context allows a blended claim to pass whenever its components appear somewhere in the pool. Chunk-level citation assessment closes that gap by testing attribution directly, which is what distinguishes this failure from ordinary ungrounded hallucination.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of retrieved chunks per query from five to twenty so the model has more supporting evidence available.
Why it's wrong here
Retrieving more chunks enlarges the context pool, which if anything makes response-level groundedness scoring more permissive because more text is available to justify a claim. It does not introduce per-claim citation checking, so blended statements remain invisible and may even increase as the model has more material to conflate.
- ✓
Enable the chunk-level citation assessment so each claim in the response is checked against the specific retrieved chunk it references, rather than scoring the response against the concatenated context.
Why this is correct
Response-level groundedness against merged context can pass when each claim is individually supported somewhere in the pool, even if the response attributes a blended statement to the wrong source. Chunk-level citation assessment forces per-claim traceability to the cited chunk, which is exactly the compliance requirement and surfaces cross-document blending that aggregate groundedness masks.
- ✗
Raise the temperature setting on the served model to zero and re-run the evaluation suite to obtain more deterministic outputs.
Why it's wrong here
Lowering temperature reduces run-to-run variance but does not change whether the judge checks per-claim attribution. A deterministic model can still blend two policies consistently. Determinism helps reproducibility of the evaluation, yet the failure mode described is a missing attribution check, so this change leaves the compliance risk undetected.
- ✗
Add a relevance judge that scores the retrieved chunks against the user question and filter out any chunk scoring below a fixed threshold.
Why it's wrong here
Relevance measures whether retrieved chunks match the question, not whether generated claims are attributable to the cited chunk. Filtering low-relevance chunks may improve retrieval precision but does nothing to detect a response that correctly retrieves two relevant policies and then synthesizes a claim neither supports, so the compliance gap remains.
About these practice questions
Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.