Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
A team uses MLflow LLM Evaluation with an LLM judge to score a summarization agent. On a 400-row golden dataset the judge marks 96% of summaries as relevant. Spot-checking reveals the judge approves nearly every summary whenever the summary is fluent, even when key facts are missing. The team wants a defensible quality signal before approving a release. Which action best addresses this judge weakness?
⚠ Common exam trap
The trap here is trying to fix a systematically lenient judge by enlarging the dataset or lowering rigor, when the real need is to calibrate the judge against human labels.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Add a small human-labeled calibration set and compare the judge's verdicts against human labels before trusting the automated relevance scores.
An automated judge that approves nearly every fluent summary is a biased measurement instrument, and the fix is to calibrate it rather than to change its sampling or the dataset size. Human-labeling a representative subset and measuring judge-human agreement exposes the leniency and tells the team how much to trust the scores. Only then can thresholds be set or a complementary fact-coverage check be added, producing a release signal that can be defended.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Replace the relevance metric with a cheaper lexical overlap score such as ROUGE against the reference summaries.
Why it's wrong here
Lexical overlap rewards surface word matching and is notoriously insensitive to missing key facts, since a fluent paraphrase can share few n-grams with the reference while omitting essential content. Swapping one flawed instrument for another does not validate the quality signal and may trade leniency for a different bias. It also discards the judge's ability to assess semantic relevance, moving away from the goal of a defensible release decision.
- ✗
Expand the golden dataset from 400 rows to 4,000 rows and rerun the same relevance metric.
Why it's wrong here
More rows reduce the confidence interval around an estimate, but they do not remove bias. If the judge approves almost every fluent summary, a larger dataset will report the same inflated rate with a narrower confidence interval, making a wrong number look more authoritative. Scale cannot substitute for validating the measurement instrument itself, which is the actual defect here.
- ✓
Add a small human-labeled calibration set and compare the judge's verdicts against human labels before trusting the automated relevance scores.
Why this is correct
The judge exhibits a systematic bias toward fluency, so its scores cannot be trusted at face value. Labeling a representative subset by hand and measuring agreement between judge and human verdicts quantifies that bias and reveals which cases the judge mishandles. Once agreement is characterized, the team can report calibrated results, adjust thresholds, or supplement the judge with a fact-coverage metric, making the release decision defensible rather than resting on inflated automated scores.
- ✗
Raise the judge's temperature so it explores a wider range of judgments and produces a more nuanced relevance distribution.
Why it's wrong here
Higher temperature increases randomness, not accuracy, so a judge that already over-approves fluent summaries will simply become inconsistent as well as biased. The reported 96% approval rate is a systematic leniency problem, and adding sampling noise would make it harder to detect. This change degrades reproducibility without correcting the underlying preference for fluent text, so release decisions would become less defensible.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.