Courseiva

Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question

A team runs an offline evaluation of a RAG chatbot using MLflow LLM Evaluation with an LLM judge. The same evaluation dataset produces noticeably different scores when re-run on different days, even though the application code and retrieved documents are unchanged. Which action best addresses this score instability?

⚠ Common exam trap

The trap here is blaming the application or the retrieval pipeline for score drift, when unchanged inputs and code point squarely at the LLM judge's own nondeterminism.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Pin the judge model version and set temperature to zero for evaluation runs

Identical inputs producing different scores across runs indicate the judge itself is nondeterministic, typically through model version drift or sampling temperature. Pinning the judge model version and forcing temperature to zero makes scoring reproducible, so observed differences can be attributed to the application. The other options either change the system under test, destroy evaluation fidelity, or hide variance rather than eliminate it.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the number of retrieved chunks per query so the judge has more context

    Why it's wrong here

    Adding more retrieved chunks changes what the application sees and may dilute precision, but it does not make the judge deterministic. Score variance across identical inputs points to judge nondeterminism, not insufficient context. This action would alter the system under test rather than stabilize measurement, and could even introduce new noise by surfacing irrelevant passages that confuse the judge.

  • ✗

    Switch all metrics from LLM-judged to exact-match string comparison

    Why it's wrong here

    Exact-match comparison is deterministic but far too brittle for open-ended GenAI responses, where many valid phrasings exist. It would make scores stable only by making them meaningless, and it cannot evaluate groundedness or relevance. The team's problem is judge variance, which is solved by controlling the judge configuration, not by abandoning semantic evaluation entirely.

  • ✗

    Reduce the evaluation dataset to a single representative example

    Why it's wrong here

    Shrinking the dataset hides variance by removing the observations that reveal it, and a single example cannot support any aggregate quality claim. The team still needs coverage across query types to trust the evaluation. This action degrades statistical reliability instead of fixing the nondeterministic judge that causes identical inputs to score differently across days.

  • ✓

    Pin the judge model version and set temperature to zero for evaluation runs

    Why this is correct

    Varying scores on identical inputs are a classic symptom of a nondeterministic judge: the judge model version may silently change, and sampling temperature introduces randomness in its verdicts. Pinning the judge model version and setting temperature to zero makes the scoring function repeatable, so differences in scores reflect real changes in the application rather than judge noise. This directly targets the instability described.

About these practice questions

One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.