Courseiva

Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question

An engineer uses MLflow LLM Evaluation with mlflow.evaluate() to score a RAG application. The judge model is configured with a temperature of 0.9 and no explicit metric thresholds are set. Reruns of the identical evaluation dataset produce relevance scores that swing by up to 20 percentage points, and the team cannot tell whether a prompt change helped. Which change most directly improves the reliability of the evaluation comparison?

⚠ Common exam trap

The trap here is treating evaluation variance as a dataset-size problem when the root cause is a high-temperature judge producing different verdicts on identical inputs.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Set the judge model's temperature to 0 and define fixed pass/fail thresholds for each metric before comparing runs.

Score swings across identical runs are caused by nondeterministic judge sampling, so the fix must remove that randomness and make the decision rule stable. Setting the judge temperature to zero makes each judgment reproducible, and predefined thresholds turn noisy continuous scores into consistent pass/fail outcomes. Enlarging the judge, duplicating rows, or hiding per-row results leaves the stochasticity in place and keeps run-to-run comparisons unreliable.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the evaluation dataset size by duplicating each question three times so variance averages out across more judge calls.

    Why it's wrong here

    Duplicating identical rows inflates the row count without adding independent information, so the judge sees the same inputs repeatedly. Because the variance comes from the judge's high-temperature sampling, repeating the same items reproduces the same noisy judgments rather than canceling them. Statistical averaging assumes independent samples; cloned questions violate that assumption and can even bias aggregate scores toward whichever verdict the judge happened to favor on that item.

  • ✓

    Set the judge model's temperature to 0 and define fixed pass/fail thresholds for each metric before comparing runs.

    Why this is correct

    The instability originates from stochastic sampling in the judge, and temperature 0 makes the judge's verdicts deterministic for a given input, removing run-to-run score drift. Pairing that with predefined metric thresholds converts continuous, noisy scores into stable pass/fail decisions, so a prompt change can be attributed to real quality movement instead of sampling noise. This combination directly addresses reproducibility and comparability of evaluation runs.

  • ✗

    Switch the evaluation to a larger judge model with more parameters while keeping the existing sampling settings.

    Why it's wrong here

    A larger judge may be more accurate on average, but it still samples at a high temperature, so identical inputs can produce different verdicts on each run. The reported 20-point swings are a determinism problem, not a capability problem. Without lowering temperature or fixing decision thresholds, the team would get a more expensive but equally irreproducible comparison and still could not attribute changes to the prompt edit.

  • ✗

    Log only the aggregate mean relevance score per run and discard the per-row judgments from the MLflow run artifacts.

    Why it's wrong here

    Discarding per-row results hides the very detail needed to diagnose whether variance comes from a handful of ambiguous questions or from the judge overall. Aggregates alone also remove the ability to audit individual verdicts or recompute metrics with different thresholds. This reduces transparency without touching the sampling temperature that causes the swings, so the underlying reproducibility problem remains fully intact.

About these practice questions

Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.