Courseiva
Evaluation and Monitoring →mediumMultiple Choice

Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question

A team runs an LLM-as-a-judge evaluation on Databricks using the built-in `mlflow.evaluate()` with `model_type="databricks-agent"`. They notice that the judge model, a serving endpoint, is producing scores that are consistently inflated compared to human review. Which configuration change should the team make first to improve the reliability of the evaluation?

⚠ Common exam trap

The trap here is assuming that increasing the judge model's temperature will make it more objective, when in fact it adds randomness and worsens reproducibility without fixing bias.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Switch the judge to a different, stronger model endpoint and calibrate it using a small human-labeled dataset with known ground-truth scores.

LLM judges can exhibit systematic bias, such as consistently inflated scores. The most effective first step is to use a stronger judge model and calibrate it against human-labeled data with known scores. This improves reliability and aligns with Databricks recommendations for validating judges before using them in production monitoring. Simply changing temperature or filtering data does not address the root cause.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Switch the judge to a different, stronger model endpoint and calibrate it using a small human-labeled dataset with known ground-truth scores.

    Why this is correct

    Using a more capable judge model and calibrating it against human-labeled examples is the recommended way to reduce systematic bias. Calibration with ground-truth scores helps detect and correct consistent over- or under-scoring. This approach aligns with MLflow's guidance to validate LLM judges before relying on their outputs in production monitoring.

  • ✗

    Reduce the number of evaluation examples to only those where the judge and human agree, then report the average score.

    Why it's wrong here

    Filtering out disagreement hides the bias rather than fixing it and produces misleadingly optimistic metrics. The evaluation would no longer represent the full distribution of inputs. This is a form of cherry-picking that undermines the purpose of monitoring. The correct approach is to improve judge reliability, not to discard inconvenient data.

  • ✗

    Increase the temperature parameter of the judge model to 1.0 to encourage more diverse scoring.

    Why it's wrong here

    Raising the temperature increases randomness in the judge's outputs, which would make scores less stable and less reproducible. It does not address the systematic bias toward inflated scores. Higher temperature is useful for creative generation, not for evaluation consistency. The team needs a more deterministic and better-calibrated judge, not more stochasticity.

  • ✗

    Replace the LLM judge with a BLEU score comparison against a reference answer for every question.

    Why it's wrong here

    BLEU measures n-gram overlap and is poorly suited for open-ended generative answers. It would penalize valid paraphrases and not capture semantic quality. While it is deterministic, it does not solve the bias problem and would introduce a different, often larger, evaluation error. The team needs a judge that understands meaning, not surface form.

About these practice questions

This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.