Databricks-GenAI-Assoc • Practice Test 20
Free Databricks-GenAI-Assoc practice test — 15 questions with explanations. Set 20. No signup required.
An engineer uses MLflow LLM Evaluation with mlflow.evaluate() to score a RAG application. The judge model is configured with a temperature of 0.9 and no explicit metric thresholds are set. Reruns of the identical evaluation dataset produce relevance scores that swing by up to 20 percentage points, and the team cannot tell whether a prompt change helped. Which change most directly improves the reliability of the evaluation comparison?
Choose an answer to begin — your selection is scored in the full session.
15 questions · instant feedback and full explanations after every question.