Courseiva
Evaluation and Monitoring →mediumMultiple Select

Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question

A fintech company runs a customer-support RAG assistant on Databricks. Before promoting a new prompt template from staging to production, the ML team must demonstrate that the change does not regress answer quality. Which TWO evaluation practices should they apply in Mosaic AI Agent Evaluation to make this promotion decision defensible? (Choose two.)

⚠ Common exam trap

The trap here is chasing a passing aggregate score by enlarging or re-tuning the judge, instead of holding evaluation conditions constant and comparing the same questions under both prompt templates.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Pin the evaluation dataset, judge model version, and retrieval configuration so both templates are scored under identical conditions and results are reproducible.

A defensible promotion gate requires a controlled comparison where the prompt template is the only variable, which means running both templates over the same dataset with the judge model and retrieval configuration pinned, and inspecting per-question deltas instead of relying on averages that can mask localized regressions. Together these practices attribute any score movement to the change under test and make the decision reproducible.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Pin the evaluation dataset, judge model version, and retrieval configuration so both templates are scored under identical conditions and results are reproducible.

    Why this is correct

    Holding dataset, judge version, and retrieval settings constant ensures any score delta is attributable to the prompt template alone. Without this control, a judge upgrade or index refresh between runs confounds the comparison. Pinning also makes the promotion decision reproducible for auditors and lets the team re-run the gate after future infrastructure changes.

  • ✗

    Deploy the candidate template to a small percentage of production traffic and rely on user thumbs feedback alone as the promotion signal.

    Why it's wrong here

    Online canary feedback is valuable but sparse, slow, and self-selected, so a small traffic slice rarely accumulates enough signal to detect a modest regression before the decision deadline. Offline paired evaluation on a curated dataset provides denser, faster, and more controlled evidence; the canary is a complement, not the primary gate.

  • ✓

    Run the evaluation dataset against both the current production template and the candidate template, then compare per-question scores rather than only aggregate averages.

    Why this is correct

    A paired run over the same dataset isolates the template as the only variable, and per-question comparison reveals regressions that averages hide, such as gains on easy questions masking failures on a critical subset. This is the core of a defensible promotion gate because it attributes score movement to the change under test.

  • ✗

    Increase the judge model size until its scores on the candidate template exceed the production template's scores.

    Why it's wrong here

    Scaling the judge upward until a desired outcome appears is tuning the measurement to the conclusion. Judge capability affects score calibration, not whether the candidate genuinely regressed. This practice invalidates the comparison because the two templates would no longer be measured under identical conditions, and it invites threshold gaming rather than evidence.

  • ✗

    Evaluate only the questions where the candidate template produced different answers, skipping the rest to save judge tokens.

    Why it's wrong here

    Skipping unchanged answers biases the sample toward cases where behavior moved, which may be exactly the regressions or exactly the improvements. It also removes the baseline needed to compute a per-question delta and can hide that an unchanged answer became incorrect due to a retrieval shift. Cost savings do not justify an unrepresentative evaluation set.

About these practice questions

This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.