Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
A team wants every MLflow evaluation run of their RAG chain to be reproducible months later, including the exact prompt template, model parameters, and retrieved context used. Which practice best ensures this reproducibility?
⚠ Common exam trap
The trap here is equating saved metric numbers with reproducibility, when replaying a run actually requires the inputs and configuration to be captured too.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Log the evaluation as an MLflow run with the model parameters, prompt template, and input dataset version recorded as run metadata and artifacts.
Reproducibility depends on capturing inputs and configuration, not just scores. Logging the run with parameters, the prompt template, and a versioned dataset as artifacts and metadata ties every result to the exact conditions that produced it. That lets anyone rerun or audit the evaluation later, even after library upgrades or dataset changes.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Log the evaluation as an MLflow run with the model parameters, prompt template, and input dataset version recorded as run metadata and artifacts.
Why this is correct
An MLflow run records parameters, metrics, tags, and artifacts tied to a run ID. Capturing the prompt template, decoding parameters, and the versioned evaluation dataset as artifacts and metadata makes it possible to reconstruct exactly what was scored later, which is the core requirement for reproducibility.
- ✗
Export the evaluation results to a CSV file stored on a personal laptop so the engineer can compare future runs manually.
Why it's wrong here
Local CSV files are not versioned, shared, or linked to run metadata, so prompt and parameter provenance is lost. This practice also risks data loss and prevents teammates or automated tooling from reproducing the run, failing the reproducibility requirement despite preserving the numeric outputs.
- ✗
Rely on the built-in default judge prompts so that no prompt configuration needs to be tracked between runs.
Why it's wrong here
Default judge prompts may change between library versions, and relying on implicit defaults means the exact scoring behavior is not pinned. Reproducibility requires recording which prompts and parameters were used, not assuming defaults remain constant across upgrades and environments.
- ✗
Store only the final aggregate metric values in a shared spreadsheet, since the underlying outputs can always be regenerated on demand.
Why it's wrong here
Aggregate metrics without inputs or configuration cannot be regenerated deterministically, especially if the dataset, prompt, or model version changes. This approach records outcomes but discards the provenance needed to replay the evaluation, so it does not satisfy reproducibility.
About these practice questions
One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.