Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
A team runs a RAG application on Mosaic AI Model Serving and logs all requests to an inference table. Reviewers report that answers are sometimes fluent but contradict the retrieved documents. The team wants a recurring, automated check that quantifies this contradiction on production traffic and alerts when it exceeds a threshold. Which approach best fits?
⚠ Common exam trap
The trap here is treating operational telemetry such as latency or error rate as a proxy for answer quality, when semantically wrong answers still return successful responses.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Schedule a Databricks job that runs mlflow.evaluate() with the groundedness judge over recent inference-table records and writes results to a Delta table for alerting.
Groundedness evaluation compares generated answers against retrieved context, which is precisely the failure mode described. Running it as a scheduled job over recent inference-table records produces a quantified, recurring metric, and persisting results to Delta supports alerting on threshold breaches. Operational metrics and manual review cannot detect semantically unsupported but fluent responses.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Enable inference table payload logging and inspect raw prompt and response text manually each week.
Why it's wrong here
Payload logging captures the data needed for evaluation, but manual weekly inspection does not quantify contradiction at scale or produce an alertable metric. Human review is slow, inconsistent, and cannot keep pace with production volume, so it fails the requirement for automated, recurring measurement with thresholds.
- ✗
Increase the number of retrieved chunks per query so the model has more context to draw from.
Why it's wrong here
Adding chunks can introduce more contradictory or irrelevant material and dilute the context window, sometimes worsening unsupported answers. It is a retrieval tuning change, not a measurement, so it provides no quantification or alerting. The team needs to detect the problem, not guess at a fix.
- ✓
Schedule a Databricks job that runs mlflow.evaluate() with the groundedness judge over recent inference-table records and writes results to a Delta table for alerting.
Why this is correct
A scheduled job that applies the groundedness judge to recent inference-table records turns production traffic into a repeatable quality measurement. Writing scores to Delta enables threshold-based alerts and trend analysis. This directly targets the contradiction symptom because groundedness measures whether the answer is supported by retrieved context.
- ✗
Monitor the endpoint's p95 latency and error rate in the serving endpoint metrics and alert when either degrades.
Why it's wrong here
Latency and error rates are operational signals. Fluent but contradictory answers are successful HTTP responses, so they will not move these metrics. Alerting on latency or errors would miss the quality problem entirely, which is why operational monitoring alone cannot satisfy this requirement.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.