Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
A generative AI team runs a RAG chatbot whose Mosaic AI Model Serving endpoint is monitored in Unity Catalog inference tables. Over two weeks, the percentage of user questions that receive a refusal answer ('I don't have enough information') climbs from 4% to 31%, while retrieval latency and token counts stay flat. The team wants the earliest actionable signal that the retrieval corpus has gone stale rather than the prompt or model. Which monitoring signal should they inspect first?
⚠ Common exam trap
The trap here is assuming a rising refusal rate must be a prompt-engineering or model problem, when stable latency and token counts actually point to degraded retrieval relevance.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The distribution of retrieval scores for the top-k chunks returned by the Vector Search index, tracked over time in the inference table.
Refusals rising while latency and token counts hold steady isolates the problem to retrieval quality rather than the model or the serving tier. Retrieval score distributions logged to the inference table show whether top-k chunks are becoming weaker matches, which is the earliest signal that the indexed corpus no longer covers current user questions. Checking capacity, token cost, or generation time cannot distinguish stale content from a healthy pipeline.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The average generation time per completion recorded by the serving endpoint's request logs.
Why it's wrong here
Generation time measures how long the model took to produce output, which is governed by output length, model size, and hardware rather than by whether retrieved chunks were useful. The scenario says latency is flat, so this metric is unchanged and cannot be the cause. It would be relevant to performance tuning, but it gives no visibility into whether retrieval returned relevant documents.
- ✗
The endpoint's provisioned concurrency utilization and queue depth metrics from the serving endpoint logs.
Why it's wrong here
Concurrency and queue depth describe capacity and throughput, not answer quality. The scenario explicitly states retrieval latency and token counts stayed flat, so the endpoint is neither saturating nor slowing down. Rising refusals with stable latency cannot be explained by capacity pressure; these metrics would only help if requests were timing out or being throttled before retrieval completed.
- ✓
The distribution of retrieval scores for the top-k chunks returned by the Vector Search index, tracked over time in the inference table.
Why this is correct
A rising refusal rate with unchanged latency and token counts points at the retrieval stage, not generation. If top-k similarity scores drift downward or their spread collapses, the index is returning progressively weaker matches for the same query mix, which is the earliest measurable evidence that the corpus has gone stale. Inspecting score distributions in the logged inference table isolates retrieval quality before blaming prompt or model changes.
- ✗
The count of prompt tokens consumed per request, compared week over week in the billing usage table.
Why it's wrong here
Prompt token volume shows how much context was assembled and how much the request costs, but it does not reveal whether that context was relevant. The scenario already states token counts are flat, so a billing comparison adds no discriminating information. Token accounting is a cost and sizing signal, not a retrieval-quality signal, and cannot explain an eightfold increase in refusals.
About these practice questions
One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.