Courseiva

Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question

A logistics company monitors a Databricks-hosted RAG assistant that answers questions about shipping regulations. Over one week, the retrieval index was rebuilt after a document refresh, and the team observes that the rate of answers flagged as unsupported by retrieved context rose sharply while retrieval relevance scores stayed flat. Which metric should they inspect first to determine whether the regression originates in the retrieval stage or the generation stage?

⚠ Common exam trap

The trap here is assuming flat relevance scores exonerate the retriever, when relevance measures topical similarity and can remain high while authoritative chunks are silently dropped after an index rebuild.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Context recall, measured against a labeled set of questions with known ground-truth source documents.

Context recall against a labeled ground-truth set directly tests whether the retriever brought back the documents that contain the answer. Because relevance measures topical similarity and can stay high even when authoritative chunks are missed, only recall can reveal that the rebuilt index now returns plausible but non-supporting passages, which would explain unsupported answers without any change in generation behavior.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The number of thumbs-up reactions collected from end users during the same week.

    Why it's wrong here

    Explicit user feedback is subjective, sparse, and lagging, and it does not localize a failure to a pipeline stage. The team already has an automated unsupported-answer signal; adding a noisy sentiment count will not reveal whether the retriever stopped surfacing authoritative chunks or the generator began ignoring them.

  • ✗

    Endpoint request latency, comparing the p95 before and after the index rebuild.

    Why it's wrong here

    Latency reflects serving performance, not answer quality or evidence coverage. The scenario already states that answer support dropped while relevance held steady, which is a quality signal unrelated to how quickly responses return. Inspecting p95 latency would consume time without distinguishing retrieval from generation as the cause.

  • ✓

    Context recall, measured against a labeled set of questions with known ground-truth source documents.

    Why this is correct

    Context recall checks whether the retriever surfaced the documents that actually contain the answer. Since relevance stayed flat but unsupported answers rose after an index rebuild, the retriever may now be returning topically similar but non-authoritative chunks, which recall against ground-truth sources detects immediately and separates retrieval failure from generation failure.

  • ✗

    Token usage per response, broken down by prompt and completion tokens on the serving endpoint.

    Why it's wrong here

    Token counts describe cost and verbosity, not whether the retrieved chunks contain the needed evidence. A post-rebuild regression in groundedness can occur with identical token profiles because the same number of chunks was retrieved, just the wrong ones. Token metrics would confirm nothing about retrieval versus generation origin.

About these practice questions

Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.