Courseiva
Evaluation and Monitoring →mediumMultiple Select

Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question

A RAG assistant built with Mosaic AI Agent Framework is in production. Product managers complain that answers are sometimes plausible but unsupported by the retrieved documents, and separately that the assistant sometimes ignores documents that clearly contain the answer. The team wants automated, scheduled checks in MLflow LLM Evaluation that separate these two failure modes so each can be triaged independently. Which TWO evaluation metrics should they add to the evaluation suite? (Choose two.)

⚠ Common exam trap

The trap here is reaching for performance metrics such as latency or token counts to explain answer-quality complaints, when only faithfulness and retrieval-coverage metrics can separate hallucination from missing context.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Groundedness, which judges whether the claims in the generated response are supported by the retrieved context.

The two complaints map to distinct pipeline stages. A groundedness or faithfulness metric detects responses that are plausible but unsupported by the retrieved context, catching generation-side hallucination. A context recall metric detects whether retrieval supplied the evidence at all, catching retrieval-side gaps that make the model answer without support. Scoring both in the same scheduled evaluation run separates the failure modes so each can be fixed at its own stage, while latency and token metrics carry no information about either complaint.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    P95 response time measured across all endpoint requests in the evaluation window.

    Why it's wrong here

    A tail-latency percentile describes the slowest requests in a window and is purely an operational measure of responsiveness. It provides no signal about whether claims are supported by context or whether retrieved chunks contained the needed answer. Including it in a quality suite would dilute the evaluation with performance data that cannot be acted on by prompt or retrieval changes.

  • ✗

    Latency, measured as the wall-clock time from request receipt to final token of the response.

    Why it's wrong here

    Latency is a performance metric and says nothing about whether responses are supported by retrieved documents or whether retrieval found the right evidence. Adding it would not help triage either complaint, since a fast answer can be equally ungrounded and a slow one can be perfectly faithful. It belongs in a separate performance dashboard rather than in a quality evaluation suite aimed at faithfulness and retrieval coverage.

  • ✓

    Groundedness, which judges whether the claims in the generated response are supported by the retrieved context.

    Why this is correct

    Groundedness isolates the failure mode where the answer sounds plausible but is not backed by the retrieved documents, because the judge compares each claim in the response against the supplied context. High groundedness means the generation stayed faithful to the evidence; low groundedness flags hallucinated or extrapolated content. This directly addresses the first complaint and can be scored automatically in a scheduled MLflow evaluation run.

  • ✓

    Context recall, which measures whether the retrieved context contains the information needed to answer the question.

    Why this is correct

    Context recall targets the second complaint, where the assistant ignores documents that do contain the answer. It checks whether the retrieved chunks collectively include the ground-truth answer content, exposing retrieval gaps that cause the model to respond without evidence. Together with a faithfulness metric, it cleanly separates 'context missing' from 'context present but ignored', enabling independent triage of retrieval versus generation fixes.

  • ✗

    Total token count, computed as prompt tokens plus completion tokens per request.

    Why it's wrong here

    Token count is a cost and context-window metric; it quantifies how much text was sent and generated but not whether that text was relevant or faithful. A response can be short and hallucinated or long and fully grounded, so token accounting cannot distinguish the two complaints. It is useful for capacity planning and cost control, not for separating retrieval failures from generation failures.

About these practice questions

This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.