A RAG assistant built with Mosaic AI Agent Framework is in production. Product managers complain that answers are sometimes plausible but unsupported by the retrieved documents, and separately that the assistant sometimes ignores documents that clearly contain the answer. The team wants automated, scheduled checks in MLflow LLM Evaluation that separate these two failure modes so each can be triaged independently. Which TWO evaluation metrics should they add to the evaluation suite? (Choose two.)
Groundedness isolates the failure mode where the answer sounds plausible but is not backed by the retrieved documents, because the judge compares each claim in the response against the supplied context. High groundedness means the generation stayed faithful to the evidence; low groundedness flags hallucinated or extrapolated content. This directly addresses the first complaint and can be scored automatically in a scheduled MLflow evaluation run.
Why this answer
The two complaints map to distinct pipeline stages. A groundedness or faithfulness metric detects responses that are plausible but unsupported by the retrieved context, catching generation-side hallucination. A context recall metric detects whether retrieval supplied the evidence at all, catching retrieval-side gaps that make the model answer without support.
Scoring both in the same scheduled evaluation run separates the failure modes so each can be fixed at its own stage, while latency and token metrics carry no information about either complaint.
Exam trap
The trap here is reaching for performance metrics such as latency or token counts to explain answer-quality complaints, when only faithfulness and retrieval-coverage metrics can separate hallucination from missing context.