Courseiva

Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question

A Generative AI engineer is setting up MLflow LLM Evaluation for a RAG pipeline on Databricks that answers questions about product manuals. The team wants to evaluate both the retrieval stage and the generation stage in a single evaluation run. Which TWO metric groups should they include in the evaluation configuration? (Choose two.)

⚠ Common exam trap

The trap here is mixing infrastructure or training telemetry into an evaluation run, when retrieval and generation quality metrics are the only ones that measure pipeline output.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Retrieval metrics such as retrieval_relevance and retrieval_groundedness

A complete RAG evaluation run needs metrics for both stages. Retrieval metrics score the chunks returned by the retriever, while generation metrics such as groundedness and relevance score the final answer against the question and context. Together they let the team pinpoint whether a quality issue originates in retrieval or in generation.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    System metrics such as CPU utilization and memory consumption of the serving endpoint

    Why it's wrong here

    CPU and memory utilization describe infrastructure health, not the quality of retrieval or generation. While useful for capacity planning, they do not measure whether the right documents were retrieved or whether the answer is grounded, so they do not belong in an evaluation run focused on pipeline quality.

  • ✓

    Retrieval metrics such as retrieval_relevance and retrieval_groundedness

    Why this is correct

    Retrieval metrics score the chunks returned by the retriever against the question, so they directly evaluate the retrieval stage. Including them in the same run lets the team see whether the right passages were fetched before the generator produced its answer, which is necessary to attribute quality problems to retrieval rather than generation.

  • ✗

    Training metrics such as loss and learning rate from the fine-tuning job

    Why it's wrong here

    Loss and learning rate belong to model training and describe optimization progress, not the runtime behavior of a RAG pipeline. They cannot indicate whether retrieval returned relevant chunks or whether responses are grounded, so including them would not help evaluate the product-manual assistant's quality.

  • ✗

    Data freshness metrics such as table update timestamps in Unity Catalog

    Why it's wrong here

    Table update timestamps indicate when source data changed, which is useful for lineage and refresh scheduling but not for scoring retrieval or generation quality. They say nothing about whether the retriever surfaced the correct manual sections or whether the answer was supported, so they are not evaluation metrics for this run.

  • ✓

    Generation metrics such as groundedness and relevance

    Why this is correct

    Generation metrics judge the final response against the question and the retrieved context, covering the generation stage. Adding them alongside retrieval metrics gives a complete picture of the pipeline, so the team can tell whether a poor answer came from missing context or from the generator mishandling context that was present.

About these practice questions

One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.