Courseiva
Application Development →mediumMultiple Select

Databricks-GenAI-Assoc Application Development Practice Question

A team is evaluating their RAG application using Mosaic AI Model Evaluation. Which TWO metrics are most relevant for assessing the quality of the generated responses?

⚠ Common exam trap

Candidates often select traditional ML metrics like accuracy or F1-score, forgetting that LLM-based RAG evaluation requires specialized generative metrics like faithfulness and answer relevance.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Faithfulness.

Model evaluation requires quantitative metrics to measure both the accuracy of retrieved context and the quality of the final output. Faithfulness ensures the model sticks to the provided context, while answer relevance checks if the model actually addresses the user's intent. Using these metrics allows developers to iterate on their prompt engineering and retrieval strategy, ensuring the application remains accurate and useful for end-users while minimizing the risk of hallucinations.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Faithfulness.

    Why this is correct

    Faithfulness measures the degree to which the generated answer is derived exclusively from the retrieved context. This is the primary metric for detecting hallucinations, ensuring that the model does not invent information, which is a critical requirement for enterprise applications that prioritize truthfulness and reliability in their generative outputs.

  • ✗

    CPU utilization percentage.

    Why it's wrong here

    CPU utilization is a system performance metric, not a quality metric for generative models. While important for infrastructure monitoring and cost management, it provides zero insight into whether the model is producing accurate, relevant, or hallucination-free content, making it irrelevant for assessing the quality of the generative application.

  • ✓

    Answer relevance.

    Why this is correct

    Answer relevance evaluates how well the generated response directly addresses the user's question. This metric is essential for measuring the utility of the application, as a model can be factually correct but still fail to provide a useful answer, which would lead to poor user engagement and satisfaction.

  • ✗

    Network latency in milliseconds.

    Why it's wrong here

    Network latency measures the responsiveness of the infrastructure, not the quality of the model's output. While essential for user experience, it does not assess whether the model is actually performing its task of providing accurate and relevant information, and it should be treated as a separate operational concern.

  • ✗

    Total storage cost of the index.

    Why it's wrong here

    Storage cost is a financial metric related to resource management and index size. It does not provide any information regarding the accuracy or semantic quality of the responses generated by the RAG system, making it completely unsuitable for evaluating the effectiveness of the model's reasoning or the retrieval accuracy.

About these practice questions

Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.