Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
When evaluating LLM outputs, which TWO metrics are most appropriate for measuring the 'quality' of a response in a RAG system? (Choose two)
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Faithfulness
Answering quality in RAG requires looking at two distinct dimensions: the correctness of the content (faithfulness) and the relevance of the answer to the user's intent. Faithfulness ensures the model remains grounded in the provided evidence, while answer relevance captures whether the model actually addressed the specific user question. These metrics provide a balanced view, ensuring the system is both accurate and useful, which is essential for end-user satisfaction.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Faithfulness
Why this is correct
Faithfulness checks if the generated answer is derived from the retrieved context. This is the primary metric for preventing hallucinations, which are the biggest risk in RAG deployments. A faithful model stays within the bounds of its provided information, ensuring reliability and accuracy for the end user.
- ✗
Inference Latency
Why it's wrong here
Inference latency is a performance metric, not a quality metric. While it is important for the operational health of the application, it does not inform you about the factual accuracy, relevance, or grammatical correctness of the content produced by the generative model during the interaction.
- ✓
Answer Relevance
Why this is correct
Answer relevance measures how well the generated response addresses the user's initial query. Even if an answer is factually correct and faithful to the source, it must still be relevant to the user's specific intent. This metric ensures the RAG pipeline is providing a genuinely helpful user experience.
- ✗
GPU Utilization
Why it's wrong here
GPU utilization is a hardware-level monitoring metric used for infrastructure scaling and cost optimization. It provides zero insight into the quality of the generative AI outputs. Focusing on utilization while ignoring output quality is a common pitfall that overlooks the core objective of the AI application's success.
- ✗
Training Iterations
Why it's wrong here
Training iterations are a parameter of the fine-tuning process. They reflect the effort spent during development but tell you nothing about the quality of the resulting model's outputs in a production RAG environment. Monitoring must be based on inference-time results, not internal training process statistics.
About these practice questions
Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.