Courseiva

Databricks-GenAI-Assoc · topic practice

Evaluation and Monitoring practice questions

This domain covers observability, quality measurement, and evaluation tooling for generative AI apps on Databricks. Questions test MLflow tracing and evaluation, Mosaic AI Agent Evaluation, judge-based scoring, and RAG quality metrics. Expect scenario items asking you to pick the right approach for diagnosing latency, comparing evaluation methods, or interpreting baseline and metric choices.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Evaluation and Monitoring

What the exam tests

What to know about Evaluation and Monitoring

Be able to instrument a RAG or agent pipeline with MLflow Tracing, run Mosaic AI Agent Evaluation with appropriate judges and metrics, and interpret baseline comparisons. The key is matching each metric to what it actually measures: retrieval quality versus response quality versus safety.

MLflow Tracing to capture spans across retrieval, LLM calls, and tool steps for latency debugging

Mosaic AI Agent Evaluation with LLM judges scoring correctness, groundedness, relevance, and safety

MLflow Model Evaluation APIs that compute metrics programmatically over evaluation datasets at scale

Baseline datasets in Mosaic AI Model Evaluation used to compare candidate model outputs against reference behavior

Watch out for

Common Evaluation and Monitoring exam traps

  • ▸Treating manual spot-checks or generic logging as sufficient observability instead of instrumenting the pipeline with MLflow Tracing spans.
  • ▸Confusing retrieval metrics like context relevance with answer-quality metrics like groundedness and correctness when selecting RAG evaluation metrics.
  • ▸Assuming a baseline dataset is training data or a fine-tuning input rather than a reference for comparing evaluation results.

Practice set

Evaluation and Monitoring questions

20 questions · select your answer, then reveal the explanation

An enterprise team is preparing to deploy a custom Large Language Model for a customer-facing support application on Databricks. They need to evaluate the model's responses for hallucination and groundedness using MLflow. Which specific evaluation utility should the ML Engineer utilize?

An AI engineer is configuring MLflow LLM Evaluation on Databricks to assess a Retrieval-Augmented Generation (RAG) pipeline. Which TWO evaluation metrics require both the retrieved context and the ground truth answer to compute accurately? (Select TWO)

Refer to the exhibit. The engineer notices that while inference logs are being captured, the standard Databricks monitoring dashboard for the endpoint is not populating any quality metrics. What is the most likely cause?

Exhibit

{
  "endpoint_name": "llm-chat-bot",
  "inference_tables": {
    "enabled": true,
    "catalog": "main",
    "schema": "monitoring",
    "table_name": "inference_logs"
  },
  "monitoring_enabled": false
}

An organization is using 'LLM-as-a-judge' to evaluate their RAG application. What is the primary risk associated with this approach if the judge model is not carefully selected?

Refer to the exhibit. An engineer evaluates a RAG pipeline and sees the provided results. What is the most reasonable conclusion regarding the system's performance?

Exhibit

{
  "evaluation_run": "eval-789",
  "model": "gpt-4-turbo",
  "dataset": "validation_set_v2",
  "results": {
    "faithfulness": 0.95,
    "answer_relevance": 0.60
  }
}

A security auditor requires proof that the LLM application is safe for production. Which THREE of the following monitoring/evaluation tasks would satisfy this requirement? (Choose three)

Refer to the exhibit. An LLM-powered RAG system returns these metrics. Which action should the team prioritize to improve performance?

Exhibit

{
  "evaluation_metrics": {
    "ground_truth_relevance": 0.82,
    "context_precision": 0.45,
    "faithfulness": 0.91
  }
}

A GenAI engineer runs mlflow.evaluate() with model_type="databricks-agent" on a RAG chain registered as a Unity Catalog function. The run completes and the evaluation results table is written, but the engineer needs to compare this run against the previous production baseline and programmatically gate the next deployment on it inside a Databricks job. Which action should the engineer take?

A team is building an offline regression suite for a RAG agent using MLflow LLM Evaluation on Databricks. They want the evaluation dataset to support both retrieval quality and answer quality judgments and to make future runs comparable. Which TWO dataset design choices should they make? (Choose two.)

A GenAI engineer at a retail bank has deployed a RAG assistant on Mosaic AI Model Serving. Compliance requires that any answer about account balances must be grounded strictly in the retrieved policy documents. During evaluation, the engineer notices that some answers include invented interest rates not present in any retrieved chunk. Which metric should the engineer inspect FIRST to quantify this grounding failure?

A data engineering team is deploying a RAG application using Mosaic AI Model Serving. They need to monitor the quality of the model's responses in production. Which Databricks feature should they use to capture and analyze inference data, such as requests, responses, and latency metrics?

Which of the following metrics is most effective for evaluating a RAG-based chatbot's ability to retrieve relevant context from a vector database during production monitoring?

A Generative AI engineer is configuring a Mosaic AI Model Serving endpoint for a production-grade LLM. Which TWO of the following tasks are necessary to ensure effective monitoring and evaluation of the endpoint? (Choose two)

When evaluating a generative AI model using the Mosaic AI Model Evaluation tool, what is the primary purpose of providing a 'baseline' dataset?

Which THREE of the following are common challenges when monitoring LLM applications in production that differ significantly from traditional ML monitoring? (Choose three)

When monitoring a RAG application, you notice a high discrepancy between the retrieved context and the generated answer. Which metric would specifically help identify if the model is ignoring the provided context?

A team is using MLflow to track their model evaluation runs. They want to ensure that every evaluation run is reproducible. What is the best practice to achieve this in Databricks?

When evaluating LLM outputs, which TWO metrics are most appropriate for measuring the 'quality' of a response in a RAG system? (Choose two)

Which of the following describes the 'drift' phenomenon in the context of LLM monitoring?

Why should you integrate Mosaic AI Model Evaluation with Unity Catalog?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Evaluation and Monitoring sessions

Start a Evaluation and Monitoring only practice session

Every question in these sessions is drawn from the Evaluation and Monitoring domain — nothing else.

Related practice questions

Related Databricks-GenAI-Assoc topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-GenAI-Assoc exam test about Evaluation and Monitoring?
Be able to instrument a RAG or agent pipeline with MLflow Tracing, run Mosaic AI Agent Evaluation with appropriate judges and metrics, and interpret baseline comparisons. The key is matching each metric to what it actually measures: retrieval quality versus response quality versus safety.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Evaluation and Monitoring questions in a focused session?
Yes — the session launcher on this page draws every question from the Evaluation and Monitoring domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-GenAI-Assoc topics?
Use the topic links above to move to related areas, or go back to the Databricks-GenAI-Assoc question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-GenAI-Assoc exam covers. They are not copied from any real exam or dump site.