Courseiva
Evaluation and Monitoring →mediumMultiple Choice

Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question

A GenAI engineering team has deployed a customer-support RAG chain on Databricks and registered it in Unity Catalog. They now want MLflow 3 to automatically score every production request for groundedness and relevance without writing custom scoring code, and to persist those assessments against the logged traces. Which approach should they use?

⚠ Common exam trap

The trap here is assuming that any Databricks observability surface, such as system tables or dashboards, can produce semantic quality scores, when only MLflow scorers attached to the agent compute groundedness and relevance.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Attach MLflow scorers to the deployed agent so that assessments are computed on live traces and stored in the MLflow experiment.

MLflow 3 production monitoring is designed for exactly this need: built-in scorers are attached to a deployed agent rather than invoked manually, and assessments are recorded on live traces inside the MLflow experiment. This removes custom scoring code, keeps quality signals tied to each request, and enables continuous monitoring without batch jobs or dashboard workarounds.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Configure the endpoint's autoscaling policy to collect evaluation metrics alongside throughput metrics.

    Why it's wrong here

    Autoscaling policies control replica counts based on traffic, not quality evaluation. They have no mechanism to compute groundedness or relevance and do not write assessments to traces. Confusing scaling configuration with evaluation configuration is a category error, and this option will not deliver any quality scoring for the RAG chain.

  • ✗

    Create a Databricks SQL dashboard that re-runs mlflow.evaluate() on the inference table every five minutes.

    Why it's wrong here

    A SQL dashboard is a visualization layer, not an evaluator. Re-running mlflow.evaluate() on the inference table repeatedly is a batch pattern that adds latency and cost, and it does not attach per-request assessments to the individual traces. It also requires custom orchestration, which the team explicitly wants to avoid.

  • ✗

    Enable the system table system.serving.endpoint_usage and query it for groundedness scores.

    Why it's wrong here

    system.serving.endpoint_usage contains operational metrics such as request counts, latency, and error rates for serving endpoints. It does not contain semantic quality scores like groundedness or relevance. Querying it cannot produce the per-trace quality assessments the team needs, making this approach ineffective for their requirement.

  • ✓

    Attach MLflow scorers to the deployed agent so that assessments are computed on live traces and stored in the MLflow experiment.

    Why this is correct

    MLflow 3 monitoring lets you attach built-in scorers such as groundedness and relevance directly to a deployed agent. The scorers run asynchronously on live traces and write assessment records back to the MLflow experiment, so no custom scoring code is required and results stay linked to each trace.

About these practice questions

Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.