Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
A healthcare analytics team has a RAG application that must not reveal protected health information from other patients. They want to continuously monitor production traffic on their Mosaic AI Model Serving endpoint and alert when responses contain unsafe content. Which Databricks capability should they configure to evaluate each logged request and response against safety criteria and route flagged records for review?
⚠ Common exam trap
The trap here is assuming that governance features like column masks or time travel provide content-level safety evaluation, when they only control access or retain history.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Mosaic AI Agent Evaluation with a scheduled evaluation job over inference table data
Mosaic AI Agent Evaluation supplies judges for safety and other quality dimensions, and it can be scheduled against inference table data so every production interaction is scored. Records that fail the safety criteria can be written to a separate table for review, giving the team continuous, automated monitoring rather than manual spot checks.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Unity Catalog column masks on the inference table
Why it's wrong here
Column masks restrict which users can see specific column values at query time, but they do not evaluate the semantic content of model responses or detect unsafe text. Applying a mask would hide data from analysts, not flag responses containing protected health information, so it does not satisfy the monitoring requirement.
- ✓
Mosaic AI Agent Evaluation with a scheduled evaluation job over inference table data
Why this is correct
Agent Evaluation provides built-in safety and correctness judges, and it can run as a scheduled job over data captured in inference tables. That combination continuously scores production requests and responses for unsafe content and can write flagged records to a review table, which matches the requirement to monitor live traffic without manual sampling.
- ✗
MLflow autologging during model training
Why it's wrong here
Autologging captures parameters, metrics, and artifacts while a model is being trained or when a framework's APIs are called. It operates during development, not on live serving traffic, so it cannot score each production response for safety or route flagged records for human review in this monitoring scenario.
- ✗
Delta Lake time travel on the inference table
Why it's wrong here
Time travel lets you query earlier versions of a Delta table for auditing or rollback, but it performs no semantic analysis of the stored responses. It cannot determine whether a response contains protected health information, so it fails to provide the safety evaluation and alerting the team requires.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.