Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question
A Generative AI engineer is configuring a Mosaic AI Model Serving endpoint for a production-grade LLM. Which TWO of the following tasks are necessary to ensure effective monitoring and evaluation of the endpoint? (Choose two)
⚠ Common exam trap
Candidates often select 'Model Training' or 'Feature Store' options, missing that inference monitoring requires capturing live request/response data via Inference Tables for evaluation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable Inference Tables on the serving endpoint.
Effective monitoring requires both capturing data and evaluating that data against ground truth or automated metrics. Enabling inference tables provides the necessary raw data, while integrating with evaluation tools allows for systematic scoring of output quality. Together, these steps form the backbone of a robust monitoring strategy, ensuring that production drift is detected and quality is maintained according to business standards.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable Inference Tables on the serving endpoint.
Why this is correct
Enabling inference tables is the mandatory first step to store production request/response logs. Without this, you lack the raw material needed for post-hoc evaluation, analysis of failure modes, or detection of data drift. It creates a persistent data asset in Unity Catalog for ongoing oversight.
- ✗
Set up an automatic retraining job on the endpoint.
Why it's wrong here
Automatic retraining of large language models is rarely performed directly on a serving endpoint. LLMs require significant compute and complex fine-tuning processes outside the scope of simple model serving. Monitoring should focus on evaluation and RAG optimization rather than automated model re-training of the base LLM.
- ✓
Implement an automated evaluation pipeline to score inference logs.
Why this is correct
An automated evaluation pipeline uses the captured logs to run metrics such as faithfulness or relevance. This process transforms raw logs into actionable insights about model performance. By scoring these logs, the team can identify regressions or quality degradation before they severely impact end-user experience.
- ✗
Increase the GPU count to maximize inference throughput.
Why it's wrong here
Scaling compute resources impacts operational latency and throughput, but it does not contribute to the evaluation or monitoring of model quality. While performance is a component of monitoring, the question specifically targets the evaluation and oversight of model behavior and output quality in production.
- ✗
Delete all request logs after 24 hours to save storage.
Why it's wrong here
Deleting logs immediately destroys the historical record needed for long-term trend analysis and model evaluation. Monitoring requires a longitudinal view to detect subtle shifts in model behavior or data drift, making the retention of inference logs a critical requirement for any production AI system.
About these practice questions
One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.