Courseiva

CCNA Evaluation and Monitoring Questions

54 questions · Evaluation and Monitoring · All types, answers revealed

1
Multi-Selectmedium

A RAG assistant built with Mosaic AI Agent Framework is in production. Product managers complain that answers are sometimes plausible but unsupported by the retrieved documents, and separately that the assistant sometimes ignores documents that clearly contain the answer. The team wants automated, scheduled checks in MLflow LLM Evaluation that separate these two failure modes so each can be triaged independently. Which TWO evaluation metrics should they add to the evaluation suite? (Choose two.)

Select 2 answers
A.P95 response time measured across all endpoint requests in the evaluation window.
B.Latency, measured as the wall-clock time from request receipt to final token of the response.
C.Groundedness, which judges whether the claims in the generated response are supported by the retrieved context.
D.Context recall, which measures whether the retrieved context contains the information needed to answer the question.
E.Total token count, computed as prompt tokens plus completion tokens per request.
AnswersC, D

Groundedness isolates the failure mode where the answer sounds plausible but is not backed by the retrieved documents, because the judge compares each claim in the response against the supplied context. High groundedness means the generation stayed faithful to the evidence; low groundedness flags hallucinated or extrapolated content. This directly addresses the first complaint and can be scored automatically in a scheduled MLflow evaluation run.

Why this answer

The two complaints map to distinct pipeline stages. A groundedness or faithfulness metric detects responses that are plausible but unsupported by the retrieved context, catching generation-side hallucination. A context recall metric detects whether retrieval supplied the evidence at all, catching retrieval-side gaps that make the model answer without support.

Scoring both in the same scheduled evaluation run separates the failure modes so each can be fixed at its own stage, while latency and token metrics carry no information about either complaint.

Exam trap

The trap here is reaching for performance metrics such as latency or token counts to explain answer-quality complaints, when only faithfulness and retrieval-coverage metrics can separate hallucination from missing context.

2
MCQeasy

A team deploys a customer-support assistant on a Mosaic AI Model Serving endpoint and enables inference table logging to Unity Catalog. Compliance requires that every production response be traceable back to the exact request, the retrieved context, and the model version that produced it, and that reviewers can query this history with SQL months later. Which capability satisfies this requirement?

A.The endpoint's request and response payloads written to a Unity Catalog inference table, queried with SQL and joined to model version metadata.
B.Real-time dashboards built on the endpoint's built-in metrics such as request count and error rate.
C.MLflow experiment runs recorded during offline evaluation of the assistant before release.
D.Client-side application logs streamed to a log analytics workspace and retained for thirty days.
AnswerA

Inference tables persist the full request, retrieved context, response, and associated metadata as Delta tables governed by Unity Catalog, so reviewers can run SQL against historical traffic long after the fact. Because the logged rows include the served model version and timestamp, each response can be traced to its exact inputs and the model that generated it, which is precisely the auditability the compliance team requires.

Why this answer

Compliance-grade traceability demands durable, queryable records of each request, the context supplied to the model, the response, and the serving model version. Unity Catalog inference tables store exactly that as governed Delta tables, so SQL queries and joins to model version metadata answer audit questions months after the traffic occurred. Client logs, offline experiment runs, and aggregate dashboards all lack either the payload detail or the retention and governance needed.

Exam trap

The trap here is confusing operational dashboards or client-side logs with auditable records, when only governed payload-level logging preserves request, context, and model version together.

3
MCQhard

A GenAI engineer monitors a customer-facing RAG assistant hosted on Databricks. After a routine re-indexing job, groundedness scores from the MLflow LLM Evaluation job drop sharply while answer relevance stays flat. The application prompt and the LLM serving endpoint were untouched. Which conclusion is best supported by these signals?

A.The LLM judge model was silently upgraded and is scoring more harshly
B.Users are asking harder questions than before, lowering groundedness
C.The re-indexing job degraded retrieval, so the generator receives weaker supporting context
D.The generation prompt has drifted and needs to be rewritten
AnswerC

Groundedness depends on the retrieved context supporting the answer, so a sharp drop right after re-indexing implicates the retrieval pipeline. Flat relevance shows the generator still addresses the question, but now without adequate evidence, producing less grounded answers. Chunking changes, embedding mismatches, or partial index builds during re-indexing are the likely culprits to investigate first.

Why this answer

The drop begins exactly when the index was rebuilt, and relevance remains flat, which isolates the change to the evidence supplied to the generator rather than to the prompt, the LLM, or the judge. Degraded retrieval after re-indexing, from altered chunking, embedding mismatches, or an incomplete build, is the best-supported conclusion and the right place to investigate first.

Exam trap

The trap here is attributing a groundedness drop to the model or judge, when a sharp drop synchronized with re-indexing and stable relevance points to the retrieval context instead.

4
MCQmedium

A media company's RAG assistant answers questions about its streaming catalog. Users report that the assistant often returns answers that ignore the retrieved documents and instead rely on outdated information from the base model. The team wants an MLflow LLM Evaluation metric that measures how much of the final answer is actually derived from the retrieved context rather than the model's prior knowledge. Which metric best fits this need?

A.relevance
B.groundedness
C.context_sufficiency
D.answer_correctness
AnswerB

Groundedness measures the degree to which the response's statements are supported by the retrieved context. When the assistant ignores retrieved documents and falls back on outdated knowledge, its claims will not be traceable to the context, producing low groundedness scores that expose exactly the behavior the team wants to quantify.

Why this answer

Groundedness directly assesses whether the response is supported by the retrieved context, so a pattern of ignoring documents and using stale base-model knowledge will show up as low groundedness. That makes it the metric that quantifies the behavior the team is investigating in their catalog assistant.

Exam trap

The trap here is assuming that answer correctness or relevance reveals whether the model used retrieved context, when only groundedness measures the link between response and context.

5
MCQhard

A media company uses an LLM-as-a-judge evaluation pipeline on Databricks to score a summarization assistant nightly. The judge is the same model family as the assistant and was given a rubric that rewards stylistic fluency. Over two months, nightly scores climb steadily while editor spot-checks find summaries increasingly omit key facts. Which corrective action best restores the evaluation's ability to detect this regression?

A.Revise the judge rubric to weight factual coverage against the source document, add a small human-labeled calibration set, and periodically verify judge agreement with editors.
B.Increase the judge's temperature so its scores vary more between runs and regressions become statistically visible.
C.Run the judge more frequently, moving from nightly to hourly scoring, so trends are captured with finer granularity.
D.Switch the judge to the largest available frontier model and remove the rubric so the judge can apply its own judgment freely.
AnswerA

The judge is self-preferring a fluent sibling model and rewarding style over substance, so the rubric must explicitly score factual coverage and be validated against human labels. A calibration set and periodic agreement checks detect when the judge drifts away from editorial standards, which is the only way the nightly signal can be trusted again.

Why this answer

The rising scores reflect a judge that rewards fluency because it shares a model family with the assistant and was given a style-oriented rubric, so the instrument itself is blind to factual omission. Restoring detection requires redefining the rubric around factual coverage of the source, anchoring it with human-labeled examples, and monitoring judge-human agreement so drift in the judge is caught before it masks another regression.

Exam trap

The trap here is treating rising judge scores as improving quality when the judge and the assistant share a family and the rubric rewards fluency, so the metric drifts upward while factual omissions grow.

6
MCQhard

A team uses an LLM judge to score their RAG agent on Databricks and sees scores that fluctuate by several points between identical evaluation runs. They need more stable, reproducible quality signals for release gating. Which action best addresses the root cause?

A.Increase the number of evaluation examples until the average stabilizes across runs.
B.Average the scores from three different judge models and use the mean for gating.
C.Switch from an LLM judge to BLEU or ROUGE string-overlap metrics.
D.Pin the judge model version and set its temperature to zero, then validate the judge against a human-labeled sample.
AnswerD

Judge variance often comes from sampling temperature and model version changes. Setting temperature to zero and pinning the judge version makes scoring deterministic and reproducible, while human-labeled validation confirms the judge's scores are trustworthy before they gate releases.

Why this answer

Fluctuating judge scores usually stem from nondeterministic sampling and drifting judge model versions. Pinning the judge version and setting temperature to zero makes each scoring pass reproducible, and validating the judge against human labels confirms that the stabilized scores are also accurate enough to gate releases reliably.

Exam trap

The trap here is chasing aggregate stability by enlarging the dataset or blending judges, when the per-example variance originates from judge sampling settings and version drift.

7
MCQhard

A Generative AI engineer is using MLflow Tracing to monitor a RAG application deployed on Mosaic AI Model Serving. They want to capture the retrieved documents, the final prompt, and the model's response for each request to debug a quality issue. Which approach should they use to ensure all three are logged in a single trace?

A.Instrument the application with the MLflow tracing SDK, creating spans for retrieval, prompt construction, and generation within the same trace context.
B.Use the `mlflow.evaluate()` API with a custom evaluator that logs retrieved documents, prompt, and response as metrics.
C.Enable inference table logging on the serving endpoint and query the payload column for the retrieved documents and response.
D.Configure the serving endpoint to log all requests to a Delta table and then join with a separate table of retrieved documents using a request ID.
AnswerA

MLflow Tracing allows manual instrumentation with spans that share a trace context. By wrapping retrieval, prompt construction, and generation in spans, all three are captured in one trace. This provides end-to-end visibility, which is essential for debugging RAG quality issues. It works with Mosaic AI Model Serving and integrates with the MLflow UI for inspection.

Why this answer

To capture retrieved documents, the final prompt, and the model response in a single trace, the engineer should use MLflow Tracing with manual instrumentation. Spans for each step share a trace context, providing a unified view. This is the native Databricks approach for debugging RAG applications and integrates with Mosaic AI Model Serving for production monitoring.

Exam trap

The trap here is assuming that inference table logging automatically includes intermediate retrieval steps and the exact prompt, when it only captures request and response payloads.

8
MCQmedium

A team has deployed a RAG chatbot on Databricks and enabled inference table logging. They want to set up automated monitoring to detect when the average response length increases significantly compared to the baseline. Which Databricks feature should they use to create this monitor?

A.Mosaic AI Model Serving endpoint logs analyzed with the `mlflow.evaluate()` API on a schedule.
B.Databricks SQL alerts on a scheduled query that calculates the average response length from the inference table.
C.Lakehouse Monitoring for the inference table, with a custom metric for response length.
D.MLflow Model Registry webhooks to trigger a retraining pipeline when response length changes.
AnswerC

Lakehouse Monitoring can track statistical properties of Delta tables, including inference tables. By defining a custom metric for response length, the team can monitor drift and receive alerts when the average deviates from the baseline. This is the native Databricks solution for monitoring data quality and model performance over time. It integrates with the inference table without additional ETL.

Why this answer

Lakehouse Monitoring is designed to monitor Delta tables, including inference tables, for data quality and drift. By creating a monitor with a custom metric for response length, the team can automatically detect significant changes from the baseline and receive alerts. This is the most direct and integrated approach on Databricks for this monitoring requirement.

Exam trap

The trap here is confusing batch evaluation tools like `mlflow.evaluate()` with continuous monitoring features like Lakehouse Monitoring, which are purpose-built for tracking drift in production tables.

9
MCQmedium

When monitoring a RAG application, you notice a high discrepancy between the retrieved context and the generated answer. Which metric would specifically help identify if the model is ignoring the provided context?

A.Context Precision
B.Faithfulness
C.Retrieval Recall
D.Semantic Similarity
AnswerB

Faithfulness specifically assesses whether the answer is logically derived from the provided context. High faithfulness indicates the model is respecting the context; low faithfulness suggests the model is generating responses based on its own training data, which leads to hallucinations and incorrect information in RAG systems.

Why this answer

Faithfulness measures whether the generated answer is derived exclusively from the retrieved context. If a model generates information not supported by the source text, it is 'hallucinating' or ignoring the context. Monitoring faithfulness is crucial for RAG systems because it directly detects when the model drifts away from the ground truth provided by the internal knowledge base, which is the primary value proposition of a RAG architecture.

Exam trap

Candidates often confuse faithfulness with context relevance, failing to realize that faithfulness specifically evaluates whether the answer is derived directly from the retrieved context.

10
Multi-Selectmedium

A fintech company runs a customer-support RAG assistant on Databricks. Before promoting a new prompt template from staging to production, the ML team must demonstrate that the change does not regress answer quality. Which TWO evaluation practices should they apply in Mosaic AI Agent Evaluation to make this promotion decision defensible? (Choose two.)

Select 2 answers
A.Pin the evaluation dataset, judge model version, and retrieval configuration so both templates are scored under identical conditions and results are reproducible.
B.Deploy the candidate template to a small percentage of production traffic and rely on user thumbs feedback alone as the promotion signal.
C.Run the evaluation dataset against both the current production template and the candidate template, then compare per-question scores rather than only aggregate averages.
D.Increase the judge model size until its scores on the candidate template exceed the production template's scores.
E.Evaluate only the questions where the candidate template produced different answers, skipping the rest to save judge tokens.
AnswersA, C

Holding dataset, judge version, and retrieval settings constant ensures any score delta is attributable to the prompt template alone. Without this control, a judge upgrade or index refresh between runs confounds the comparison. Pinning also makes the promotion decision reproducible for auditors and lets the team re-run the gate after future infrastructure changes.

Why this answer

A defensible promotion gate requires a controlled comparison where the prompt template is the only variable, which means running both templates over the same dataset with the judge model and retrieval configuration pinned, and inspecting per-question deltas instead of relying on averages that can mask localized regressions. Together these practices attribute any score movement to the change under test and make the decision reproducible.

Exam trap

The trap here is chasing a passing aggregate score by enlarging or re-tuning the judge, instead of holding evaluation conditions constant and comparing the same questions under both prompt templates.

11
MCQhard

A team runs a RAG application on Mosaic AI Model Serving and logs all requests to an inference table. Reviewers report that answers are sometimes fluent but contradict the retrieved documents. The team wants a recurring, automated check that quantifies this contradiction on production traffic and alerts when it exceeds a threshold. Which approach best fits?

A.Enable inference table payload logging and inspect raw prompt and response text manually each week.
B.Increase the number of retrieved chunks per query so the model has more context to draw from.
C.Schedule a Databricks job that runs mlflow.evaluate() with the groundedness judge over recent inference-table records and writes results to a Delta table for alerting.
D.Monitor the endpoint's p95 latency and error rate in the serving endpoint metrics and alert when either degrades.
AnswerC

A scheduled job that applies the groundedness judge to recent inference-table records turns production traffic into a repeatable quality measurement. Writing scores to Delta enables threshold-based alerts and trend analysis. This directly targets the contradiction symptom because groundedness measures whether the answer is supported by retrieved context.

Why this answer

Groundedness evaluation compares generated answers against retrieved context, which is precisely the failure mode described. Running it as a scheduled job over recent inference-table records produces a quantified, recurring metric, and persisting results to Delta supports alerting on threshold breaches. Operational metrics and manual review cannot detect semantically unsupported but fluent responses.

Exam trap

The trap here is treating operational telemetry such as latency or error rate as a proxy for answer quality, when semantically wrong answers still return successful responses.

12
MCQeasy

A team wants every MLflow evaluation run of their RAG chain to be reproducible months later, including the exact prompt template, model parameters, and retrieved context used. Which practice best ensures this reproducibility?

A.Log the evaluation as an MLflow run with the model parameters, prompt template, and input dataset version recorded as run metadata and artifacts.
B.Export the evaluation results to a CSV file stored on a personal laptop so the engineer can compare future runs manually.
C.Rely on the built-in default judge prompts so that no prompt configuration needs to be tracked between runs.
D.Store only the final aggregate metric values in a shared spreadsheet, since the underlying outputs can always be regenerated on demand.
AnswerA

An MLflow run records parameters, metrics, tags, and artifacts tied to a run ID. Capturing the prompt template, decoding parameters, and the versioned evaluation dataset as artifacts and metadata makes it possible to reconstruct exactly what was scored later, which is the core requirement for reproducibility.

Why this answer

Reproducibility depends on capturing inputs and configuration, not just scores. Logging the run with parameters, the prompt template, and a versioned dataset as artifacts and metadata ties every result to the exact conditions that produced it. That lets anyone rerun or audit the evaluation later, even after library upgrades or dataset changes.

Exam trap

The trap here is equating saved metric numbers with reproducibility, when replaying a run actually requires the inputs and configuration to be captured too.

13
MCQeasy

A Generative AI engineer at a retail company has deployed a RAG chatbot on a Databricks Mosaic AI Model Serving endpoint. The team wants to capture per-request evaluation data, including the user's question, the retrieved context chunks, and the model's response, so they can analyze response quality over time. Which Databricks feature should they use to collect this data for downstream evaluation?

A.MLflow Model Registry
B.Delta Live Tables expectations
C.Inference tables
D.Unity Catalog lineage
AnswerC

Inference tables automatically log the request payload and the response payload for each call to a Mosaic AI Model Serving endpoint, which includes the user's question, retrieved context, and generated answer when the endpoint serves a RAG chain. This gives the team the raw per-request data needed to run evaluation and monitoring jobs later, without changing the client application.

Why this answer

Inference tables are the Databricks mechanism that logs request and response payloads for Mosaic AI Model Serving endpoints. Because the logged payloads include the user question, retrieved context, and generated answer, the team gains the raw material needed to run MLflow LLM Evaluation or custom monitoring jobs later against real production traffic.

Exam trap

The trap here is assuming that any Databricks governance or tracking feature automatically captures serving payloads, when only inference tables log request and response content for an endpoint.

14
MCQmedium

A team runs a RAG chatbot whose MLflow evaluation with the built-in groundedness judge previously scored well. After they swap the retriever for a new embedding model, groundedness scores drop sharply even though the generator model and prompt are unchanged. They confirm the judge model itself is unchanged. Which action should they take FIRST to diagnose the regression?

A.Roll back the embedding model change immediately and re-run the evaluation to confirm the previous score returns.
B.Inspect the retrieved context chunks logged per evaluation row to verify whether the new retriever is returning passages that no longer support the generated claims.
C.Retrain the judge model on a fresh labeled dataset so its scoring distribution aligns with the new embedding model.
D.Increase the temperature of the generator model to encourage more diverse phrasing that the groundedness judge can score more reliably.
AnswerB

Groundedness judges whether each claim in the response is supported by the retrieved context. If the retriever changed, the logged retrieved chunks are the first artifact to inspect, since they directly feed the judge. Reviewing per-row traces confirms whether the regression originates in retrieval rather than generation.

Why this answer

Groundedness measures whether the answer is supported by the retrieved context, so a change to retrieval is the most likely culprit. Inspecting the logged retrieved chunks per evaluation row reveals whether the new embedding model returns less relevant or contradictory passages, which the judge then flags as unsupported claims. This isolates retrieval as the cause before any remediation.

Exam trap

The trap here is assuming a groundedness drop always indicates a hallucinating generator, when the metric is just as sensitive to degraded retrieved context.

15
MCQhard

A team uses MLflow LLM Evaluation with an LLM judge to score a summarization agent. On a 400-row golden dataset the judge marks 96% of summaries as relevant. Spot-checking reveals the judge approves nearly every summary whenever the summary is fluent, even when key facts are missing. The team wants a defensible quality signal before approving a release. Which action best addresses this judge weakness?

A.Replace the relevance metric with a cheaper lexical overlap score such as ROUGE against the reference summaries.
B.Expand the golden dataset from 400 rows to 4,000 rows and rerun the same relevance metric.
C.Add a small human-labeled calibration set and compare the judge's verdicts against human labels before trusting the automated relevance scores.
D.Raise the judge's temperature so it explores a wider range of judgments and produces a more nuanced relevance distribution.
AnswerC

The judge exhibits a systematic bias toward fluency, so its scores cannot be trusted at face value. Labeling a representative subset by hand and measuring agreement between judge and human verdicts quantifies that bias and reveals which cases the judge mishandles. Once agreement is characterized, the team can report calibrated results, adjust thresholds, or supplement the judge with a fact-coverage metric, making the release decision defensible rather than resting on inflated automated scores.

Why this answer

An automated judge that approves nearly every fluent summary is a biased measurement instrument, and the fix is to calibrate it rather than to change its sampling or the dataset size. Human-labeling a representative subset and measuring judge-human agreement exposes the leniency and tells the team how much to trust the scores. Only then can thresholds be set or a complementary fact-coverage check be added, producing a release signal that can be defended.

Exam trap

The trap here is trying to fix a systematically lenient judge by enlarging the dataset or lowering rigor, when the real need is to calibrate the judge against human labels.

16
MCQmedium

Which of the following metrics is most effective for evaluating a RAG-based chatbot's ability to retrieve relevant context from a vector database during production monitoring?

A.Model Training Loss
B.Context Relevance
C.Inference Latency
D.Parameter Count
AnswerB

Context relevance measures how much of the retrieved information is actually necessary to answer the user query. Low scores indicate a failure in the embedding or vector search configuration, providing actionable insights for tuning the retriever component, which is a critical part of RAG evaluation.

Why this answer

Context relevance and faithfulness are key RAG metrics. Specifically, evaluating the retrieved context's alignment with the user query is essential for identifying retrieval failures. If the context is irrelevant, the model cannot generate a correct answer even with high-quality generative capabilities.

Monitoring this metric helps distinguish between errors caused by the retriever versus errors caused by the generator model during the production lifecycle.

Exam trap

Candidates frequently select generation metrics like faithfulness when asked specifically about evaluating how well the vector database retrieves information for the user query.

17
Multi-Selecthard

A Generative AI engineer is setting up MLflow LLM Evaluation for a RAG pipeline on Databricks that answers questions about product manuals. The team wants to evaluate both the retrieval stage and the generation stage in a single evaluation run. Which TWO metric groups should they include in the evaluation configuration? (Choose two.)

Select 2 answers
A.System metrics such as CPU utilization and memory consumption of the serving endpoint
B.Retrieval metrics such as retrieval_relevance and retrieval_groundedness
C.Training metrics such as loss and learning rate from the fine-tuning job
D.Data freshness metrics such as table update timestamps in Unity Catalog
E.Generation metrics such as groundedness and relevance
AnswersB, E

Retrieval metrics score the chunks returned by the retriever against the question, so they directly evaluate the retrieval stage. Including them in the same run lets the team see whether the right passages were fetched before the generator produced its answer, which is necessary to attribute quality problems to retrieval rather than generation.

Why this answer

A complete RAG evaluation run needs metrics for both stages. Retrieval metrics score the chunks returned by the retriever, while generation metrics such as groundedness and relevance score the final answer against the question and context. Together they let the team pinpoint whether a quality issue originates in retrieval or in generation.

Exam trap

The trap here is mixing infrastructure or training telemetry into an evaluation run, when retrieval and generation quality metrics are the only ones that measure pipeline output.

18
MCQmedium

A team runs an LLM-as-a-judge evaluation on Databricks using the built-in `mlflow.evaluate()` with `model_type="databricks-agent"`. They notice that the judge model, a serving endpoint, is producing scores that are consistently inflated compared to human review. Which configuration change should the team make first to improve the reliability of the evaluation?

A.Switch the judge to a different, stronger model endpoint and calibrate it using a small human-labeled dataset with known ground-truth scores.
B.Reduce the number of evaluation examples to only those where the judge and human agree, then report the average score.
C.Increase the temperature parameter of the judge model to 1.0 to encourage more diverse scoring.
D.Replace the LLM judge with a BLEU score comparison against a reference answer for every question.
AnswerA

Using a more capable judge model and calibrating it against human-labeled examples is the recommended way to reduce systematic bias. Calibration with ground-truth scores helps detect and correct consistent over- or under-scoring. This approach aligns with MLflow's guidance to validate LLM judges before relying on their outputs in production monitoring.

Why this answer

LLM judges can exhibit systematic bias, such as consistently inflated scores. The most effective first step is to use a stronger judge model and calibrate it against human-labeled data with known scores. This improves reliability and aligns with Databricks recommendations for validating judges before using them in production monitoring.

Simply changing temperature or filtering data does not address the root cause.

Exam trap

The trap here is assuming that increasing the judge model's temperature will make it more objective, when in fact it adds randomness and worsens reproducibility without fixing bias.

19
MCQeasy

A logistics company monitors a Databricks-hosted RAG assistant that answers questions about shipping regulations. Over one week, the retrieval index was rebuilt after a document refresh, and the team observes that the rate of answers flagged as unsupported by retrieved context rose sharply while retrieval relevance scores stayed flat. Which metric should they inspect first to determine whether the regression originates in the retrieval stage or the generation stage?

A.The number of thumbs-up reactions collected from end users during the same week.
B.Endpoint request latency, comparing the p95 before and after the index rebuild.
C.Context recall, measured against a labeled set of questions with known ground-truth source documents.
D.Token usage per response, broken down by prompt and completion tokens on the serving endpoint.
AnswerC

Context recall checks whether the retriever surfaced the documents that actually contain the answer. Since relevance stayed flat but unsupported answers rose after an index rebuild, the retriever may now be returning topically similar but non-authoritative chunks, which recall against ground-truth sources detects immediately and separates retrieval failure from generation failure.

Why this answer

Context recall against a labeled ground-truth set directly tests whether the retriever brought back the documents that contain the answer. Because relevance measures topical similarity and can stay high even when authoritative chunks are missed, only recall can reveal that the rebuilt index now returns plausible but non-supporting passages, which would explain unsupported answers without any change in generation behavior.

Exam trap

The trap here is assuming flat relevance scores exonerate the retriever, when relevance measures topical similarity and can remain high while authoritative chunks are silently dropped after an index rebuild.

20
MCQmedium

An enterprise LLM application on Databricks is experiencing high latency and inconsistent responses. Which approach best enables observability to identify the root cause of these performance bottlenecks within the LLM pipeline?

A.Enable standard Spark UI logs to monitor cluster resource utilization during inference.
B.Review the Databricks SQL query history to identify slow retrieval operations.
C.Use MLflow Tracing to record inputs, outputs, and execution duration of each step in the chain.
D.Deploy an external monitoring agent to scan the model server network traffic.
AnswerC

MLflow Tracing captures end-to-end execution details for LLM workflows. It allows developers to visualize the entire dependency graph and latency of each step, enabling precise identification of bottlenecks in chains, prompt templates, or vector search lookups.

Why this answer

Integrating MLflow Tracing provides granular visibility into the execution flow, including model inputs, outputs, and latency per step. This is critical for diagnosing complex chains where intermediate calls may be stalling. By instrumenting the code, engineers can identify which specific retrieval step or prompt generation block is contributing to the high latency, allowing for targeted optimization of the RAG pipeline components.

Exam trap

Candidates often choose standard cluster log analysis or simple end-to-end latency metrics instead of MLflow Tracing to inspect individual steps within an LLM chain.

21
MCQmedium

A financial services team runs a RAG assistant on Databricks that answers questions about internal policy documents. During a review, the team finds that for many questions the answer is factually correct but cites a document that does not actually contain the supporting statement. They want an MLflow LLM Evaluation metric that specifically detects when the response is not supported by the retrieved context. Which metric should they add to their evaluation run?

A.answer_similarity
B.groundedness
C.retrieval_precision
D.toxicity
AnswerB

The groundedness metric in MLflow LLM Evaluation judges whether the claims in the response are supported by the retrieved context. In this scenario the answers are correct but not backed by the cited documents, which is exactly the failure mode groundedness is designed to surface, making it the right metric to add to the evaluation run.

Why this answer

Groundedness evaluates whether the response's claims can be traced back to the retrieved context, which is precisely the gap identified in this review. Adding it to the MLflow LLM Evaluation run gives the team a metric that flags answers which sound correct but are not actually supported by the cited policy documents.

Exam trap

The trap here is confusing factual correctness with grounding, assuming that an accurate answer must also be supported by the retrieved context.

22
Multi-Selecthard

A Generative AI engineer is configuring a Mosaic AI Model Serving endpoint for a production-grade LLM. Which TWO of the following tasks are necessary to ensure effective monitoring and evaluation of the endpoint? (Choose two)

Select 2 answers
A.Enable Inference Tables on the serving endpoint.
B.Set up an automatic retraining job on the endpoint.
C.Implement an automated evaluation pipeline to score inference logs.
D.Increase the GPU count to maximize inference throughput.
E.Delete all request logs after 24 hours to save storage.
AnswersA, C

Enabling inference tables is the mandatory first step to store production request/response logs. Without this, you lack the raw material needed for post-hoc evaluation, analysis of failure modes, or detection of data drift. It creates a persistent data asset in Unity Catalog for ongoing oversight.

Why this answer

Effective monitoring requires both capturing data and evaluating that data against ground truth or automated metrics. Enabling inference tables provides the necessary raw data, while integrating with evaluation tools allows for systematic scoring of output quality. Together, these steps form the backbone of a robust monitoring strategy, ensuring that production drift is detected and quality is maintained according to business standards.

Exam trap

Candidates often select 'Model Training' or 'Feature Store' options, missing that inference monitoring requires capturing live request/response data via Inference Tables for evaluation.

23
MCQmedium

A data engineering team is deploying a RAG application using Mosaic AI Model Serving. They need to monitor the quality of the model's responses in production. Which Databricks feature should they use to capture and analyze inference data, such as requests, responses, and latency metrics?

A.Databricks SQL Alerts
B.Mosaic AI Inference Tables
C.Delta Live Tables Audit Logs
D.MLflow Experiment Tracking
AnswerB

Inference tables automatically log inference requests and responses to a Delta table in Unity Catalog. This enables seamless integration with monitoring tools for assessing model performance, latency, and throughput. It is the standardized method for capturing production data required to perform comprehensive model evaluation and drift detection.

Why this answer

Mosaic AI Model Serving provides built-in inference tables to automatically capture request and response logs. By enabling these tables, engineers can export data to a Unity Catalog table for analysis. This is critical for monitoring performance, data drift, and model quality over time.

Without this feature, teams lack the visibility required for production-grade LLM governance and continuous improvement cycles within the Databricks ecosystem.

Exam trap

Candidates often confuse inference tables with MLflow experiments or standard Delta tables created manually. They forget that Mosaic AI Inference Tables are a native, built-in feature specifically for capturing model serving request and response logs.

24
MCQmedium

A team's RAG evaluation shows high context recall but low answer correctness. Retrieved documents contain the needed facts, yet the generated answers frequently contradict them. Which single metric should they examine next to pinpoint whether the generator is ignoring or misusing the retrieved context?

A.Retrieval latency, to determine whether slow document fetches cause the generator to fall back on parametric memory.
B.Token count of the retrieved chunks, to verify whether the context window is overflowing and truncating documents.
C.Context precision, to check whether the most relevant chunks are ranked at the top of the retrieved list.
D.Groundedness, to check whether the claims in the generated answer are actually supported by the retrieved context.
AnswerD

Groundedness evaluates whether each statement in the answer is entailed by the retrieved context. High context recall with low correctness and low groundedness points to a generator that is not faithfully using the provided passages, whereas high groundedness would suggest a different failure mode such as a flawed ground-truth label.

Why this answer

With context recall high, retrieval supplied the necessary facts, so the failure lies downstream in generation. Groundedness directly measures whether the answer's claims are supported by the retrieved context, so a low groundedness score localizes the problem to the generator ignoring or misreading context. That distinguishes a generation-faithfulness issue from a retrieval-ranking or labeling issue.

Exam trap

The trap here is reaching for another retrieval metric when retrieval already proved sufficient via high context recall, so the investigation should move downstream.

25
MCQmedium

A GenAI engineering team has deployed a customer-support RAG chain on Databricks and registered it in Unity Catalog. They now want MLflow 3 to automatically score every production request for groundedness and relevance without writing custom scoring code, and to persist those assessments against the logged traces. Which approach should they use?

A.Configure the endpoint's autoscaling policy to collect evaluation metrics alongside throughput metrics.
B.Create a Databricks SQL dashboard that re-runs mlflow.evaluate() on the inference table every five minutes.
C.Enable the system table system.serving.endpoint_usage and query it for groundedness scores.
D.Attach MLflow scorers to the deployed agent so that assessments are computed on live traces and stored in the MLflow experiment.
AnswerD

MLflow 3 monitoring lets you attach built-in scorers such as groundedness and relevance directly to a deployed agent. The scorers run asynchronously on live traces and write assessment records back to the MLflow experiment, so no custom scoring code is required and results stay linked to each trace.

Why this answer

MLflow 3 production monitoring is designed for exactly this need: built-in scorers are attached to a deployed agent rather than invoked manually, and assessments are recorded on live traces inside the MLflow experiment. This removes custom scoring code, keeps quality signals tied to each request, and enables continuous monitoring without batch jobs or dashboard workarounds.

Exam trap

The trap here is assuming that any Databricks observability surface, such as system tables or dashboards, can produce semantic quality scores, when only MLflow scorers attached to the agent compute groundedness and relevance.

26
MCQmedium

A Databricks Generative AI engineer has deployed a RAG application and is now setting up production monitoring. They want to automatically detect when the distribution of incoming user questions diverges from the distribution seen during development, so they can trigger retraining or prompt adjustments. Which Databricks capability should they configure?

A.Mosaic AI Agent Evaluation with a custom LLM judge for every request
B.Delta Live Tables expectations on the raw question stream
C.Inference tables with Lakehouse Monitoring for the endpoint's payload and response columns
D.MLflow Model Registry stage transitions with webhook notifications
AnswerC

Inference tables log the request payload and model response for a Mosaic AI Model Serving endpoint, and Lakehouse Monitoring can profile those tables to compute drift metrics on the input distribution. This directly addresses detecting when live user questions diverge from the development baseline, enabling automated alerts and retraining triggers.

Why this answer

Inference tables capture the actual request and response payloads served by the endpoint, and Lakehouse Monitoring can build profiles and drift metrics over those tables. Together they provide the automated distributional comparison needed to detect when live questions diverge from the development baseline, which is exactly the monitoring goal.

Exam trap

The trap here is assuming that any monitoring feature can detect input drift, when only inference tables combined with Lakehouse Monitoring profile the actual request payload distribution.

27
Multi-Selecthard

A Generative AI engineer is designing an evaluation harness for a customer-support RAG agent on Databricks. They need metrics that specifically assess the RETRIEVAL stage rather than the generation stage. (Choose two.)

Select 2 answers
A.Context recall, measuring whether the retrieved passages collectively contain the information present in the ground-truth answer.
B.Answer relevance, measuring how well the final generated response addresses the user's original question.
C.Toxicity, measuring whether the generated response contains harmful or offensive content.
D.Groundedness, measuring whether the answer's claims are supported by the retrieved context.
E.Context precision, measuring the proportion of retrieved chunks that are actually relevant and how highly they are ranked.
AnswersA, E

Context recall compares retrieved passages against the reference answer to determine whether all needed information was retrieved. It is computed purely from retrieval outputs and ground truth, so it isolates the retriever's ability to fetch relevant content independent of how the generator phrases its response.

Why this answer

Context recall and context precision are computed from the retrieved passages and ground-truth relevance, so they quantify what the retriever returned and in what order. Answer relevance, groundedness, and toxicity all depend on the generated response, which means they blend generation behavior into the score. Selecting the two context-based metrics keeps the harness focused on retrieval.

Exam trap

The trap here is treating groundedness as a retrieval metric because it references context, when it actually compares generated claims to context and thus spans both stages.

28
MCQhard

A healthcare analytics team has a RAG application that must not reveal protected health information from other patients. They want to continuously monitor production traffic on their Mosaic AI Model Serving endpoint and alert when responses contain unsafe content. Which Databricks capability should they configure to evaluate each logged request and response against safety criteria and route flagged records for review?

A.Unity Catalog column masks on the inference table
B.Mosaic AI Agent Evaluation with a scheduled evaluation job over inference table data
C.MLflow autologging during model training
D.Delta Lake time travel on the inference table
AnswerB

Agent Evaluation provides built-in safety and correctness judges, and it can run as a scheduled job over data captured in inference tables. That combination continuously scores production requests and responses for unsafe content and can write flagged records to a review table, which matches the requirement to monitor live traffic without manual sampling.

Why this answer

Mosaic AI Agent Evaluation supplies judges for safety and other quality dimensions, and it can be scheduled against inference table data so every production interaction is scored. Records that fail the safety criteria can be written to a separate table for review, giving the team continuous, automated monitoring rather than manual spot checks.

Exam trap

The trap here is assuming that governance features like column masks or time travel provide content-level safety evaluation, when they only control access or retain history.

29
MCQmedium

Which of the following describes the 'drift' phenomenon in the context of LLM monitoring?

A.The model's weights change during inference.
B.A shift in user query patterns over time.
C.The model generating the same answer too often.
D.Increased latency in the serving endpoint.
AnswerB

Drift in generative AI often manifests as a divergence in user query patterns, where the model encounters prompts that are significantly different from its training or fine-tuning data. This shift leads to degradation in performance, as the model was not optimized for these new interaction patterns or domains.

Why this answer

LLM drift occurs when the input data distribution or the nature of user queries changes over time, causing the model to perform worse than during validation. Because user expectations and language usage evolve, a model that performed well at launch may become less accurate as queries drift into new domains or change in tone. Monitoring this drift is essential for proactive maintenance and model updating.

Exam trap

Candidates often confuse LLM drift with traditional feature drift, missing that LLM drift specifically refers to evolving user query patterns and intents over time.

30
MCQhard

A healthcare analytics team deployed a RAG assistant on Databricks whose answers cite clinical policy documents. Compliance requires that every production answer be attributable to a specific retrieved chunk. During evaluation, the team notices the groundedness judge scores are high, but manual review finds answers that blend two policies into a statement neither document supports. Which change to their Mosaic AI Agent Evaluation configuration best detects this failure mode?

A.Increase the number of retrieved chunks per query from five to twenty so the model has more supporting evidence available.
B.Enable the chunk-level citation assessment so each claim in the response is checked against the specific retrieved chunk it references, rather than scoring the response against the concatenated context.
C.Raise the temperature setting on the served model to zero and re-run the evaluation suite to obtain more deterministic outputs.
D.Add a relevance judge that scores the retrieved chunks against the user question and filter out any chunk scoring below a fixed threshold.
AnswerB

Response-level groundedness against merged context can pass when each claim is individually supported somewhere in the pool, even if the response attributes a blended statement to the wrong source. Chunk-level citation assessment forces per-claim traceability to the cited chunk, which is exactly the compliance requirement and surfaces cross-document blending that aggregate groundedness masks.

Why this answer

The compliance requirement is per-answer attribution to a specific chunk, so the evaluation must verify each claim against the chunk it cites. Scoring the whole response against concatenated context allows a blended claim to pass whenever its components appear somewhere in the pool. Chunk-level citation assessment closes that gap by testing attribution directly, which is what distinguishes this failure from ordinary ungrounded hallucination.

Exam trap

The trap here is assuming a high groundedness score means every claim is properly attributed, when response-level groundedness against merged context can pass blended statements whose parts each appear somewhere in the retrieved pool.

31
MCQmedium

A GenAI team at a retail bank runs a RAG assistant on a Databricks Mosaic AI Model Serving endpoint. During a pilot, they captured end-user thumbs-up/down feedback in a Delta table but never joined it to the trace payloads. Six weeks later, hallucination complaints spike, yet the aggregate thumbs-down rate is unchanged. Which approach best resolves this discrepancy using Databricks-native tooling?

A.Increase the endpoint's provisioned concurrency and enable autoscaling so responses are returned faster during peak hours.
B.Switch the judge from a smaller open model to a larger frontier model and re-run the nightly evaluation job over the last 30 days.
C.Retrain the embedding model on the latest product documentation and redeploy the index to the Vector Search endpoint.
D.Join the Delta feedback table to MLflow traces on request_id and segment quality metrics by user cohort, prompt template version, and retrieved document source.
AnswerD

Joining feedback rows to MLflow traces via the shared request_id lets the team slice the unchanged aggregate rate by the dimensions that actually moved, exposing whether the spike is confined to one prompt template version or document source. Aggregates hide this because a large silent-majority cohort with stable thumbs-up dilutes the regressed segment.

Why this answer

Aggregate feedback rates are diluted by cohort mix, so an unchanged overall thumbs-down percentage can coexist with a severe regression inside one prompt template version, document source, or user cohort. Correlating the stored human feedback with MLflow traces on request_id and then slicing by those dimensions surfaces where the hallucination spike actually lives, which is the prerequisite for any targeted fix.

Exam trap

The trap here is treating an unchanged aggregate feedback rate as proof the application is stable, when the real issue is that feedback was never correlated with trace attributes so segment-level regressions stay invisible.

32
Multi-Selectmedium

Which THREE of the following are common challenges when monitoring LLM applications in production that differ significantly from traditional ML monitoring? (Choose three)

Select 3 answers
A.Non-deterministic output behavior.
B.Difficulty in defining a single 'ground truth' for responses.
C.Lack of high-throughput API endpoints.
D.High cost of manual evaluation at scale.
E.The model's inability to connect to internet data.
AnswersA, B, D

LLMs can produce different outputs for the same prompt due to temperature settings or inherent stochasticity. This makes simple equality checks useless for evaluation. Monitoring must account for this variance, requiring probabilistic or semantic similarity metrics rather than static, exact-match validation common in traditional classification tasks.

Why this answer

LLMs introduce unique challenges because their outputs are non-deterministic, high-dimensional, and often lack a single 'correct' answer. Traditional monitoring focuses on numerical drift and binary classification metrics like precision/recall. LLM evaluation requires semantic understanding, nuance assessment, and the ability to handle unstructured text, necessitating specialized tools like LLM-as-a-judge or human-in-the-loop validation to manage the subjectivity inherent in generative AI outputs.

Exam trap

Candidates often apply traditional ML monitoring assumptions, failing to account for LLM non-determinism, subjective evaluation, and the absence of a single ground truth.

33
Multi-Selectmedium

A GenAI engineer is instrumenting a production RAG application on Databricks to detect quality degradation before users complain. The team wants signals that reveal problems in the retrieval stage specifically, rather than in the generation stage. Which TWO signals should they track? (Choose two.)

Select 2 answers
A.Recall of the retrieval step measured against a curated question-to-document mapping
B.Toxicity score of the final generated answer
C.Average completion token count per response
D.Time to first token on the serving endpoint
E.Precision at k of the retrieved chunks against a labeled relevance set
AnswersA, E

Recall against a curated mapping reveals whether the chunks needed to answer each question were retrieved at all. When recall falls, the generator is starved of the correct evidence and cannot produce grounded answers. This is a retrieval-stage metric by construction and localizes failures to indexing, chunking, or embedding rather than to the LLM.

Why this answer

Retrieval-stage quality is measured by whether the right documents are found and how many of the returned ones are relevant. Precision at k and recall against a curated question-to-document mapping capture those two dimensions directly, so regressions localize to indexing, chunking, or embeddings. Token count, time to first token, and toxicity all describe the generator or the serving layer, not retrieval.

Exam trap

The trap here is treating any production signal, such as latency or output toxicity, as evidence about retrieval quality, when only metrics computed against the retrieved chunk set can isolate the retrieval stage.

34
MCQeasy

Why should you integrate Mosaic AI Model Evaluation with Unity Catalog?

A.To increase the model's training speed.
B.To provide centralized governance and lineage.
C.To automatically delete old evaluation data.
D.To bypass the need for prompt engineering.
AnswerB

Unity Catalog provides a single source of truth for governance, lineage, and access controls. Integrating evaluation with it ensures that all model performance data can be linked to specific training and data versions, providing the full transparency and accountability required for enterprise-grade AI risk management.

Why this answer

Unity Catalog acts as the central governance layer for all Databricks data. By integrating evaluation with Unity Catalog, you ensure that all evaluation datasets, results, and models are discoverable, lineage-tracked, and governed. This is essential for compliance, ensuring that every model deployment is backed by a verifiable audit trail of its evaluation performance, which is a core requirement for enterprise AI deployments that must adhere to strict internal and external standards.

Exam trap

Candidates often focus on 'performance optimization' or 'cost reduction', missing the primary purpose of Unity Catalog integration, which is centralized governance, auditability, and lineage tracking.

35
MCQhard

A team is using MLflow to track their model evaluation runs. They want to ensure that every evaluation run is reproducible. What is the best practice to achieve this in Databricks?

A.Only log the final scalar metrics.
B.Log the model URI, dataset version, and code hash.
C.Hardcode the dataset paths in the notebook.
D.Use a global variable for model versioning.
AnswerB

Logging these identifiers provides a complete trail of the experiment setup. This allows engineers to re-run the exact configuration to verify results or investigate anomalies. It is the standard best practice for maintaining a rigorous, reproducible research and evaluation pipeline in professional Databricks ML environments.

Why this answer

Reproducibility in evaluation requires logging not only the results but also the exact artifacts, environment, and code version used. MLflow's ability to log the model URI, the specific evaluation dataset version, and the code state (via git hash) ensures that an evaluation run can be perfectly recreated. This is mandatory for auditing and verifying improvements in LLM performance over time, preventing 'black box' results that cannot be validated.

Exam trap

Test-takers often think saving just the model weights is sufficient, forgetting that data versioning and code hashes are required for complete reproducibility.

36
MCQeasy

A GenAI engineer notices that a RAG agent's retrieval stage is returning relevant chunks, but the generated answers frequently omit key facts present in those chunks. The team wants a single evaluation metric that isolates whether the generator is using the provided context. Which metric should they focus on?

A.Answer correctness against a reference answer.
B.Groundedness, which checks whether claims in the answer are supported by the retrieved context.
C.p95 latency of the serving endpoint.
D.Context recall of the retriever.
AnswerB

Groundedness evaluates each claim in the generated answer against the retrieved context. When retrieval is good but the answer omits supported facts or invents unsupported ones, groundedness drops. It isolates the generator's use of context, which is exactly the failure described.

Why this answer

Groundedness is the metric that inspects whether the answer's claims are supported by the retrieved context. Since the retriever is already delivering relevant chunks, a low groundedness score points squarely at the generator ignoring or misusing that evidence, giving the team a precise diagnostic for the stage that is failing.

Exam trap

The trap here is choosing answer correctness because the final output is wrong, when the diagnostic question is which pipeline stage is failing and only groundedness isolates the generator's use of context.

37
MCQmedium

When monitoring a production LLM, you detect a drift where the model's responses are becoming increasingly verbose and less helpful compared to the baseline. Which strategy is most effective for detecting this quality decay?

A.Monitor system-level CPU and memory usage of the inference cluster.
B.Set up alerts for high request volume spikes in the API logs.
C.Implement automated LLM-as-a-judge evaluations on a sample of production inputs.
D.Require manual reviews for every single response generated by the model.
AnswerC

Using an LLM as a judge allows for automated, scalable evaluation of qualitative metrics. By comparing production outputs against predefined criteria or reference answers, you can effectively track quality drift over time and identify when model responses deteriorate.

Why this answer

Model quality decay in LLMs is best detected through continuous automated evaluation using a 'judge' model. By running a set of evaluation prompts through the production model and grading them against an LLM-as-a-judge, you can quantify performance trends. This proactive monitoring allows teams to identify when model behavior deviates from expected standards, triggering alerts before the issues impact a large number of end-users or require significant architectural changes.

38
Multi-Selectmedium

Which TWO evaluation approaches are most effective for measuring the quality of a RAG pipeline's retrieval stage?

Select 2 answers
A.Context Precision
B.Model Answer Faithfulness
C.Context Recall
D.Answer Semantic Similarity
E.LLM Latency Monitoring
AnswersA, C

Context Precision measures the proportion of retrieved chunks that are genuinely relevant to the query, isolating the retrieval stage's ranking quality independently of generation. This satisfies the stem's requirement to evaluate retrieval specifically, revealing whether irrelevant context is being surfaced before the LLM processes it.

Why this answer

Evaluating the retrieval stage requires assessing how well the system identifies relevant documents from the vector store before the LLM generates a response. Context Precision and Context Recall provide quantitative measures to ensure the retrieval engine is performing correctly. Without these metrics, it is impossible to determine if the LLM is failing due to poor generation or simply because it lacked the correct information to answer the query.

39
MCQhard

A retail company runs a customer-facing RAG assistant on a Mosaic AI Model Serving endpoint. The team wants every production request, response, and retrieved context to be captured automatically into a Unity Catalog Delta table so they can monitor quality and latency trends over time. Which action should they take?

A.Configure MLflow autologging on the notebook that calls the serving endpoint.
B.Register the model in Unity Catalog and enable lineage tracking on the catalog.
C.Enable inference tables on the Model Serving endpoint.
D.Enable Model Serving request logging to a cloud storage bucket.
AnswerC

Inference tables are the Databricks feature that automatically logs the request payload, response, and metadata from a Model Serving endpoint into a Unity Catalog Delta table. Enabling them on the endpoint requires no application code changes and captures retrieved context when the agent passes it through the payload, giving the team the monitoring substrate they want for trend analysis.

Why this answer

Inference tables are the native Databricks mechanism for capturing production traffic from a Model Serving endpoint into a governed Unity Catalog Delta table. They log payloads, responses, and metadata automatically, which is exactly what the team needs to analyze quality and latency over time. Other options either log only development activity or record lineage instead of runtime traffic, so they cannot satisfy the monitoring requirement.

Exam trap

The trap here is assuming that MLflow autologging or Unity Catalog lineage will capture live endpoint traffic, when only inference tables persist request and response payloads into a Delta table.

40
MCQmedium

Your organization requires an auditable record of all LLM evaluation results for compliance. Which Databricks feature provides the best centralized storage for these evaluation runs?

A.Databricks File System (DBFS) text logs
B.MLflow Experiments
C.Unity Catalog volume for raw CSV storage
D.Git commit history of the inference code
AnswerB

MLflow Experiments provide a structured environment to log parameters, metrics, and artifacts. This creates a highly auditable, searchable, and version-controlled record of every evaluation, which is ideal for compliance and tracking the history of model performance.

Why this answer

MLflow Experiments act as the centralized repository for all model development and evaluation runs. By logging evaluation results to MLflow, teams create a permanent, version-controlled audit trail. This allows stakeholders to compare different iterations of models, review evaluation metrics, and verify that quality gates were met before deployment, ensuring compliance with internal AI governance policies and external regulatory frameworks for model transparency.

41
MCQeasy

What is the primary benefit of using MLflow Model Evaluation for generative AI applications compared to manual evaluation methods?

A.It automatically rewrites the prompt templates based on user feedback.
B.It enables consistent, repeatable evaluation metrics across different model versions.
C.It removes the need for ground truth datasets during the testing phase.
D.It creates the vector index automatically from raw unstructured data.
AnswerB

MLflow offers a consistent framework to calculate standard metrics like faithfulness and relevance. This ensures that every model update is evaluated using the same methodology, providing a reliable baseline for comparing performance improvements or regressions across versions.

Why this answer

MLflow Model Evaluation provides a standardized framework to calculate metrics like RAGAS and perform side-by-side comparisons of model versions. By automating this process, organizations can ensure consistency across releases, track performance history over time, and reduce human bias. This systematic approach allows teams to make data-driven decisions when deciding whether to promote a model to production, ensuring reliability and auditability in the AI lifecycle.

Exam trap

Candidates often assume MLflow Model Evaluation automatically tunes hyper-parameters, overlooking its core role in providing consistent, repeatable evaluation metrics across model versions.

42
Multi-Selectmedium

When evaluating LLM outputs, which TWO metrics are most appropriate for measuring the 'quality' of a response in a RAG system? (Choose two)

Select 2 answers
A.Faithfulness
B.Inference Latency
C.Answer Relevance
D.GPU Utilization
E.Training Iterations
AnswersA, C

Faithfulness checks if the generated answer is derived from the retrieved context. This is the primary metric for preventing hallucinations, which are the biggest risk in RAG deployments. A faithful model stays within the bounds of its provided information, ensuring reliability and accuracy for the end user.

Why this answer

Answering quality in RAG requires looking at two distinct dimensions: the correctness of the content (faithfulness) and the relevance of the answer to the user's intent. Faithfulness ensures the model remains grounded in the provided evidence, while answer relevance captures whether the model actually addressed the specific user question. These metrics provide a balanced view, ensuring the system is both accurate and useful, which is essential for end-user satisfaction.

43
MCQeasy

When evaluating a generative AI model using the Mosaic AI Model Evaluation tool, what is the primary purpose of providing a 'baseline' dataset?

A.To increase the number of tokens the model can generate.
B.To provide a reference for comparative performance analysis.
C.To act as a secondary training set for fine-tuning.
D.To optimize the vector database search latency.
AnswerB

A baseline allows developers to quantitatively compare current model performance against known good results. This comparison is vital for validating that new model versions or updated RAG configurations maintain or exceed performance standards, facilitating data-driven decisions about whether to promote a new version to production.

Why this answer

A baseline dataset serves as a ground-truth reference point. By comparing the model's current outputs against this known high-quality set, engineers can objectively measure performance degradation or improvement. This benchmarking approach is fundamental to scientific model evaluation, ensuring that updates to the model or data retrieval logic do not silently introduce regressions in the quality of the generated answers during the deployment process.

Exam trap

Candidates incorrectly assume a baseline dataset is used for fine-tuning the model, rather than understanding its purpose as a reference point for comparative evaluation.

44
MCQhard

An engineer registers a GenAI agent to Unity Catalog and enables inference tables on its Mosaic AI Model Serving endpoint. They want to automatically detect when response quality degrades in production without waiting for human review. Which capability should they configure to achieve continuous automated quality monitoring?

A.Configure autoscaling on the serving endpoint so additional model replicas absorb traffic spikes and stabilize response quality.
B.Set up a scheduled job that calls the endpoint with synthetic prompts and records latency percentiles to a Delta table for trending.
C.A Databricks Lakehouse Monitoring monitor over the inference table, joined with a labeled evaluation dataset to compute quality metrics on a schedule.
D.Enable model signature enforcement on the endpoint so requests that violate the expected schema are rejected before scoring.
AnswerC

Inference tables capture request and response payloads in a Delta table, and Lakehouse Monitoring can build a monitor over that table to compute profile and drift metrics on a schedule. Joining with labeled ground-truth data lets the monitor compute quality metrics continuously, providing automated degradation detection without manual review of each response.

Why this answer

Inference tables persist production request and response payloads as Delta, which makes them queryable by Lakehouse Monitoring. Building a monitor over that table, especially when joined with labeled data, yields scheduled quality and drift metrics that surface degradation automatically. This combination is the supported path for continuous, automated quality monitoring on Databricks rather than ad hoc manual inspection.

Exam trap

The trap here is confusing operational telemetry such as latency and autoscaling with semantic quality monitoring, which requires payload capture plus a monitor.

45
MCQeasy

A team has deployed a RAG application on Databricks and wants to monitor the quality of responses in production. They have enabled inference table logging. Which built-in Databricks capability allows them to periodically evaluate the logged requests and responses for quality metrics like groundedness?

A.Databricks SQL dashboards that display the count of requests per hour.
B.Delta Live Tables expectations that validate response length.
C.Mosaic AI Agent Evaluation, which can be run on the inference table to compute quality metrics.
D.MLflow Tracking with custom metrics logged during model serving.
AnswerC

Mosaic AI Agent Evaluation is designed to evaluate agent and RAG applications. It can be run on logged inference tables to compute metrics such as groundedness and relevance. This allows continuous monitoring of production quality without manual intervention. It integrates with MLflow and Databricks workflows, making it the appropriate built-in capability for this scenario.

Why this answer

Mosaic AI Agent Evaluation is the built-in Databricks capability for evaluating RAG and agent applications. It can process inference tables and compute quality metrics such as groundedness, relevance, and answer correctness. This enables periodic, automated monitoring of production quality without custom code.

It is the correct tool for the described scenario.

Exam trap

The trap here is assuming that any monitoring tool, like SQL dashboards or Delta Live Tables expectations, can compute semantic quality metrics, when only Agent Evaluation provides LLM-specific judges.

46
Multi-Selecthard

Which THREE components are critical to include in a comprehensive evaluation strategy for a RAG-based Generative AI application?

Select 3 answers
A.Retrieval metrics (e.g., Hit Rate, MRR)
B.Generation metrics (e.g., Faithfulness, Answer Relevance)
C.End-to-end performance metrics (e.g., Answer Correctness)
D.Hardware temperature monitoring for GPUs
E.Network packet loss percentage
AnswersA, B, C

Retrieval metrics are essential because they confirm whether the system is successfully finding the correct documents. If retrieval fails, no amount of prompt engineering or fine-tuning will lead the generator to provide a correct, factual answer.

Why this answer

A robust RAG evaluation requires a holistic view that covers the entire pipeline. Evaluating just the generation or just the retrieval is insufficient because the quality of the final response is dependent on both stages. By measuring retrieval performance, generation faithfulness, and end-to-end answer quality, engineers obtain a full picture of where the application is succeeding and where it needs tuning to meet user requirements.

47
MCQeasy

A Generative AI engineer is deploying a new version of a RAG chain to a Mosaic AI Model Serving endpoint. Before promoting it to production, they want to run an evaluation that checks whether the generated answers are faithful to the retrieved documents. Which Mosaic AI Agent Evaluation metric should they examine?

A.chunk_relevance
B.latency
C.request_count
D.groundedness
AnswerD

groundedness measures whether the generated answer is supported by the retrieved context, directly assessing faithfulness to the documents. This is exactly the quality dimension the engineer needs to verify before promotion. It helps catch hallucinations where the model adds information not present in the retrieved chunks.

Why this answer

To verify that generated answers are faithful to the retrieved documents, the evaluation must compare the answer against the context. groundedness is the Agent Evaluation metric that performs this check, flagging responses that include unsupported claims. It is the appropriate pre-promotion quality gate for a RAG chain.

Exam trap

The trap here is confusing retrieval-stage metrics like chunk_relevance with answer-stage metrics, when faithfulness of the final answer is measured by groundedness.

48
MCQhard

An engineer uses MLflow LLM Evaluation with mlflow.evaluate() to score a RAG application. The judge model is configured with a temperature of 0.9 and no explicit metric thresholds are set. Reruns of the identical evaluation dataset produce relevance scores that swing by up to 20 percentage points, and the team cannot tell whether a prompt change helped. Which change most directly improves the reliability of the evaluation comparison?

A.Increase the evaluation dataset size by duplicating each question three times so variance averages out across more judge calls.
B.Set the judge model's temperature to 0 and define fixed pass/fail thresholds for each metric before comparing runs.
C.Switch the evaluation to a larger judge model with more parameters while keeping the existing sampling settings.
D.Log only the aggregate mean relevance score per run and discard the per-row judgments from the MLflow run artifacts.
AnswerB

The instability originates from stochastic sampling in the judge, and temperature 0 makes the judge's verdicts deterministic for a given input, removing run-to-run score drift. Pairing that with predefined metric thresholds converts continuous, noisy scores into stable pass/fail decisions, so a prompt change can be attributed to real quality movement instead of sampling noise. This combination directly addresses reproducibility and comparability of evaluation runs.

Why this answer

Score swings across identical runs are caused by nondeterministic judge sampling, so the fix must remove that randomness and make the decision rule stable. Setting the judge temperature to zero makes each judgment reproducible, and predefined thresholds turn noisy continuous scores into consistent pass/fail outcomes. Enlarging the judge, duplicating rows, or hiding per-row results leaves the stochasticity in place and keeps run-to-run comparisons unreliable.

Exam trap

The trap here is treating evaluation variance as a dataset-size problem when the root cause is a high-temperature judge producing different verdicts on identical inputs.

49
MCQmedium

Refer to the exhibit. An engineer observes an unexpected drop in "relevance" for the latest deployment. What is the most likely cause related to the evaluation process itself?

A.The model was trained on too much data.
B.The evaluation dataset is misaligned with current production query trends.
C.The inference cluster has too much compute capacity.
D.The faithfulness score increased too much.
AnswerB

If production traffic evolves but the evaluation set remains static, the metrics will reflect outdated user needs. A drop in relevance usually signifies that the model is no longer performing well on the types of questions users are currently asking.

Why this answer

A drop in relevance often suggests that the evaluation dataset (eval_set_v4) may have become outdated or misaligned with the current production use cases. If the user base or query patterns shift, the original evaluation set may no longer accurately reflect the actual performance requirements. Regularly refreshing the evaluation dataset to match current production trends is essential to ensure that metrics remain valid indicators of system performance.

Exam trap

Candidates often blame model degradation first, overlooking that an outdated or misaligned evaluation dataset itself can cause an unexpected drop in metric scores.

50
MCQhard

A team is using MLflow LLM Evaluation with the built-in answer_correctness metric to compare two prompt templates for a question-answering application. They notice that answer_correctness scores are nearly identical, but manual review shows one template produces answers that are factually correct yet omit key supporting details. Which additional built-in metric should they add to their evaluation to surface this difference?

A.groundedness
B.answer_completeness
C.toxicity
D.answer_similarity
AnswerB

answer_completeness evaluates whether the response covers all the key information present in the ground truth answer. Since the issue is that one template omits supporting details while remaining factually correct, this metric directly measures that gap and will differentiate the templates. It is the built-in metric designed to assess thoroughness against the reference.

Why this answer

When answers are factually correct but differ in how much supporting detail they include, a completeness metric is required. answer_completeness compares the response against the ground truth to check whether all key points are covered, making it the right built-in metric to expose the omission of supporting details that answer_correctness alone misses.

Exam trap

The trap here is treating factual correctness and completeness as the same thing, when a response can be correct yet incomplete and needs a dedicated completeness metric to detect it.

51
Multi-Selecthard

A Generative AI engineer is evaluating a RAG pipeline using MLflow LLM Evaluation on Databricks. They want to assess both the retrieval quality and the generation quality. Which TWO built-in evaluation metrics should they use? (Choose two.)

Select 2 answers
A.`perplexity`
B.`toxicity`
C.`ROUGE`
D.`relevance`
E.`groundedness`
AnswersD, E

Relevance assesses how well the retrieved documents match the user's query, directly evaluating retrieval quality. In MLflow LLM Evaluation, the relevance metric can be computed for the retrieval step. It helps identify whether the retriever is fetching appropriate context. This is essential for diagnosing retrieval failures in a RAG pipeline.

Why this answer

For a RAG pipeline, retrieval quality is assessed by relevance, which measures how well retrieved documents match the query. Generation quality is assessed by groundedness, which checks if the answer is supported by the retrieved context. Both are built-in metrics in MLflow LLM Evaluation and are specifically designed for RAG evaluation on Databricks.

Exam trap

The trap here is selecting common NLP metrics like ROUGE or perplexity, which are not part of MLflow's built-in RAG evaluation suite and do not measure retrieval or groundedness.

52
MCQmedium

A generative AI team runs a RAG chatbot whose Mosaic AI Model Serving endpoint is monitored in Unity Catalog inference tables. Over two weeks, the percentage of user questions that receive a refusal answer ('I don't have enough information') climbs from 4% to 31%, while retrieval latency and token counts stay flat. The team wants the earliest actionable signal that the retrieval corpus has gone stale rather than the prompt or model. Which monitoring signal should they inspect first?

A.The average generation time per completion recorded by the serving endpoint's request logs.
B.The endpoint's provisioned concurrency utilization and queue depth metrics from the serving endpoint logs.
C.The distribution of retrieval scores for the top-k chunks returned by the Vector Search index, tracked over time in the inference table.
D.The count of prompt tokens consumed per request, compared week over week in the billing usage table.
AnswerC

A rising refusal rate with unchanged latency and token counts points at the retrieval stage, not generation. If top-k similarity scores drift downward or their spread collapses, the index is returning progressively weaker matches for the same query mix, which is the earliest measurable evidence that the corpus has gone stale. Inspecting score distributions in the logged inference table isolates retrieval quality before blaming prompt or model changes.

Why this answer

Refusals rising while latency and token counts hold steady isolates the problem to retrieval quality rather than the model or the serving tier. Retrieval score distributions logged to the inference table show whether top-k chunks are becoming weaker matches, which is the earliest signal that the indexed corpus no longer covers current user questions. Checking capacity, token cost, or generation time cannot distinguish stale content from a healthy pipeline.

Exam trap

The trap here is assuming a rising refusal rate must be a prompt-engineering or model problem, when stable latency and token counts actually point to degraded retrieval relevance.

53
MCQhard

A team runs an offline evaluation of a RAG chatbot using MLflow LLM Evaluation with an LLM judge. The same evaluation dataset produces noticeably different scores when re-run on different days, even though the application code and retrieved documents are unchanged. Which action best addresses this score instability?

A.Increase the number of retrieved chunks per query so the judge has more context
B.Switch all metrics from LLM-judged to exact-match string comparison
C.Reduce the evaluation dataset to a single representative example
D.Pin the judge model version and set temperature to zero for evaluation runs
AnswerD

Varying scores on identical inputs are a classic symptom of a nondeterministic judge: the judge model version may silently change, and sampling temperature introduces randomness in its verdicts. Pinning the judge model version and setting temperature to zero makes the scoring function repeatable, so differences in scores reflect real changes in the application rather than judge noise. This directly targets the instability described.

Why this answer

Identical inputs producing different scores across runs indicate the judge itself is nondeterministic, typically through model version drift or sampling temperature. Pinning the judge model version and forcing temperature to zero makes scoring reproducible, so observed differences can be attributed to the application. The other options either change the system under test, destroy evaluation fidelity, or hide variance rather than eliminate it.

Exam trap

The trap here is blaming the application or the retrieval pipeline for score drift, when unchanged inputs and code point squarely at the LLM judge's own nondeterminism.

54
MCQeasy

A support team operates a Databricks-hosted RAG assistant and wants end-user feedback to feed their monitoring dashboards. They plan to add a thumbs-up and thumbs-down control to the chat UI and log each vote alongside the request ID. What is the primary value of collecting this feedback for the evaluation and monitoring workflow?

A.It guarantees that groundedness scores will improve over time
B.It provides real-world signal that can be joined to traces for triage and dataset curation
C.It replaces the need for any offline evaluation dataset
D.It automatically fine-tunes the underlying foundation model
AnswerB

Thumbs votes are weak labels, but joined to request IDs and traces they point directly at production interactions that went wrong or right. Teams use negatively rated requests to build triage queues and to curate new evaluation examples, closing the loop from production back into offline testing. This is the core value: grounding monitoring and dataset growth in actual user experience.

Why this answer

End-user thumbs votes are weak but real labels tied to specific production requests. Joined with traces and request IDs, they drive triage of failing interactions and curation of new offline evaluation examples, connecting production monitoring to the evaluation dataset. They do not replace datasets, train models automatically, or improve metrics on their own.

Exam trap

The trap here is overvaluing raw user feedback as a complete evaluation or training signal, when it is sparse, biased, and useful mainly for triage and dataset curation.

Ready to test yourself?

Try a timed practice session using only Evaluation and Monitoring questions.