Courseiva

Databricks Certified Generative AI Engineer Associate (Databricks-GenAI-Assoc) — Questions 151–225

330 questions total · 5pages · All types, answers revealed

Page 2

Page 3 of 5

Page 4
151
Multi-Selecthard

Which THREE conditions must be met for a user to successfully create a new table in a Unity Catalog schema?

Select 3 answers
A.The user must have the USAGE privilege on the catalog.
B.The user must have the USAGE privilege on the schema.
C.The user must have the CREATE TABLE privilege on the schema.
D.The user must have the MODIFY privilege on the catalog.
E.The user must have the OWNER role for the entire metastore.
AnswersA, B, C

Access to any object within a catalog starts with the USAGE privilege at the catalog level. Without this fundamental permission, a user cannot navigate into the catalog to access schemas or perform any operations like creating new tables, even if they have other schema-level permissions.

Why this answer

Creating a table in Unity Catalog involves a specific permission chain. The user must have USAGE on the catalog, USAGE on the schema, and the CREATE TABLE privilege on the schema. These requirements ensure that governance is maintained at every level of the hierarchy, preventing unauthorized data creation in sensitive or restricted namespaces within the organization's data architecture.

Exam trap

Many candidates forget the 'USAGE' privilege on the parent catalog and schema, incorrectly assuming that 'CREATE TABLE' permission on the schema is sufficient to perform the action.

152
Multi-Selectmedium

An engineer is packaging a custom PyFunc model that wraps an open-source LLM and must deploy it to a Databricks Model Serving endpoint. The endpoint must load the model from Unity Catalog and expose it through a REST API. Which two actions are required to make the deployment succeed? (Choose two.)

Select 2 answers
A.Register the logged model in Unity Catalog and grant the serving identity USE CATALOG, USE SCHEMA, and EXECUTE or SELECT privileges on the model version.
B.Set the endpoint's scale-to-zero configuration to disabled so the model stays in memory permanently.
C.Attach the model to a Databricks cluster that stays running so the endpoint can proxy requests to it.
D.Log the model with mlflow.pyfunc.log_model, including a pip_requirements entry for every runtime dependency the wrapper imports.
E.Convert the PyFunc wrapper to a scikit-learn estimator so the serving container recognizes the flavor.
AnswersA, D

Model Serving resolves the model by its Unity Catalog three-level name, and the endpoint's identity must be able to read that model version. Without the catalog, schema, and model privileges, endpoint creation fails with a permission error even though the artifact itself is valid. Granting these is mandatory for Unity Catalog-based serving.

Why this answer

A successful custom model deployment requires a correctly logged artifact with all dependencies declared, and a Unity Catalog registration with privileges granted to the serving identity. Environment build failures and permission errors are the two dominant causes of endpoint creation failure. Flavor conversion, attached clusters, and scale-to-zero settings are unrelated to whether the deployment itself can complete.

Exam trap

The trap here is focusing on endpoint scaling or compute choices, when deployment success actually hinges on declared dependencies and Unity Catalog permissions.

153
MCQmedium

A Generative AI engineer is packaging a retrieval-augmented generation application for deployment with Databricks Asset Bundles. The bundle must provision a Databricks job that periodically refreshes a Delta table used as the vector index, and the job requires a specific cluster node type that differs between the development and production workspaces. Which approach correctly handles the node type difference while keeping a single bundle definition?

A.Create a separate bundle directory for each environment with entirely duplicated resource files and deploy each one independently.
B.Define the node type in a databricks.yml variable and override its value in a target-specific block for each environment.
C.Use a notebook widget inside the job to read the node type from a table at runtime and restart the cluster with the new value.
D.Hard-code the production node type in the job resource and rely on the development workspace to silently substitute a smaller node type.
AnswerB

Databricks Asset Bundles support variables declared at the top level and overridden per target, so the same job definition can use a development node type and a production node type without duplicating the resource YAML. This is the intended mechanism for environment-specific values.

Why this answer

The bundle should declare a variable for the node type and override it inside each target, because Databricks Asset Bundles resolve variables per target at deploy time. This keeps one authoritative job definition while letting development and production use different compute, and it avoids duplicated YAML that drifts over time.

Exam trap

The trap here is assuming a single bundle cannot express environment differences, which pushes candidates toward duplicating resource files instead of using target-scoped variable overrides.

154
MCQmedium

A Databricks Generative AI engineer has deployed a RAG application and is now setting up production monitoring. They want to automatically detect when the distribution of incoming user questions diverges from the distribution seen during development, so they can trigger retraining or prompt adjustments. Which Databricks capability should they configure?

A.Mosaic AI Agent Evaluation with a custom LLM judge for every request
B.Delta Live Tables expectations on the raw question stream
C.Inference tables with Lakehouse Monitoring for the endpoint's payload and response columns
D.MLflow Model Registry stage transitions with webhook notifications
AnswerC

Inference tables log the request payload and model response for a Mosaic AI Model Serving endpoint, and Lakehouse Monitoring can profile those tables to compute drift metrics on the input distribution. This directly addresses detecting when live user questions diverge from the development baseline, enabling automated alerts and retraining triggers.

Why this answer

Inference tables capture the actual request and response payloads served by the endpoint, and Lakehouse Monitoring can build profiles and drift metrics over those tables. Together they provide the automated distributional comparison needed to detect when live questions diverge from the development baseline, which is exactly the monitoring goal.

Exam trap

The trap here is assuming that any monitoring feature can detect input drift, when only inference tables combined with Lakehouse Monitoring profile the actual request payload distribution.

155
MCQhard

An AI engineer is deploying a RAG application using Databricks Model Serving. They need to ensure the endpoint can handle high traffic with low latency and automatically scale based on demand. Which configuration should they use?

A.Configure autoscaling with a minimum of 1 and maximum of 10 replicas, and set appropriate concurrency per replica.
B.Enable scale-to-zero and set a maximum concurrency of 1.
C.Disable autoscaling and manually provision 20 replicas to handle peak load.
D.Use a single large replica with maximum concurrency set to 100.
AnswerA

Autoscaling dynamically adjusts the number of replicas based on load, ensuring low latency during traffic spikes and cost savings during idle periods. Setting a minimum of 1 avoids cold starts, while a maximum of 10 caps resource usage. Concurrency per replica should be tuned to the model's throughput. This meets the requirements for high traffic and automatic scaling.

Why this answer

Autoscaling with a minimum and maximum replica count allows the endpoint to scale out during high traffic and scale in during low traffic, optimizing both latency and cost. Setting appropriate concurrency per replica ensures efficient utilization. This configuration provides the elasticity required for high traffic with low latency, while avoiding the pitfalls of fixed provisioning or overly restrictive concurrency limits.

Exam trap

The trap here is focusing only on scale-to-zero for cost savings without considering the need for sufficient concurrency and autoscaling to handle high traffic.

156
MCQmedium

Which Unity Catalog object should be used to manage access to a folder of JSON files that are not part of a formal Delta table?

A.Managed Table
B.External Location
C.Unity Catalog Volume
D.Global View
AnswerC

Volumes are the designated Unity Catalog objects for managing non-tabular data. They allow users to read, write, and manage file-based assets using a familiar path-based interface, with full support for Unity Catalog's granular access controls at the volume level.

Why this answer

Volumes in Unity Catalog are designed specifically for managing non-tabular data, such as JSON files, CSVs, or machine learning models. By creating a volume, you can govern access to these files using the same security model as tables. This allows for consistent data governance across all types of assets in your environment, not just those formatted as Delta tables.

Exam trap

Candidates often incorrectly select 'External Location' or 'Managed Table', failing to realize that Volumes are the specific Unity Catalog object built for governance of non-tabular file assets like JSON or CSV.

157
MCQmedium

A team is deploying a GenAI agent built with Mosaic AI Agent Framework. The agent must be exposed as a REST API that automatically scales with traffic and records inference logs to Unity Catalog. Which deployment mechanism meets these requirements with the least operational overhead?

A.Deploy the agent to a Databricks Model Serving endpoint using the agents.deploy() API.
B.Deploy the agent on a single-node cluster with a public IP and expose it via a Flask server.
C.Package the agent as a Python wheel and run it on a Databricks job cluster with a scheduled trigger.
D.Register the agent as a Unity Catalog model and call it directly from a notebook using mlflow.pyfunc.load_model().
AnswerA

The agents.deploy() API in the Mosaic AI Agent Framework is purpose-built to deploy agents to a Model Serving endpoint. It automatically provisions a scalable REST endpoint, enables inference table logging to Unity Catalog, and integrates with the agent's evaluation and monitoring stack, requiring minimal manual configuration.

Why this answer

The agents.deploy() API is the native Mosaic AI Agent Framework mechanism for taking an agent from development to a production Model Serving endpoint. It handles endpoint creation, scaling, and integrates with Unity Catalog inference tables for logging, so it directly satisfies all stated requirements with minimal manual work.

Exam trap

The trap here is assuming any compute that can run Python can serve an agent, when only a Model Serving endpoint provides autoscaling REST access with built-in inference logging.

158
Multi-Selecthard

A Generative AI engineer is designing an evaluation harness for a customer-support RAG agent on Databricks. They need metrics that specifically assess the RETRIEVAL stage rather than the generation stage. (Choose two.)

Select 2 answers
A.Context recall, measuring whether the retrieved passages collectively contain the information present in the ground-truth answer.
B.Answer relevance, measuring how well the final generated response addresses the user's original question.
C.Toxicity, measuring whether the generated response contains harmful or offensive content.
D.Groundedness, measuring whether the answer's claims are supported by the retrieved context.
E.Context precision, measuring the proportion of retrieved chunks that are actually relevant and how highly they are ranked.
AnswersA, E

Context recall compares retrieved passages against the reference answer to determine whether all needed information was retrieved. It is computed purely from retrieval outputs and ground truth, so it isolates the retriever's ability to fetch relevant content independent of how the generator phrases its response.

Why this answer

Context recall and context precision are computed from the retrieved passages and ground-truth relevance, so they quantify what the retriever returned and in what order. Answer relevance, groundedness, and toxicity all depend on the generated response, which means they blend generation behavior into the score. Selecting the two context-based metrics keeps the harness focused on retrieval.

Exam trap

The trap here is treating groundedness as a retrieval metric because it references context, when it actually compares generated claims to context and thus spans both stages.

159
MCQeasy

Which tool in the Databricks ecosystem is best suited for developers to experiment with prompt engineering and tool-calling logic iteratively?

A.Databricks SQL Editor.
B.Databricks Notebooks.
C.Unity Catalog explorer.
D.Cluster Management UI.
AnswerB

Notebooks offer the most flexible and interactive environment for iterating on LLM prompts and testing code-based logic. The ability to run segments of code, visualize results, and maintain a history of experiments makes it the preferred tool for developers building and refining their generative AI application logic.

Why this answer

Databricks notebooks provide an interactive environment that is perfectly suited for iterative experimentation with prompts and chains. Developers can easily test different variations, observe the model's output in real-time, and integrate various tools, making it the ideal workbench for rapid prototyping in the early stages of the AI development lifecycle before committing to a final, production-ready implementation of their generative AI application.

Exam trap

Candidates often select heavy CI/CD deployment pipelines or MLflow registry tools for initial prompt experimentation, ignoring the need for rapid interactive testing.

160
MCQhard

A healthcare analytics team has a RAG application that must not reveal protected health information from other patients. They want to continuously monitor production traffic on their Mosaic AI Model Serving endpoint and alert when responses contain unsafe content. Which Databricks capability should they configure to evaluate each logged request and response against safety criteria and route flagged records for review?

A.Unity Catalog column masks on the inference table
B.Mosaic AI Agent Evaluation with a scheduled evaluation job over inference table data
C.MLflow autologging during model training
D.Delta Lake time travel on the inference table
AnswerB

Agent Evaluation provides built-in safety and correctness judges, and it can run as a scheduled job over data captured in inference tables. That combination continuously scores production requests and responses for unsafe content and can write flagged records to a review table, which matches the requirement to monitor live traffic without manual sampling.

Why this answer

Mosaic AI Agent Evaluation supplies judges for safety and other quality dimensions, and it can be scheduled against inference table data so every production interaction is scored. Records that fail the safety criteria can be written to a separate table for review, giving the team continuous, automated monitoring rather than manual spot checks.

Exam trap

The trap here is assuming that governance features like column masks or time travel provide content-level safety evaluation, when they only control access or retain history.

161
MCQhard

A team is using `databricks.agents.deploy()` to publish a Mosaic AI Agent to a serving endpoint. They must expose an environment-specific Vector Search index name and the endpoint name to the deployment without hardcoding values in the notebook, and the same notebook must run in dev and prod. Which approach should the engineer use?

A.Store the values in a Delta table and have the notebook query that table at deploy time.
B.Read the values from `dbutils.widgets` defined at the top of the notebook and pass them into `databricks.agents.deploy()`.
C.Detect the environment by reading the current workspace URL inside the notebook and branch on it.
D.Hardcode the prod values and use a separate copied notebook for dev with the dev values.
AnswerB

Notebook widgets provide parameterized inputs that a Databricks job or bundle can supply per environment, so the same notebook runs in dev and prod with different index and endpoint names. The values flow into `databricks.agents.deploy()` at runtime without editing code. This is the standard Databricks pattern for environment-parameterized deployments and satisfies the no-hardcoding requirement.

Why this answer

The cleanest way to run one notebook across dev and prod is to parameterize it with notebook widgets, which jobs and Databricks Asset Bundles can populate per target. The values then flow into `databricks.agents.deploy()` so the Vector Search index and endpoint names differ by environment without code changes. Duplicated notebooks, host-string branching, and Delta-table config stores all introduce drift or fragility that the scenario's requirements exclude.

Exam trap

The trap here is reaching for ad hoc environment detection instead of using Databricks' built-in parameterization for deployment targets.

162
MCQmedium

Which of the following describes the 'drift' phenomenon in the context of LLM monitoring?

A.The model's weights change during inference.
B.A shift in user query patterns over time.
C.The model generating the same answer too often.
D.Increased latency in the serving endpoint.
AnswerB

Drift in generative AI often manifests as a divergence in user query patterns, where the model encounters prompts that are significantly different from its training or fine-tuning data. This shift leads to degradation in performance, as the model was not optimized for these new interaction patterns or domains.

Why this answer

LLM drift occurs when the input data distribution or the nature of user queries changes over time, causing the model to perform worse than during validation. Because user expectations and language usage evolve, a model that performed well at launch may become less accurate as queries drift into new domains or change in tone. Monitoring this drift is essential for proactive maintenance and model updating.

Exam trap

Candidates often confuse LLM drift with traditional feature drift, missing that LLM drift specifically refers to evolving user query patterns and intents over time.

163
MCQeasy

A GenAI engineer has registered a fine-tuned Llama 3 model in Unity Catalog as `prod.ml.models.support_llm` and now needs to expose it as a REST endpoint for a customer-facing chatbot. The team wants Databricks to manage the serving infrastructure, GPU autoscaling, and version upgrades with minimal operational overhead. Which action should the engineer take?

A.Publish the model artifact to a Delta table and query it with a SQL warehouse.
B.Deploy the model to an all-purpose cluster and expose it through a notebook's REST API calls.
C.Register the model in the workspace model registry and call it from a Databricks job schedule.
D.Create a Mosaic AI Model Serving endpoint that serves the Unity Catalog model version.
AnswerD

Mosaic AI Model Serving is the managed path for exposing Unity Catalog-registered models as REST endpoints. It handles GPU provisioning, autoscaling, and workload isolation, and can serve a specific registered model version with automatic scale-to-zero and rolling upgrades. Because the model already lives in Unity Catalog, the engineer only needs to create the endpoint and point it at the model version, satisfying the low-overhead requirement.

Why this answer

Serving a Unity Catalog-registered model through Mosaic AI Model Serving is the standard Databricks pattern for production GenAI endpoints. It provides managed GPU compute, autoscaling, and versioned deployments without the engineer building infrastructure. The other choices either use batch or interactive compute that lacks an inference endpoint, or confuse storage with serving, so none can deliver a low-latency REST API for the chatbot.

Exam trap

The trap here is assuming any Databricks compute that can load the model can also serve it as a production REST endpoint.

164
MCQhard

A team maintains a Mosaic AI Agent application whose endpoint must call an external LLM provider through a secret stored in a Databricks secret scope. During a deployment pipeline run, the endpoint build step fails while resolving the credential, even though the secret scope exists and the notebook test works. Which configuration should the engineer verify first?

A.That the secret scope is replicated into the Unity Catalog metastore attached to the serving endpoint.
B.That the endpoint has been configured with a personal access token embedded directly in the model signature.
C.That the serving endpoint's service principal or the deploying user has READ permission on the secret scope and that the code reads the secret at load time using the correct scope and key.
D.That the secret scope is mounted as a Databricks secret scope in the serving endpoint's environment variables.
AnswerC

Serving endpoints run under a service principal or the deploying identity, and secret access requires explicit READ permission on the scope. If the notebook author had access but the endpoint identity does not, or the scope or key name is wrong in the model code, credential resolution fails during load. Verifying permissions and the exact scope/key names targets the most common cause.

Why this answer

Credential resolution in Model Serving depends on the identity that loads the model having READ access to the secret scope and on the code requesting the correct scope and key. A notebook author with access may succeed while the endpoint identity fails, and a typo in scope or key produces the same symptom. Checking permissions and the retrieval call resolves the failure without restructuring the deployment.

Exam trap

The trap here is assuming a working notebook proves the endpoint identity can read the secret, when serving runs under a different principal with its own ACLs.

165
MCQmedium

Which design pattern is best for protecting the LLM from prompt injection attacks when building a customer-facing chatbot on Databricks?

A.Embedding all user inputs into the vector database.
B.Using a guardrail service to validate and sanitize user inputs.
C.Relying on the LLM’s internal safety training to reject injections.
D.Sending all inputs through a standard SQL query first.
AnswerB

Guardrail services or classification models act as a security gateway, analyzing inputs for patterns indicative of injection attacks. By blocking or sanitizing these inputs before they are passed to the primary LLM, the system prevents unauthorized instructions from executing, which is a foundational requirement for securing public-facing chatbots.

Why this answer

Implementing an input-filtering layer using a separate, lightweight classification model or guardrail service effectively detects and blocks malicious prompts before they reach the LLM. This proactive defense is vital for securing AI applications, as prompt injection can lead to unauthorized data access or malicious behaviors. This layering approach creates a defense-in-depth architecture that keeps the generative model safe from adversarial inputs.

Exam trap

Candidates often assume that prompt engineering or system instructions are sufficient to prevent injection, forgetting that malicious users can easily bypass these through prompt manipulation.

166
MCQeasy

Which component of the Databricks platform allows you to bundle together notebooks, model serving configurations, and pipeline definitions for repeatable deployment?

A.Databricks SQL Warehouse.
B.Unity Catalog.
C.Databricks Asset Bundles.
D.MLflow Tracking.
AnswerC

Databricks Asset Bundles (DABs) are specifically engineered to package together code, infrastructure, and configuration into a single deployable unit. They provide a declarative way to define resources, making it easy to version, test, and deploy complex applications using standard CLI commands and CI/CD best practices.

Why this answer

Databricks Asset Bundles (DABs) are the standardized way to manage and deploy project artifacts. By bundling these components, teams can maintain version control over their entire application ecosystem. This approach is fundamental for implementing robust CI/CD pipelines, ensuring that the development, testing, and production environments remain synchronized and predictable throughout the release cycle.

Exam trap

Test-takers frequently choose workspace files or generic Git integration tools instead of Databricks Asset Bundles, missing that DABs are specifically designed to bundle and deploy multi-component projects declaratively.

167
MCQhard

A healthcare analytics team deployed a RAG assistant on Databricks whose answers cite clinical policy documents. Compliance requires that every production answer be attributable to a specific retrieved chunk. During evaluation, the team notices the groundedness judge scores are high, but manual review finds answers that blend two policies into a statement neither document supports. Which change to their Mosaic AI Agent Evaluation configuration best detects this failure mode?

A.Increase the number of retrieved chunks per query from five to twenty so the model has more supporting evidence available.
B.Enable the chunk-level citation assessment so each claim in the response is checked against the specific retrieved chunk it references, rather than scoring the response against the concatenated context.
C.Raise the temperature setting on the served model to zero and re-run the evaluation suite to obtain more deterministic outputs.
D.Add a relevance judge that scores the retrieved chunks against the user question and filter out any chunk scoring below a fixed threshold.
AnswerB

Response-level groundedness against merged context can pass when each claim is individually supported somewhere in the pool, even if the response attributes a blended statement to the wrong source. Chunk-level citation assessment forces per-claim traceability to the cited chunk, which is exactly the compliance requirement and surfaces cross-document blending that aggregate groundedness masks.

Why this answer

The compliance requirement is per-answer attribution to a specific chunk, so the evaluation must verify each claim against the chunk it cites. Scoring the whole response against concatenated context allows a blended claim to pass whenever its components appear somewhere in the pool. Chunk-level citation assessment closes that gap by testing attribution directly, which is what distinguishes this failure from ordinary ungrounded hallucination.

Exam trap

The trap here is assuming a high groundedness score means every claim is properly attributed, when response-level groundedness against merged context can pass blended statements whose parts each appear somewhere in the retrieved pool.

168
MCQmedium

A GenAI team at a retail bank runs a RAG assistant on a Databricks Mosaic AI Model Serving endpoint. During a pilot, they captured end-user thumbs-up/down feedback in a Delta table but never joined it to the trace payloads. Six weeks later, hallucination complaints spike, yet the aggregate thumbs-down rate is unchanged. Which approach best resolves this discrepancy using Databricks-native tooling?

A.Increase the endpoint's provisioned concurrency and enable autoscaling so responses are returned faster during peak hours.
B.Switch the judge from a smaller open model to a larger frontier model and re-run the nightly evaluation job over the last 30 days.
C.Retrain the embedding model on the latest product documentation and redeploy the index to the Vector Search endpoint.
D.Join the Delta feedback table to MLflow traces on request_id and segment quality metrics by user cohort, prompt template version, and retrieved document source.
AnswerD

Joining feedback rows to MLflow traces via the shared request_id lets the team slice the unchanged aggregate rate by the dimensions that actually moved, exposing whether the spike is confined to one prompt template version or document source. Aggregates hide this because a large silent-majority cohort with stable thumbs-up dilutes the regressed segment.

Why this answer

Aggregate feedback rates are diluted by cohort mix, so an unchanged overall thumbs-down percentage can coexist with a severe regression inside one prompt template version, document source, or user cohort. Correlating the stored human feedback with MLflow traces on request_id and then slicing by those dimensions surfaces where the hallucination spike actually lives, which is the prerequisite for any targeted fix.

Exam trap

The trap here is treating an unchanged aggregate feedback rate as proof the application is stable, when the real issue is that feedback was never correlated with trace attributes so segment-level regressions stay invisible.

169
MCQeasy

An AI engineer is using MLflow to track experiments for a generative AI application. They want to log parameters, metrics, and artifacts for each run, and later compare runs to select the best model. Which MLflow component should they use to organize runs into a named group for a specific project?

A.MLflow Model Registry
B.MLflow Run
C.MLflow Experiment
D.MLflow Tracking Server
AnswerC

An MLflow Experiment is a logical grouping of runs for a specific project or objective. It allows you to organize and compare runs, view metrics, and manage artifacts. By creating an experiment, you can log multiple runs and later query them to find the best model based on metrics. This is the core organizational unit in MLflow.

Why this answer

MLflow Experiments are designed to group runs for a specific project. They allow you to log and compare multiple runs, making it easy to track progress and select the best model. The experiment is the primary organizational unit in MLflow, and it is where runs are created and stored.

Exam trap

The trap here is confusing the organizational unit (Experiment) with the execution unit (Run) or the model management component (Model Registry), which serve different purposes.

170
Multi-Selecthard

A GenAI engineer is preparing a large text corpus for fine-tuning an LLM. The corpus contains many near-duplicate documents and documents in multiple languages. They need to reduce redundancy and ensure language consistency. Which two steps should be performed during data preparation? (Choose two.)

Select 2 answers
A.Use the language detection function from the spark-nlp library to filter documents to a single target language.
B.Train a custom tokenizer on the entire corpus to better handle multilingual text.
C.Use Delta Lake time travel to revert to a previous version of the dataset if duplicates are found.
D.Compute MinHash signatures and apply locality-sensitive hashing to identify and remove near-duplicate documents.
E.Apply a fixed-size chunking strategy to split all documents into 512-token segments before deduplication.
AnswersA, D

Language detection identifies the language of each document, allowing you to filter out documents that do not match the target language. This ensures consistency and prevents the model from learning from irrelevant languages, which is critical when fine-tuning for a specific language task.

Why this answer

To reduce redundancy and ensure language consistency, the engineer should use MinHash with LSH to remove near-duplicate documents and apply language detection to filter to the target language. These steps directly address the stated goals and are scalable for large corpora.

Exam trap

The trap here is confusing tokenization or versioning features with actual data cleaning steps, leading to choices that do not remove duplicates or filter languages.

171
Multi-Selectmedium

Which THREE of the following are common challenges when monitoring LLM applications in production that differ significantly from traditional ML monitoring? (Choose three)

Select 3 answers
A.Non-deterministic output behavior.
B.Difficulty in defining a single 'ground truth' for responses.
C.Lack of high-throughput API endpoints.
D.High cost of manual evaluation at scale.
E.The model's inability to connect to internet data.
AnswersA, B, D

LLMs can produce different outputs for the same prompt due to temperature settings or inherent stochasticity. This makes simple equality checks useless for evaluation. Monitoring must account for this variance, requiring probabilistic or semantic similarity metrics rather than static, exact-match validation common in traditional classification tasks.

Why this answer

LLMs introduce unique challenges because their outputs are non-deterministic, high-dimensional, and often lack a single 'correct' answer. Traditional monitoring focuses on numerical drift and binary classification metrics like precision/recall. LLM evaluation requires semantic understanding, nuance assessment, and the ability to handle unstructured text, necessitating specialized tools like LLM-as-a-judge or human-in-the-loop validation to manage the subjectivity inherent in generative AI outputs.

Exam trap

Candidates often apply traditional ML monitoring assumptions, failing to account for LLM non-determinism, subjective evaluation, and the absence of a single ground truth.

172
MCQeasy

Which Databricks feature allows developers to track experiment parameters, model artifacts, and evaluation metrics in a structured way?

A.Unity Catalog
B.MLflow Tracking
C.Delta Lake
D.Databricks SQL
AnswerB

MLflow Tracking provides a robust API for logging experiments, including parameters, metrics, and models. This functionality allows developers to systematically track their progress, reproduce results, and compare different iterations, which is essential for managing the complexity of model development and ensuring that model quality is documented and verifiable.

Why this answer

MLflow Tracking is the industry-standard component integrated within Databricks for recording experiments. It allows developers to log parameters (e.g., learning rate), metrics (e.g., accuracy), and artifacts (e.g., model weights) during the development process. This structured approach is vital for reproducibility and comparing model versions, ensuring that the most effective model is identified before promotion to production, which is a key requirement for any serious ML engineering project.

Exam trap

Candidates often confuse MLflow with general project management tools. MLflow Tracking is the specific, structured component for recording parameters and metrics necessary for model reproducibility and comparison.

173
MCQmedium

A GenAI engineer is developing a RAG application that uses a Vector Search index. The index is created from a Delta table that contains sensitive information. The security team requires that the engineer can audit which users have queried the index and what data they retrieved. Which Unity Catalog feature should the engineer enable to capture this information?

A.Unity Catalog lineage
B.Vector Search index metadata
C.Inference tables on the Vector Search endpoint
D.Audit logs in system tables
AnswerD

Audit logs in system tables capture access and query events across Databricks, including Vector Search index queries. They record the user identity, the action performed, and the timestamp. By querying the audit logs, the engineer can audit who queried the index and when, satisfying the security requirement. This is the standard mechanism for auditing access to Unity Catalog objects.

Why this answer

Audit logs in system tables capture all access and query events, including those against Vector Search indexes. They record the user, action, and timestamp, enabling the engineer to audit who queried the index and when. This is the appropriate Unity Catalog feature for meeting the security team's auditing requirement.

Exam trap

The trap here is assuming that inference tables or lineage provide user-level query auditing, but only audit logs capture that information.

174
MCQeasy

A generative AI engineer registers a model in Unity Catalog and wants a downstream application to call it for inference without granting the application broad workspace access. The engineer is told to grant a specific Unity Catalog privilege on the registered model so the application can invoke it. Which privilege should be granted?

A.SELECT
B.MODIFY
C.EXECUTE
D.USE CATALOG
AnswerC

EXECUTE is the Unity Catalog privilege that authorizes a principal to invoke a registered model or function. Granting EXECUTE on the model to the application's principal allows inference calls while keeping catalog and schema browsing rights separate, which is the least-privilege path for serving a model registered in Unity Catalog.

Why this answer

Registered models in Unity Catalog are invoked under the EXECUTE privilege, which authorizes a principal to call the model without granting ownership or modification rights. Granting EXECUTE to the application's identity, alongside any needed USE CATALOG and USE SCHEMA to resolve the path, provides least-privilege inference access for the downstream application.

Exam trap

The trap here is assuming table-style privileges like SELECT also govern model invocation, when models require EXECUTE.

175
Multi-Selectmedium

A GenAI engineer is instrumenting a production RAG application on Databricks to detect quality degradation before users complain. The team wants signals that reveal problems in the retrieval stage specifically, rather than in the generation stage. Which TWO signals should they track? (Choose two.)

Select 2 answers
A.Recall of the retrieval step measured against a curated question-to-document mapping
B.Toxicity score of the final generated answer
C.Average completion token count per response
D.Time to first token on the serving endpoint
E.Precision at k of the retrieved chunks against a labeled relevance set
AnswersA, E

Recall against a curated mapping reveals whether the chunks needed to answer each question were retrieved at all. When recall falls, the generator is starved of the correct evidence and cannot produce grounded answers. This is a retrieval-stage metric by construction and localizes failures to indexing, chunking, or embedding rather than to the LLM.

Why this answer

Retrieval-stage quality is measured by whether the right documents are found and how many of the returned ones are relevant. Precision at k and recall against a curated question-to-document mapping capture those two dimensions directly, so regressions localize to indexing, chunking, or embeddings. Token count, time to first token, and toxicity all describe the generator or the serving layer, not retrieval.

Exam trap

The trap here is treating any production signal, such as latency or output toxicity, as evidence about retrieval quality, when only metrics computed against the retrieved chunk set can isolate the retrieval stage.

176
MCQmedium

A machine learning engineer needs to deploy a real-time Model Serving endpoint. Which Databricks construct is required to manage the model's environment, dependencies, and artifacts while ensuring version control for the deployment?

A.Delta Live Tables pipelines
B.MLflow Models within Unity Catalog
C.Databricks Job clusters
D.Workspace-local Model Registry
AnswerB

MLflow Models registered within Unity Catalog provide a centralized repository for model artifacts and metadata. This integration enables Model Serving to access validated model versions, manage environment configurations, and track the full lineage of the model, which is essential for secure and reliable production deployments.

Why this answer

Databricks Model Serving relies on the Unity Catalog Model Registry to manage model versions and artifacts. By using Unity Catalog, organizations ensure governance, lineage, and consistent environment reproduction. This is critical for production deployments as it decouples the model development lifecycle from the serving infrastructure, ensuring that the exact artifact deployed in staging is the one running in production, thereby maintaining model integrity and reproducibility across environments.

Exam trap

Candidates frequently choose raw MLflow runs or workspace artifact paths, failing to realize that production serving mandates version-controlled MLflow Models stored within the Unity Catalog.

177
MCQeasy

Why should you integrate Mosaic AI Model Evaluation with Unity Catalog?

A.To increase the model's training speed.
B.To provide centralized governance and lineage.
C.To automatically delete old evaluation data.
D.To bypass the need for prompt engineering.
AnswerB

Unity Catalog provides a single source of truth for governance, lineage, and access controls. Integrating evaluation with it ensures that all model performance data can be linked to specific training and data versions, providing the full transparency and accountability required for enterprise-grade AI risk management.

Why this answer

Unity Catalog acts as the central governance layer for all Databricks data. By integrating evaluation with Unity Catalog, you ensure that all evaluation datasets, results, and models are discoverable, lineage-tracked, and governed. This is essential for compliance, ensuring that every model deployment is backed by a verifiable audit trail of its evaluation performance, which is a core requirement for enterprise AI deployments that must adhere to strict internal and external standards.

Exam trap

Candidates often focus on 'performance optimization' or 'cost reduction', missing the primary purpose of Unity Catalog integration, which is centralized governance, auditability, and lineage tracking.

178
MCQhard

A team is using MLflow to track their model evaluation runs. They want to ensure that every evaluation run is reproducible. What is the best practice to achieve this in Databricks?

A.Only log the final scalar metrics.
B.Log the model URI, dataset version, and code hash.
C.Hardcode the dataset paths in the notebook.
D.Use a global variable for model versioning.
AnswerB

Logging these identifiers provides a complete trail of the experiment setup. This allows engineers to re-run the exact configuration to verify results or investigate anomalies. It is the standard best practice for maintaining a rigorous, reproducible research and evaluation pipeline in professional Databricks ML environments.

Why this answer

Reproducibility in evaluation requires logging not only the results but also the exact artifacts, environment, and code version used. MLflow's ability to log the model URI, the specific evaluation dataset version, and the code state (via git hash) ensures that an evaluation run can be perfectly recreated. This is mandatory for auditing and verifying improvements in LLM performance over time, preventing 'black box' results that cannot be validated.

Exam trap

Test-takers often think saving just the model weights is sufficient, forgetting that data versioning and code hashes are required for complete reproducibility.

179
MCQeasy

Which Databricks asset is best suited for scheduling and orchestrating a multi-step GenAI pipeline that includes data ingestion, vector index updating, and model evaluation?

A.Databricks Feature Store
B.Databricks SQL Warehouse
C.Databricks Workflows
D.Unity Catalog
AnswerC

Databricks Workflows provides a robust orchestration engine to schedule and run multi-step pipelines. It supports various task types, including notebooks, JARs, and SQL queries, and allows for complex branching and dependencies, making it the ideal tool for orchestrating the end-to-end lifecycle of a GenAI data and model pipeline.

Why this answer

Databricks Workflows is the unified tool for orchestrating multi-step pipelines. It allows developers to define dependencies between tasks, manage retries, and monitor job execution. For GenAI, this is crucial because pipelines often involve sequential dependencies—such as ensuring vector database updates complete before a model is evaluated.

Leveraging Workflows ensures reliability and observability in automated production pipelines, which are essential for maintaining the integrity of data and models in an enterprise environment.

Exam trap

Candidates often select MLflow or Unity Catalog for pipeline scheduling, confusing artifact tracking and governance with the actual task orchestration capabilities provided by Databricks Workflows.

180
Multi-Selecthard

A team is designing a Databricks RAG application that must return answers with citations to source documents. They plan to use Databricks Vector Search and want the LLM to reference specific chunks. Which two design choices are required to produce reliable citations? (Choose two.)

Select 2 answers
A.Increase the LLM temperature to encourage the model to generate more detailed citations.
B.Store a stable document identifier and chunk metadata alongside each embedding in the Vector Search index.
C.Enable inference tables on the Vector Search endpoint to capture citation metadata.
D.Set the Vector Search index to use a cosine similarity metric instead of dot product.
E.Include the retrieved chunk identifiers and source references in the prompt context so the LLM can cite them in its response.
AnswersB, E

Citations require mapping a retrieved chunk back to its source document. Including a stable document ID and chunk metadata in the index payload allows the retriever to return that identifier with each result. Without this, the LLM cannot reference a specific source, and citations become guesses or are omitted entirely.

Why this answer

Reliable citations require two things: the retriever must return source identifiers, and those identifiers must reach the LLM in the prompt. Storing document IDs and chunk metadata in the Vector Search index enables the first, and including them in the prompt context enables the second. Similarity metrics, temperature, and inference tables do not provide the linkage needed for citations.

Exam trap

The trap here is focusing on retrieval tuning knobs such as similarity metrics while overlooking that citations depend on carrying source metadata from the index into the prompt.

181
MCQeasy

A team wants to review what a Databricks Asset Bundle deployment would change in the production workspace before it actually creates or modifies any resources. Which command should the engineer run first?

A.`databricks bundle plan -t prod`
B.`databricks bundle deploy -t prod --dry-run`
C.`databricks bundle validate -t prod`
D.`databricks bundle summary -t prod`
AnswerA

`databricks bundle plan` computes and displays the changes the deployment would make to the target workspace, including resources to create, update, or delete, without actually applying them. Running it with `-t prod` shows the production plan. This is exactly the preview the team needs before committing the deployment, making it the correct first step in a safe release process.

Why this answer

The `databricks bundle plan` command produces a deployment plan that lists the resources the deployment would create, update, or delete in the target workspace, without applying any change. Running it with `-t prod` previews the production deployment. This lets the team review the impact before executing `databricks bundle deploy`, which is the recommended safe release workflow.

Exam trap

The trap here is assuming `databricks bundle deploy` supports a `--dry-run` flag, when the supported way to preview changes is the separate `databricks bundle plan` command.

182
MCQmedium

Refer to the exhibit. The Databricks Vector Search index is currently in an 'ONLINE' status with incremental sync. If a developer performs a DELETE operation on the source Delta table, what happens to the vector index?

A.The record remains in the index until a full re-index is manually triggered.
B.The system throws an error because vector indexes do not support deletes.
C.The record is automatically removed from the index during the next sync cycle.
D.The index enters a 'PENDING_DELETE' state and stops accepting queries.
AnswerC

Incremental sync mode ensures that changes to the Delta table, including deletions, are propagated to the vector index. By monitoring the Delta log, the service identifies the deleted record and updates the index accordingly, keeping the vector store accurate with minimal latency without requiring manual maintenance.

Why this answer

When using incremental synchronization, the Databricks Vector Search service monitors the Delta table's transaction log for changes. A delete operation recorded in the Delta log will trigger a corresponding removal in the index. This ensures the index remains a true representation of the source data, which is critical for compliance and data quality, ensuring the RAG application does not retrieve outdated or removed sensitive information.

Exam trap

Candidates assume deletions require a full manual re-index of the vector store, forgetting that incremental sync automatically processes table transaction logs.

183
Multi-Selectmedium

When migrating a Databricks Asset Bundle (DAB) project from 'development' to 'production', which TWO actions should an engineer perform?

Select 2 answers
A.Update the target in the databricks.yml file to point to the production host
B.Manually recreate all jobs using the Databricks UI in production
C.Ensure production-level secrets are configured in the target environment
D.Delete all development artifacts from the local machine
E.Convert the bundle project into a standard notebook folder
AnswersA, C

Updating the target is essential to direct the deployment to the correct workspace and environment. The 'target' section in the YAML file defines the configuration for the production environment, including the workspace host, ensuring that deployment artifacts are sent to the production workspace rather than development.

Why this answer

Moving code to production requires updating the target configuration to point to production-grade environments and ensuring that environment variables or secrets are correctly scoped. These actions ensure that the production deployment uses hardened settings—such as different cluster types or access permissions—while maintaining the same logic defined in the bundle. This discipline prevents 'environment drift' and ensures that production workloads remain isolated, secure, and performant according to enterprise standards.

Exam trap

Candidates often assume that copying code manually or just changing cluster sizes is enough, missing that DAB targets and environment-specific secrets must be explicitly re-configured.

184
Multi-Selectmedium

Which TWO of the following are benefits of using Delta Lake for data preparation over standard Parquet files on cloud storage?

Select 2 answers
A.ACID compliance for reliable write operations.
B.Automatic conversion of JSON to XML format.
C.Time travel capability for auditing and debugging.
D.Native support for cross-cloud streaming ingestion.
E.Automatic hardware upgrades for compute clusters.
AnswersA, C

ACID compliance ensures that concurrent reads and writes are managed correctly, preventing data corruption during simultaneous operations. In standard Parquet, a failing write operation could leave orphan files or partial data, whereas Delta Lake ensures an 'all-or-nothing' consistency model that is critical for production pipelines.

Why this answer

Delta Lake provides ACID transactions and time travel, which are essential for robust data engineering. ACID transactions ensure that data preparation pipelines never leave tables in a partially written state, while time travel allows for auditing and reverting changes. These features significantly simplify data lifecycle management, reduce the need for custom retry logic, and enhance the overall reliability of data preparation workflows in complex, multi-user production environments.

Exam trap

Candidates often conflate Delta Lake with general data formats, failing to identify that ACID transactions and time travel are specific, functional advantages that distinguish Delta from standard Parquet files.

185
MCQmedium

A developer is creating a custom model serving endpoint that requires an external API call for data enrichment. What is the recommended way to handle sensitive API keys within the Databricks environment?

A.Store keys as environment variables in the notebook code.
B.Use the Databricks Secrets API to manage and retrieve keys.
C.Hardcode the keys directly into the model serving inference function.
D.Use the DBFS root directory to store key-value text files.
AnswerB

The Databricks Secrets API provides a secure, centralized location for managing sensitive information. It allows for role-based access control, ensuring that only authorized users or services can access the secrets, which protects the application from credential leakage and simplifies secret rotation and management across different deployment environments.

Why this answer

Secrets management is the cornerstone of secure application development in Databricks. By using the Databricks Secrets API, developers can decouple sensitive credentials from source code. This practice prevents the accidental disclosure of keys in version control systems and ensures that credentials are injected into the runtime environment securely, which is mandatory for maintaining a secure and audit-compliant AI development lifecycle in enterprise cloud environments.

Exam trap

Test-takers sometimes hardcode API credentials or configuration parameters directly inside the notebook or model artifact, compromising security and compliance standards.

186
MCQeasy

An engineer is designing a GenAI application that must call an external LLM provider's API. The provider key must not appear in notebook code, job logs, or Git. The team already uses Databricks and wants the key to be injected into the serving endpoint at runtime. Which Databricks capability should the engineer use?

A.Hard-coding the provider key in the model signature so the endpoint receives it as a request parameter.
B.Databricks secrets stored in a secret scope, referenced from the serving endpoint configuration using secret references.
C.A Unity Catalog volume containing a text file with the provider key, mounted into the serving endpoint.
D.Environment variables defined in the cluster's Spark configuration and read by the notebook at runtime.
AnswerB

Databricks secret scopes store credentials in an encrypted backend, and secret references let the serving endpoint configuration pull the key at runtime without exposing it in code, logs, or Git. This directly satisfies the requirement that the provider key never appear in notebook code or job logs while still being available to the endpoint.

Why this answer

Databricks secret scopes with secret references keep credentials encrypted and out of source control, logs, and notebooks, while still allowing Model Serving to inject them at runtime. Spark config values, Unity Catalog volume files, and model signature parameters all expose the key through metadata, logs, or request payloads and do not provide runtime injection for serving endpoints.

Exam trap

The trap here is confusing governed storage such as Unity Catalog volumes with secret management, even though volumes are not designed to protect credentials from readers.

187
MCQhard

An engineer is using Mosaic AI Agent Evaluation to score a conversational agent that calls tools. The agent sometimes answers correctly but with fabricated citations. Which evaluation approach best surfaces this specific failure mode?

A.Measure only token usage and latency to detect when the agent is hallucinating citations.
B.Run the agent without tools so it cannot produce citations at all.
C.Enable the groundedness judge and supply the retrieved context so the judge can verify claims against sources.
D.Rely solely on the correctness judge, since a correct answer cannot contain fabricated citations.
AnswerC

The groundedness judge checks whether statements in the response are supported by the provided context. Fabricated citations are exactly the failure mode it detects, because the cited source either does not exist in the context or does not contain the claim. Supplying the retrieved context is required for the judge to perform this verification accurately.

Why this answer

The groundedness judge compares response claims against the retrieved context, which is precisely how fabricated citations are detected. A correctness judge evaluates answers against references and misses unsupported citations, operational metrics say nothing about grounding, and removing tools avoids the scenario rather than evaluating it.

Exam trap

The trap here is conflating answer correctness with grounding, assuming that a factually right answer cannot also contain a fabricated citation.

188
MCQeasy

When building an application that retrieves context from Databricks Vector Search, what is the recommended data format for storing the document chunks?

A.JSON files in an S3 bucket
B.Delta tables
C.Parquet files on DBFS root
D.SQL Server database
AnswerB

Delta tables are the standard and recommended source for Vector Search indexes. They offer optimized performance, built-in change detection, and native compatibility with Databricks infrastructure. This ensures that the vector search index is always synchronized with the source data, which is crucial for building reliable and accurate AI-driven applications.

Why this answer

Delta tables are the native, optimized storage format for Databricks. When using Vector Search, the index is automatically maintained by Delta tables. This provides a performant and reliable way to sync data between the source Delta table and the vector index.

Storing data in Delta ensures that the vector search index can benefit from CDC (Change Data Capture) patterns, keeping the RAG knowledge base automatically updated and consistent.

Exam trap

Many candidates mistakenly select standalone vector databases or generic file formats like CSV/JSON instead of native Databricks storage formats that automatically sync indices.

189
MCQmedium

When deploying a model endpoint, which Databricks feature provides the capability to review and approve the model before it is promoted to production?

A.Delta Live Tables Quality Checks
B.Unity Catalog Model Status Transitions
C.Databricks Job Task Notifications
D.Git branch protection rules
AnswerB

Model status transitions in Unity Catalog allow teams to manage a model's lifecycle through stages like 'Staging' and 'Production'. This is the standard mechanism to control which models are available for production serving, enabling teams to enforce rigorous evaluation and approval gates before promoting a new model version.

Why this answer

The Unity Catalog Model Registry provides a formal lifecycle management process, including status transitions such as 'Staging', 'Production', and 'Archived'. By using these status labels, organizations can implement a human-in-the-loop approval process where models must pass specific validation criteria—or manual reviews—before they are eligible for the production inference endpoint. This promotes safety, quality, and compliance in machine learning deployments by ensuring only vetted models reach end-users.

Exam trap

Test-takers frequently confuse workspace-level permissions or generic Git pull requests with Unity Catalog's formal model lifecycle stage transitions.

190
MCQmedium

A developer needs to monitor the performance of an LLM application in production. They want to track the latency of their model endpoint. Where can they find this metric in the Databricks workspace?

A.In the Unity Catalog table lineage view.
B.In the 'Monitoring' tab of the Model Serving endpoint UI.
C.In the Databricks SQL Warehouse query history.
D.In the workspace audit logs.
AnswerB

The 'Monitoring' tab within the Model Serving endpoint interface is the dedicated location for viewing operational metrics. It displays auto-generated graphs for latency, request volume, and error rates, enabling developers to quickly assess the health and performance of their deployed models without needing additional configuration.

Why this answer

Databricks Model Serving provides built-in monitoring dashboards for every endpoint. These dashboards automatically track key performance indicators such as request latency, throughput, and error rates. Monitoring these metrics is vital for understanding the user experience and detecting potential performance bottlenecks.

By providing this information natively, Databricks helps developers maintain high service quality without requiring external observability tools, streamlining the monitoring and optimization process for deployed models.

Exam trap

Candidates frequently look for external APM tools or cluster logs, failing to recognize that Model Serving endpoints feature a dedicated built-in Monitoring tab.

191
MCQmedium

When building a RAG application, a developer wants to ensure that the retrieved context is strictly limited to documents the user has access to. Where should this security logic be enforced?

A.Inside the prompt engineering layer.
B.At the Vector Search retrieval layer.
C.Within the final response generation phase.
D.At the client-side browser application level.
AnswerB

Enforcing access control during the retrieval phase is the most effective approach. By filtering the documents returned from the vector index based on user identity or group membership, you ensure that the generative model only receives content that the specific user is authorized to view and process.

Why this answer

Access control must be enforced at the retrieval layer, ideally within the vector search engine or via a data-filtering mechanism that respects user identity. Relying on the model to enforce security is ineffective, as models cannot reliably manage authorization. Proper integration with Unity Catalog ensures that data retrieval is governed by the same permissions as the underlying tables, maintaining consistent security across the entire data-to-AI lifecycle.

Exam trap

Many candidates incorrectly assume that security logic should be handled by the LLM prompt or the application layer, ignoring that LLMs are prone to prompt injection and cannot reliably enforce data access.

192
Multi-Selectmedium

An AI engineer is designing a scalable customer support application on Databricks that integrates custom vector search indexes with a fine-tuned LLM. Which TWO architectural components are essential for enabling efficient similarity search and low-latency retrieval within the Databricks ecosystem? (Choose TWO)

Select 2 answers
A.Databricks Vector Search
B.An external third-party Hadoop cluster managed via SSH tunnels
C.Databricks Model Serving
D.A manual cron job running on a local developer laptop
E.A static CSV file stored on local driver disk storage
AnswersA, C

Databricks Vector Search is a serverless vector database natively integrated into the Databricks platform. It automatically synchronizes with Delta tables, manages embedding indexes, and provides high-performance similarity search APIs essential for retrieving contextually relevant documents in RAG applications.

Why this answer

Databricks Vector Search provides serverless vector database capabilities to store and query embeddings efficiently without managing external infrastructure. Combined with Databricks Model Serving for hosting the LLM, these native services ensure high throughput, enterprise-grade security, and low-latency inference required for modern retrieval-augmented generation applications built on the Databricks platform.

Exam trap

Candidates often select generic options like 'Delta Tables' or 'MLflow' without identifying the specific services (Vector Search and Model Serving) required for low-latency RAG performance.

193
MCQmedium

Which Databricks feature is specifically designed to monitor model quality and drift in production?

A.Unity Catalog audit logs.
B.MLflow Model Monitoring.
C.The Databricks SQL query history.
D.Cluster event logs.
AnswerB

MLflow Model Monitoring is the dedicated Databricks component for tracking the health of deployed models. It allows engineers to monitor metrics, detect performance degradation, and identify drift in model outputs. This is essential for maintaining high-quality generative AI applications and responding effectively to changes in data or user behavior.

Why this answer

MLflow Model Monitoring (now part of Mosaic AI) provides the necessary tooling to track model performance metrics and identify drift over time. For generative AI, this is critical because language models can exhibit performance degradation or shifts in behavior when input distributions change. Proactive monitoring allows engineers to detect these shifts early and trigger retraining or fine-tuning, ensuring consistent performance in the production environment.

Exam trap

Candidates often choose generic logging tools or Unity Catalog governance features when asked specifically about tracking production model performance drift.

194
MCQmedium

A GenAI engineer has built a RAG application that queries a Delta table named `prod_ai.knowledge_base.documents` through a Databricks SQL warehouse. The application uses a service principal `sp-rag-prod` to authenticate. The table contains confidential internal documents. The security team wants to ensure that the service principal can only read the table and cannot modify or delete it. Which Unity Catalog privilege should be granted to `sp-rag-prod` on the table?

A.ALL PRIVILEGES
B.USE SCHEMA
C.MODIFY
D.SELECT
AnswerD

SELECT is the minimum privilege required to read data from a table. Granting SELECT to the service principal allows the RAG application to query the table without giving it the ability to insert, update, or delete data. This follows the principle of least privilege and satisfies the security team's requirement that the service principal can only read the table.

Why this answer

The service principal needs only to read the table for the RAG application. Granting SELECT provides exactly that capability without allowing data modification. Other privileges either grant excessive rights (MODIFY, ALL PRIVILEGES) or do not permit reading table data (USE SCHEMA).

SELECT is the correct least-privilege grant for a read-only consumer in Unity Catalog.

Exam trap

The trap here is assuming that USE SCHEMA or ownership is needed for read access, when in fact SELECT alone on the table is sufficient for a service principal to query data.

195
MCQeasy

A data scientist wants to use a Unity Catalog registered model in a GenAI pipeline. The model was trained on sensitive data, and the governance team requires that the model's lineage back to the training dataset be traceable. Which Unity Catalog capability provides this traceability?

A.Unity Catalog lineage graph that automatically tracks data and model dependencies.
B.Model version tags that describe the training data.
C.Workspace notebook comments that document the training data path.
D.Model signature that records input and output schema.
AnswerA

Unity Catalog automatically captures lineage for tables, models, and notebooks, showing how a model version was derived from training datasets. This provides the required traceability without manual effort and is auditable. The lineage graph is a core governance feature that links data assets across the workspace.

Why this answer

Unity Catalog lineage automatically tracks the flow of data through notebooks, jobs, and models, including the training dataset used to produce a model version. This gives the governance team an auditable, automated trace from model to data. Manual methods such as tags or comments are not reliable for compliance.

Exam trap

The trap here is assuming that metadata such as tags or signatures can substitute for automated lineage when proving traceability to an auditor.

196
MCQhard

What is the purpose of the 'metastores.list' command in the Unity Catalog CLI?

A.To list all users registered in the metastore.
B.To list all metastores available to the user in the account.
C.To list all tables within a specific metastore.
D.To list all storage credentials linked to the metastore.
AnswerB

This command specifically retrieves a list of metastore objects. It is the standard way to inspect the metastore topology for an account, helping administrators verify which metastores are provisioned, their IDs, and their region, which is vital for effective cross-region governance management.

Why this answer

The 'metastores.list' command allows administrators to view all Unity Catalog metastores that the current user has permission to see within their Databricks account. This is essential for managing multiple metastores in a complex, multi-region architecture. Understanding which metastores exist and their properties is the first step in ensuring that data governance policies are consistently applied across all regional footprints.

Exam trap

Candidates often confuse Unity Catalog CLI commands with workspace-level CLI commands or assume 'metastores.list' operates on individual catalogs rather than account-level metastores available to the specific user.

197
MCQmedium

A team is preparing data for a RAG system and needs to remove duplicates from a large collection of PDF text extracts. What is the most efficient way to perform de-duplication in Databricks?

A.Use a Python 'set' in a local loop on the driver node.
B.Execute a distributed 'dropDuplicates' operation in Spark.
C.Manually compare every document against every other document using nested loops.
D.Ignore duplicates, as vector databases automatically handle them during insertion.
AnswerB

Spark's 'dropDuplicates' is optimized for large, distributed datasets. It efficiently identifies and removes duplicate rows across the entire cluster, making it the standard approach for large-scale de-duplication. This ensures high-quality training sets and efficient vector indices without requiring complex custom code for distributed data handling.

Why this answer

Using Spark's 'dropDuplicates()' on the text content column or calculating a hash (e.g., MD5) of the text to identify duplicates is highly scalable. This is crucial because redundant information in the vector database can cause the retriever to prioritize duplicate entries, leading to biased results and inefficiency. Removing duplicates ensures that the search index remains focused and that the retrieved context is diverse and informative.

Exam trap

Candidates often attempt single-node Python loops or pandas-based string matching on massive text corpora, resulting in out-of-memory errors on large datasets.

198
MCQmedium

A GenAI engineer is preparing a fine-tuning dataset from a Delta table in Unity Catalog that contains raw user feedback. The feedback text includes irregular capitalization, HTML tags, and excessive punctuation. The engineer needs to normalize the text using Spark NLP within a Databricks notebook, ensuring the pipeline is reproducible and scalable. Which approach should the engineer use to apply this transformation?

A.Use Delta Live Tables to define a streaming pipeline that applies the lower() and trim() functions to the text column.
B.Use the spark-nlp library's DocumentAssembler and Normalizer annotators in a Spark ML Pipeline, and save the pipeline to MLflow.
C.Use Databricks SQL's regexp_replace function in a SELECT statement to clean the text, and create a new table with the cleaned data.
D.Use pandas UDFs with Python's re module to apply custom cleaning functions to each row, and cache the resulting DataFrame.
AnswerB

Spark NLP provides DocumentAssembler and Normalizer annotators that can be combined in a Spark ML Pipeline, enabling scalable and reproducible text normalization. Saving the pipeline to MLflow ensures versioning and reproducibility. This approach leverages Spark's distributed processing and integrates with Databricks workflows, making it ideal for preparing large-scale fine-tuning datasets.

Why this answer

The correct approach uses Spark NLP's DocumentAssembler and Normalizer within a Spark ML Pipeline, which provides distributed, reproducible text normalization. Saving the pipeline to MLflow ensures version control and reusability. This method is scalable and integrates well with Databricks, making it suitable for preparing large fine-tuning datasets with complex cleaning needs.

Exam trap

The trap here is assuming that simple string functions or regex alone suffice for comprehensive text normalization, overlooking the need for specialized NLP annotators and pipeline reproducibility.

199
MCQeasy

A data science team wants to expose a RAG chain as a REST API so that an external web application can send questions and receive answers. The chain is developed with Databricks LangChain integrations and must be deployed with autoscaling and built-in monitoring. Which Databricks capability should they use?

A.A Databricks job that runs the chain on a schedule and writes answers to a Delta table.
B.Model Serving with a custom MLflow pyfunc model that wraps the chain.
C.An all-purpose cluster with a Flask app started manually on the driver node.
D.A Databricks SQL warehouse with a query that calls the chain through a UDF.
AnswerB

Databricks Model Serving deploys MLflow models as scalable REST endpoints and supports custom pyfunc models, which can wrap a LangChain chain. It provides autoscaling, request logging, and integration with inference tables for monitoring, matching the requirement to serve the chain as an API with operational visibility.

Why this answer

Model Serving is the Databricks capability for hosting models and custom Python logic as REST endpoints with autoscaling and monitoring. Wrapping the LangChain chain in an MLflow pyfunc model makes it deployable through that service, so the web application can call a stable API. Scheduled jobs, SQL warehouses, and manual Flask apps do not provide the same managed serving characteristics.

Exam trap

The trap here is equating any Databricks compute that can run Python with a production API endpoint, overlooking that Model Serving is the managed serving layer.

200
MCQmedium

A data engineer is preparing a dataset for fine-tuning a chat model. The dataset contains conversations with alternating user and assistant messages. They need to format each conversation into a single string with special tokens indicating roles. Which approach is most appropriate in Databricks?

A.Use the format_string function to insert role tokens into a template string for each message.
B.Use the explode function to flatten the messages and then collect_list to reassemble them with role tokens.
C.Use the ai_query function to call an LLM that formats the conversation into a single string.
D.Use the concat_ws function to join messages with a delimiter, and prefix each message with a role token.
AnswerD

concat_ws can join an array of strings with a delimiter. By first transforming each message into a string with a role prefix, you can create a formatted conversation string. This is a straightforward and scalable way to prepare chat data for fine-tuning.

Why this answer

The most appropriate approach is to use concat_ws to join messages after prefixing each with a role token. This creates a single formatted string per conversation efficiently and deterministically, which is ideal for fine-tuning chat models.

Exam trap

The trap here is overcomplicating the task by using an LLM or misusing aggregation functions, when a simple string concatenation function suffices.

201
MCQmedium

You need to ingest data from an external JSON source into a Delta table. The source schema is inconsistent. Which strategy is most effective for preparing this data?

A.Use inferSchema = true in the read configuration.
B.Define a rigid schema at the ingestion point.
C.Implement a bronze-silver medallion pattern.
D.Convert the JSON to CSV before loading into Delta.
AnswerC

The bronze-silver pattern enables ingestion of raw, unstructured data into a staging area, followed by rigorous cleaning and validation in a downstream silver table. This approach prevents pipeline failures during ingestion and ensures that cleaning logic is decoupled from the data acquisition layer, allowing for better maintainability.

Why this answer

The 'bronze-to-silver' pattern is the industry standard in Databricks for handling inconsistent data. By loading raw JSON into a 'Bronze' table with a schema of type 'string' (or a single JSON column), you preserve all information for audit. You then use 'Silver' tables to perform schema enforcement, data cleaning, and type casting.

This separation of concerns allows for robust error handling and iterative refinement of the data cleaning logic.

Exam trap

Candidates often suggest cleaning data directly into a final table, overlooking the importance of the Bronze layer for preserving raw historical data before applying schema enforcement in Silver.

202
MCQmedium

A team is deploying a GenAI application using Mosaic AI Model Serving. They want to ensure that the endpoint can handle sudden spikes in traffic without dropping requests. Which feature should they configure?

A.Request batching
B.Autoscaling
C.Model versioning
D.Provisioned throughput
AnswerB

Autoscaling dynamically adjusts the number of concurrent requests the endpoint can handle based on traffic. It scales resources up during spikes and down during lulls, ensuring requests are not dropped. This is the correct feature to handle sudden traffic increases while optimizing cost.

Why this answer

Autoscaling is designed to automatically adjust the serving endpoint's capacity based on incoming traffic. It ensures that the endpoint can handle sudden spikes by adding more resources and scales down when traffic decreases. This maintains availability and performance without manual intervention, making it the correct choice for handling variable loads.

Exam trap

The trap here is confusing provisioned throughput with autoscaling; provisioned throughput is fixed and does not automatically scale during unexpected spikes.

203
Multi-Selecthard

A generative AI engineer is designing a RAG application on Databricks that uses a foundation model served via Databricks Model Serving. The application must handle peak loads gracefully and provide consistent response times. The engineer is evaluating design patterns for scaling and reliability. Which TWO design choices should the engineer implement? (Choose two.)

Select 2 answers
A.Configure the Model Serving endpoint with autoscaling enabled and set a minimum and maximum number of concurrent requests per replica.
B.Deploy the model to a single large GPU replica with maximum memory to handle all traffic without autoscaling.
C.Implement client-side retries with exponential backoff and a timeout for each request to the serving endpoint.
D.Precompute all possible user questions and their answers and store them in a lookup table for instant retrieval.
E.Cache the LLM responses for identical user queries in a Delta table and serve them without invoking the model.
AnswersA, C

Autoscaling allows the endpoint to add or remove replicas based on load, which helps maintain consistent latency during peaks. Setting concurrency bounds per replica controls how many requests each replica handles, preventing overload. Together they balance throughput and response time. This is a supported configuration for Databricks Model Serving and directly addresses the reliability and scaling requirements.

Why this answer

To handle peak loads and maintain consistent response times, the endpoint must scale elastically and the client must handle transient failures. Autoscaling with concurrency limits per replica ensures the serving infrastructure adapts to load, while client-side retries with backoff and timeouts improve resilience. These two choices together provide both server-side elasticity and client-side fault tolerance, which are essential for production reliability.

Exam trap

The trap here is focusing on caching or precomputation as primary scaling strategies, when the core requirements are elastic endpoint scaling and resilient client behavior under load.

204
MCQeasy

A data engineer needs to persist a prepared instruction-tuning dataset so that downstream fine-tuning jobs can read it with ACID guarantees, time travel, and schema enforcement, and so that Unity Catalog can track column-level lineage. Which storage format and registration should be used?

A.Write the dataset as an ORC table in the Hive metastore.
B.Write the dataset as a Delta table registered in Unity Catalog.
C.Write the dataset as CSV files to an external location registered in Unity Catalog.
D.Write the dataset as Parquet files in a Unity Catalog volume.
AnswerB

Delta Lake provides ACID transactions, time travel, and schema enforcement on the prepared dataset, and registering the table in Unity Catalog enables column-level lineage and fine-grained governance. This combination directly satisfies every stated requirement, including the ability for downstream fine-tuning jobs to read a consistent, versioned table.

Why this answer

Delta Lake on Unity Catalog is the only option that simultaneously delivers ACID transactions, time travel, schema enforcement, and column-level lineage for the prepared dataset. Parquet, CSV, and ORC files lack transactional table semantics, and the Hive metastore does not provide the Unity Catalog lineage the team needs.

Exam trap

The trap here is treating a governed file path as equivalent to a governed table, when lineage and ACID properties come from the table format and catalog registration.

205
MCQeasy

A developer is packaging a GenAI chat application as an MLflow model that will be deployed to a Mosaic AI Model Serving endpoint. The application needs to load a retrieval index and a prompt template at startup so the first request is not slowed by initialization. Which MLflow logging pattern should the developer use?

A.Log the application with mlflow.pyfunc.log_model and perform the heavy initialization in the model's load_context method so resources are ready before predict is called.
B.Use mlflow.pyfunc.log_model with a signature only, and rely on the serving environment's default warm-up to populate the index and template automatically.
C.Log only the prompt template as an artifact and let the Model Serving endpoint fetch the index from Unity Catalog on each request.
D.Log the application with mlflow.pyfunc.log_model and load the index and template inside the predict function on every request.
AnswerA

MLflow's PyFunc flavor invokes load_context when the model is loaded into the serving process, before any predict call. Placing index and template initialization there ensures the endpoint pays the cost once per replica, and subsequent requests reuse the loaded objects. This is the documented pattern for stateful resources in MLflow models deployed to Mosaic AI Model Serving.

Why this answer

MLflow's PyFunc flavor calls load_context when the model is loaded, which happens before the endpoint begins serving requests. Initializing the retrieval index and prompt template there ensures they are ready once per replica and reused across calls. Loading inside predict or fetching resources per request would add latency, and a signature alone cannot prepare application state.

Exam trap

The trap here is assuming that serving platforms automatically warm up custom resources, when in fact the model author must initialize them in load_context.

206
MCQmedium

Refer to the exhibit. What is the correct way to log a custom RAG chain so it can be loaded using the provided code?

A.mlflow.log_artifact(local_path="chain.py")
B.mlflow.pyfunc.log_model(artifact_path="my_model", python_model=my_chain_object)
C.mlflow.register_model("path/to/chain")
D.mlflow.sklearn.log_model(my_chain_object)
AnswerB

This method correctly logs a custom Python object (the RAG chain) as an MLflow model. By defining the chain as a Python model, it becomes compatible with the MLflow Model Serving platform, allowing for seamless deployment to production endpoints where it can be invoked via standard API calls.

Why this answer

To load a custom chain using the 'models:/' URI, the chain must be logged as an MLflow model. Using 'mlflow.pyfunc.log_model' with the appropriate artifact parameters allows the entire RAG pipeline—retrieval logic, prompt handling, and LLM inference—to be captured as a single deployable unit. This ensures consistency between local development and production serving environments.

Exam trap

Candidates often confuse 'log_model' with 'save_model' or try to log raw files manually. Using 'mlflow.pyfunc.log_model' is the specific standard for capturing custom RAG chains as deployable artifacts.

207
Multi-Selectmedium

When deploying a model to Databricks Model Serving, which THREE of the following are best practices to ensure production reliability?

Select 3 answers
A.Deploying the 'latest' alias for all production endpoints.
B.Configuring autoscaling policies based on traffic patterns.
C.Using static credentials embedded directly in the deployment YAML.
D.Pinning to a specific registered model version.
E.Implementing VPC peering or Private Link for secure connectivity.
AnswersB, D, E

Autoscaling ensures that the model serving endpoint can dynamically adjust its compute resources based on request volume. This improves performance during peak traffic and manages costs during idle times, contributing significantly to the overall stability and operational efficiency of the model serving infrastructure.

Why this answer

Production reliability for model serving requires careful attention to environment configuration, scalability, and security. Best practices include using specific model versions to prevent unexpected updates, configuring autoscaling to handle variable traffic, and utilizing private connectivity to keep data transmission secure. These steps minimize downtime and ensure that the serving endpoint behaves predictably under various load conditions, which is critical for enterprise-grade generative AI applications.

Exam trap

Candidates often overlook 'pinning' to a specific model version, which risks production instability if the 'latest' version is updated or changed unexpectedly during the deployment process.

208
MCQeasy

Which Unity Catalog object is used to link a specific cloud storage path to a catalog, schema, or table, allowing users to create tables without managing individual storage credentials?

A.Storage Credential
B.External Location
C.Managed Table
D.Unity Catalog Volume
AnswerB

The external location object defines the specific path in cloud storage and associates it with a storage credential. It is the primary mechanism for accessing data in Unity Catalog that is stored in user-managed cloud accounts, providing a governed interface for data engineers to register data assets.

Why this answer

An External Location in Unity Catalog acts as a secure bridge between Databricks and a specific path in cloud storage (e.g., S3 or ADLS). By using a Storage Credential, the external location allows data engineers to manage storage access at a high level. This simplifies data governance by abstracting cloud-specific permissions into Databricks-native access controls.

Exam trap

Candidates often confuse 'External Location' with 'Storage Credential', failing to distinguish between the object that maps a path and the object that holds the authentication secret.

209
MCQeasy

A data science team wants to expose a fine-tuned LLM as a REST API for an internal web application. They have registered the model in Unity Catalog and want Databricks to manage the serving infrastructure, autoscaling, and request routing. Which Databricks capability should they use?

A.A Delta Live Tables pipeline that materializes model outputs into a table the web app polls.
B.Databricks Model Serving, which creates a managed endpoint that autoscales and routes requests to the model.
C.A Databricks job cluster running a notebook that starts a Flask server on the driver.
D.A Databricks SQL warehouse with a custom HTTP connection to the model.
AnswerB

Model Serving provides managed endpoints with autoscaling, load balancing, and a REST API, and it loads models directly from Unity Catalog. It matches every stated requirement: managed infrastructure, autoscaling, request routing, and REST access for the web application without the team operating servers.

Why this answer

Databricks Model Serving is the managed capability that turns registered models into autoscaling REST endpoints with routing and load balancing. The other options either run unmanaged code on clusters, execute SQL, or precompute batch outputs, none of which deliver interactive, managed model inference for a web application.

Exam trap

The trap here is equating any compute that runs model code with a managed serving endpoint, when only Model Serving provides autoscaling REST inference.

210
MCQmedium

A generative AI engineer is designing a Databricks RAG application that ingests a Delta table containing 400 million support articles into a Databricks Vector Search index. The team wants the lowest-latency online serving with the smallest possible index while preserving retrieval quality for the most common queries. Which design decision best meets these requirements?

A.Store all 400 million embeddings in a Unity Catalog volume and query them directly with a Python UDF at request time.
B.Replace Databricks Vector Search with a keyword-only BM25 index over the raw article text to avoid embedding storage costs.
C.Use Databricks Vector Search with a Delta Sync index, and configure a smaller index by filtering out low-traffic or stale articles before syncing so only high-value documents are embedded.
D.Enable the vector index with the vector search endpoint sized for the full 400 million rows and keep all embeddings in a single index.
AnswerC

Filtering the source Delta table to high-value, frequently accessed articles before the Delta Sync index is built reduces the number of embeddings stored and searched, which lowers latency and index size. Vector Search Delta Sync indexes stay current with the source table, so the team can prune low-traffic content while keeping retrieval quality for the common queries the business cares about.

Why this answer

Pruning low-value content before building a Delta Sync index directly reduces the number of embeddings stored and searched, which shrinks the index and lowers serving latency. Because Vector Search Delta Sync indexes continue to track the source Delta table, the team retains freshness and can still serve high-quality results for the queries that matter most, matching both the performance and quality goals.

Exam trap

The trap here is assuming that a RAG application must always embed every row of the source table, when pre-filtering the Delta table before the Delta Sync index is created is a valid way to shrink the index and improve latency.

211
MCQmedium

A team is deploying a Retrieval-Augmented Generation application as a Mosaic AI Agent using `databricks.agents.deploy()`. The agent must query a Vector Search index and a Delta table in Unity Catalog. During testing, the endpoint returns permission errors when accessing those resources, even though the deploying user has access. Which configuration should the engineer apply to resolve this?

A.Disable Unity Catalog enforcement on the endpoint so it can read resources with the deploying user's identity.
B.Grant the deploying user's personal access token to the endpoint and store it in the model signature.
C.Set the endpoint's environment variable `DATABRICKS_TOKEN` to the workspace owner's token.
D.Attach a serving credential via a Unity Catalog service principal and grant that principal access to the index and table.
AnswerD

Agent endpoints run under a service identity rather than the deploying user's credentials. By configuring a serving credential backed by a Unity Catalog service principal and granting that principal USE CATALOG, SELECT on the table, and access to the Vector Search index, the endpoint can authenticate to those resources at runtime. This is the supported mechanism for agent endpoints to reach governed data securely.

Why this answer

Mosaic AI Agent endpoints execute under a dedicated service identity, not the deploying user's identity, so Unity Catalog resources must explicitly grant that identity access. Configuring a serving credential tied to a service principal and granting it the necessary privileges on the Vector Search index and Delta table lets the agent authenticate at runtime. Token-based workarounds either expire or over-privilege, and Unity Catalog enforcement cannot be disabled.

Exam trap

The trap here is assuming the endpoint inherits the deploying user's permissions, when agent endpoints actually run under a separate service identity.

212
Multi-Selectmedium

A team is evaluating their RAG application using Mosaic AI Model Evaluation. Which TWO metrics are most relevant for assessing the quality of the generated responses?

Select 2 answers
A.Faithfulness.
B.CPU utilization percentage.
C.Answer relevance.
D.Network latency in milliseconds.
E.Total storage cost of the index.
AnswersA, C

Faithfulness measures the degree to which the generated answer is derived exclusively from the retrieved context. This is the primary metric for detecting hallucinations, ensuring that the model does not invent information, which is a critical requirement for enterprise applications that prioritize truthfulness and reliability in their generative outputs.

Why this answer

Model evaluation requires quantitative metrics to measure both the accuracy of retrieved context and the quality of the final output. Faithfulness ensures the model sticks to the provided context, while answer relevance checks if the model actually addresses the user's intent. Using these metrics allows developers to iterate on their prompt engineering and retrieval strategy, ensuring the application remains accurate and useful for end-users while minimizing the risk of hallucinations.

Exam trap

Candidates often select traditional ML metrics like accuracy or F1-score, forgetting that LLM-based RAG evaluation requires specialized generative metrics like faithfulness and answer relevance.

213
MCQeasy

A GenAI engineer notices that a RAG agent's retrieval stage is returning relevant chunks, but the generated answers frequently omit key facts present in those chunks. The team wants a single evaluation metric that isolates whether the generator is using the provided context. Which metric should they focus on?

A.Answer correctness against a reference answer.
B.Groundedness, which checks whether claims in the answer are supported by the retrieved context.
C.p95 latency of the serving endpoint.
D.Context recall of the retriever.
AnswerB

Groundedness evaluates each claim in the generated answer against the retrieved context. When retrieval is good but the answer omits supported facts or invents unsupported ones, groundedness drops. It isolates the generator's use of context, which is exactly the failure described.

Why this answer

Groundedness is the metric that inspects whether the answer's claims are supported by the retrieved context. Since the retriever is already delivering relevant chunks, a low groundedness score points squarely at the generator ignoring or misusing that evidence, giving the team a precise diagnostic for the stage that is failing.

Exam trap

The trap here is choosing answer correctness because the final output is wrong, when the diagnostic question is which pipeline stage is failing and only groundedness isolates the generator's use of context.

214
MCQhard

A GenAI team's Mosaic AI Agent application is deployed via a Databricks Asset Bundle and served through a Model Serving endpoint. They need the endpoint to call a Unity Catalog function as a tool during inference, and the function reads from a table the endpoint's service principal cannot currently access. What should the engineer do?

A.Grant the endpoint's service principal the required privileges on the table and on the Unity Catalog function, then redeploy the bundle.
B.Embed the service principal's personal access token in the agent code so the function can authenticate at runtime.
C.Add the function's source table to the endpoint's `auto_capture_config` so the endpoint can read it during inference.
D.Change the endpoint's `workload_type` to `GPU_LARGE` so the endpoint runs with elevated permissions.
AnswerA

When a Model Serving endpoint invokes a Unity Catalog function as a tool, the call executes under the endpoint's identity. The service principal therefore needs EXECUTE on the function and the relevant SELECT or USE privileges on the underlying table. Granting those privileges and redeploying the bundle ensures the endpoint can resolve and execute the function at inference time, which is the correct fix for the access failure.

Why this answer

Unity Catalog functions invoked as tools by a Model Serving endpoint execute under the endpoint's identity, which is a service principal. For the function to read its underlying table, that principal needs EXECUTE on the function plus the appropriate privileges on the table and its parent catalog and schema. Granting those privileges and redeploying the bundle resolves the access failure while keeping credentials out of the code.

Exam trap

The trap here is assuming that serving configuration fields like `workload_type` or `auto_capture_config` influence permissions, when access is governed by Unity Catalog grants to the endpoint's service principal.

215
MCQeasy

What is the primary function of a Storage Credential in Unity Catalog?

A.To store user passwords for Databricks.
B.To authorize Databricks to access specific cloud storage.
C.To define the schema of a table.
D.To cache table data for faster query performance.
AnswerB

A storage credential provides the necessary authentication (like a service principal or IAM role) for the Databricks platform to communicate with cloud storage. It acts as the secure identity that Unity Catalog uses to perform read and write operations on behalf of the authorized users.

Why this answer

A Storage Credential serves as a secure container for the cloud-provider credentials required to access external storage locations. By separating the credential from the data objects, Unity Catalog allows administrators to manage access to cloud storage in a centralized, governed way without exposing secrets to individual users. This is a fundamental security practice that prevents credential leakage.

Exam trap

Candidates often think Storage Credentials are used to manage table permissions directly, rather than acting as the underlying authentication mechanism for external storage locations.

216
MCQhard

A GenAI engineer is building a Databricks RAG application where the retrieval step returns the top-5 chunks for each user question. The engineer wants to add a second LLM call that evaluates whether each retrieved chunk contains enough information to answer the question, and then filters out chunks that fail this evaluation before passing the remaining chunks to the final answer-generation prompt. Which design pattern is the engineer implementing?

A.Chain-of-thought prompting
B.Retrieval-augmented generation with relevance grading
C.Fine-tuning the embedding model on domain data
D.Hybrid search combining vector similarity with BM25 keyword matching
AnswerB

The engineer is adding an LLM-based relevance grader between retrieval and generation. Each retrieved chunk is scored for whether it can actually answer the question, and only chunks that pass the grade are forwarded. This pattern improves answer precision by removing semantically similar but non-answering passages before the final prompt is assembled, which is exactly what the scenario describes.

Why this answer

The defining behavior in the scenario is a second LLM call that grades each retrieved chunk for answerability and drops the ones that fail. That is relevance grading layered on top of retrieval-augmented generation, which raises precision by keeping only context that can actually support the answer. The other techniques alter retrieval ranking, prompt reasoning, or embeddings but do not insert an evaluative filtering step between retrieval and generation.

Exam trap

The trap here is assuming that any technique which improves retrieval quality, such as hybrid search or embedding fine-tuning, also performs per-chunk relevance filtering with an LLM.

217
MCQhard

A GenAI application uses a Databricks Model Serving endpoint hosting a foundation model, and the team wants structured JSON output that conforms to a fixed schema for downstream parsing. Responses sometimes include prose or markdown fences that break the parser. Which design change is most likely to produce reliably parseable output?

A.Add a regex-based post-processor that strips markdown fences and extracts the first balanced JSON object from the raw response.
B.Lower the temperature to zero and add the phrase 'return only JSON' to the end of the user message.
C.Define the expected structure with a response format or structured-output parameter on the serving request so the endpoint constrains generation to the schema.
D.Increase max_tokens so the model has room to finish the JSON object without being truncated mid-structure.
AnswerC

Constrained decoding against a declared response schema forces the model's output to match the expected structure, eliminating prose and markdown fences at generation time rather than relying on post-processing. This is the most reliable design change because it enforces conformance at the source and keeps downstream parsing deterministic for the fixed schema.

Why this answer

Declaring a response format or structured-output schema makes the endpoint constrain decoding so the generated text matches the required structure, removing prose and markdown fences by construction. Lowering temperature, raising the token limit, or stripping fences afterward only reduce or mask the symptom without guaranteeing conformance.

Exam trap

The trap here is treating prompt wording such as 'return only JSON' as equivalent to schema enforcement, when only a declared response format constrains decoding.

218
MCQmedium

When monitoring a production LLM, you detect a drift where the model's responses are becoming increasingly verbose and less helpful compared to the baseline. Which strategy is most effective for detecting this quality decay?

A.Monitor system-level CPU and memory usage of the inference cluster.
B.Set up alerts for high request volume spikes in the API logs.
C.Implement automated LLM-as-a-judge evaluations on a sample of production inputs.
D.Require manual reviews for every single response generated by the model.
AnswerC

Using an LLM as a judge allows for automated, scalable evaluation of qualitative metrics. By comparing production outputs against predefined criteria or reference answers, you can effectively track quality drift over time and identify when model responses deteriorate.

Why this answer

Model quality decay in LLMs is best detected through continuous automated evaluation using a 'judge' model. By running a set of evaluation prompts through the production model and grading them against an LLM-as-a-judge, you can quantify performance trends. This proactive monitoring allows teams to identify when model behavior deviates from expected standards, triggering alerts before the issues impact a large number of end-users or require significant architectural changes.

219
MCQmedium

A GenAI engineer is building a retrieval-augmented generation application on Databricks. They want to store document embeddings and perform fast approximate nearest-neighbor search without managing a separate vector database. They have already created a source Delta table with columns: id (string), text (string), and embedding (array<float>). Which Databricks feature should they use to create a Vector Search index that automatically syncs with the Delta table?

A.Databricks Feature Store with a training set
B.MLflow Model Registry with a custom PyFunc model
C.Databricks Vector Search with a Delta Sync index
D.Delta Live Tables with a materialized view
AnswerC

Databricks Vector Search natively supports Delta Sync indexes that automatically keep embeddings in sync with a source Delta table. You can create an index using the embedding column and specify the source table; Databricks manages the underlying vector database and synchronization. This is the intended serverless vector search capability for RAG applications on Databricks.

Why this answer

Databricks Vector Search is the managed service designed for storing embeddings and performing fast similarity search. A Delta Sync index automatically syncs with a source Delta table, so when new documents or embeddings are added, the index updates without manual intervention. This eliminates the need to manage a separate vector database and integrates natively with Databricks workflows.

Exam trap

The trap here is assuming that any Databricks storage feature like Delta tables or Feature Store can serve as a vector index, but only Vector Search provides the required similarity search and automatic synchronization.

220
MCQhard

Refer to the exhibit. An AI engineer is configuring a Databricks Asset Bundle (DAB) to deploy a generative AI application. When executing 'databricks bundle deploy --target prod', which workspace host will the bundle resources be deployed to, and why?

A.The deployment will fail because the default target conflicts with the production target definition during workspace authentication initialization.
B.It will deploy to the development workspace because the dev target is explicitly marked with 'default: true'.
C.It will deploy to the production workspace at https://adb-987654321.azuredatabricks.net because the target flag explicitly selects the prod configuration block.
D.It will deploy to both workspaces simultaneously to maintain synchronization between development and production environments.
AnswerC

The '--target prod' CLI argument instructs the bundle deployment engine to evaluate the corresponding YAML section under targets. Consequently, the workspace host URL defined in the prod block is utilized for provisioning all associated jobs, pipelines, and endpoints.

Why this answer

Databricks Asset Bundles allow developers to define multiple deployment targets within the databricks.yml configuration file. When the user explicitly passes the --target prod flag to the CLI command, the bundle execution context switches from the default target to the specified target configuration, overriding default settings and directing deployment to the production workspace host URL defined under the prod block.

Exam trap

Candidates assume Databricks Asset Bundles always default to production configurations or local workspaces, ignoring the explicit role of the command-line target flag.

221
Multi-Selectmedium

Which TWO evaluation approaches are most effective for measuring the quality of a RAG pipeline's retrieval stage?

Select 2 answers
A.Context Precision
B.Model Answer Faithfulness
C.Context Recall
D.Answer Semantic Similarity
E.LLM Latency Monitoring
AnswersA, C

Context Precision measures the proportion of retrieved chunks that are genuinely relevant to the query, isolating the retrieval stage's ranking quality independently of generation. This satisfies the stem's requirement to evaluate retrieval specifically, revealing whether irrelevant context is being surfaced before the LLM processes it.

Why this answer

Evaluating the retrieval stage requires assessing how well the system identifies relevant documents from the vector store before the LLM generates a response. Context Precision and Context Recall provide quantitative measures to ensure the retrieval engine is performing correctly. Without these metrics, it is impossible to determine if the LLM is failing due to poor generation or simply because it lacked the correct information to answer the query.

222
MCQmedium

An engineer needs to ensure that only authorized users can access the RAG chatbot. Which Databricks security feature should be used to enforce this access control?

A.Notebook tags.
B.Cluster environment variables.
C.Unity Catalog access controls.
D.The public internet firewall.
AnswerC

Unity Catalog provides centralized access control that is essential for securing all data assets within Databricks. By defining clear permissions, engineers can ensure that only authenticated and authorized users have access to the data powering the chatbot, which is a fundamental requirement for any secure, enterprise-ready application.

Why this answer

Unity Catalog's access control lists (ACLs) provide a robust framework for managing permissions at various levels, including catalog, schema, table, and function. By integrating identity management with these controls, engineers can define granular access to the data and the endpoints behind the chatbot. This ensures that sensitive information is only accessible by verified users, maintaining compliance and enterprise-level data security throughout the system.

Exam trap

Candidates often choose workspace-level permissions or generic IAM roles. However, Unity Catalog provides the specific, granular object-level access controls required to secure data assets directly within the RAG pipeline architecture.

223
MCQeasy

A data scientist is developing a GenAI application that uses a foundation model served via Databricks Model Serving. The model endpoint is configured to log inference tables for monitoring. The data science team wants to ensure that the inference logs, which may contain sensitive user prompts, are protected according to Unity Catalog policies. Which Unity Catalog object should be used to store and govern the inference tables?

A.A managed table in a Unity Catalog schema with appropriate grants and column masks.
B.A volume in Unity Catalog with file-level access controls.
C.An external table pointing to a cloud storage location with a storage credential.
D.A view that dynamically redacts sensitive columns from the inference logs.
AnswerA

Inference tables logged by Databricks Model Serving are Delta tables that can be stored in Unity Catalog. By storing them as managed tables in a Unity Catalog schema, the team can apply Unity Catalog governance, including grants, row filters, and column masks, to protect sensitive prompt data. This integrates inference logging with the existing security model, ensuring that access to logs is controlled and auditable.

Why this answer

Inference tables logged by Databricks Model Serving are Delta tables that can be stored in Unity Catalog. Using a managed table in a Unity Catalog schema allows the team to apply Unity Catalog governance features such as grants, column masks, and row filters to protect sensitive data. This ensures that the inference logs are subject to the same security policies as other data assets, providing centralized governance and auditability.

Exam trap

The trap here is thinking that inference logs are automatically governed or that a volume is appropriate; inference logs are tabular Delta tables and should be stored as Unity Catalog tables to leverage fine-grained access controls.

224
MCQhard

A healthcare company uses Databricks to build a GenAI chatbot that answers questions from patient records stored in a Delta table. The records contain PHI, and the company must ensure that the chatbot never returns PHI to unauthorized users. The security team wants to enforce policies at the data layer so that even if the LLM is manipulated, it cannot access PHI. Which Unity Catalog feature should be used to dynamically redact PHI columns based on the user's group membership?

A.Row-level security with a filter function that excludes rows containing PHI.
B.Attribute-based access control (ABAC) policies defined on the catalog to deny access to PHI columns.
C.Column masks using a user-defined function that returns redacted values for unauthorized groups.
D.Dynamic view that joins the patient records with a permissions table and filters out PHI columns.
AnswerC

Unity Catalog column masks allow dynamic redaction of column values based on the invoking user's group membership. A mask function can check 'is_account_group_member()' and return a redacted value (e.g., 'REDACTED') for users not in an authorized group, while returning the original value for authorized users. This enforces PHI protection at the data layer, ensuring the chatbot cannot retrieve PHI for unauthorized users even if the LLM is manipulated.

Why this answer

Unity Catalog column masks are designed to dynamically redact column values based on the user's identity or group membership. By applying a mask function to PHI columns, the healthcare company can ensure that unauthorized users see redacted values while authorized users see the actual data. This enforcement occurs at the data layer, preventing PHI leakage even if the LLM is manipulated, and it applies to all queries against the table.

Exam trap

The trap here is confusing row-level security with column-level security; row filters remove entire rows, while column masks selectively redact values within columns, which is necessary when only specific fields contain PHI.

225
Multi-Selecthard

When designing an application that requires fine-tuning a small model (like Llama-3-8B) on Databricks, which THREE factors must be considered to ensure a successful training job?

Select 3 answers
A.Ensuring the training data is sufficiently representative of the target task.
B.Selecting a GPU-accelerated instance type sufficient for the model's memory.
C.Disabling checkpointing to save storage costs.
D.Configuring appropriate learning rates to prevent catastrophic forgetting.
E.Increasing the batch size to the maximum possible for all datasets.
AnswersA, B, D

Data quality and representativeness are the most critical factors for fine-tuning success. If the training data does not cover the breadth of scenarios the model will encounter, the fine-tuned model will fail to generalize. Ensuring the dataset is high-quality and well-curated is the first step toward effective model adaptation.

Why this answer

Training small models on Databricks requires careful consideration of compute, data quality, and hyperparameter tuning. By ensuring sufficient GPU memory, clean data, and properly tuned learning rates, engineers can successfully adapt the model to new domains. These factors are critical to avoid common pitfalls like catastrophic forgetting or overfitting, which can render the fine-tuned model less useful than its base version.

Exam trap

Candidates sometimes forget hyperparameter tuning specifics, ignoring learning rate configuration which directly causes catastrophic forgetting during fine-tuning.

Page 2

Page 3 of 5

Page 4

All pages