Courseiva

Databricks Certified Generative AI Engineer Associate (Databricks-GenAI-Assoc) — Questions 1–75

330 questions total · 5pages · All types, answers revealed

Page 1 of 5

Page 2
1
MCQmedium

A data scientist is building a GenAI application that uses the OpenAI API through Databricks external model endpoints. The application must use credentials stored securely in Databricks and must not expose the API key in code or logs. Which Databricks feature should the data scientist use to store and reference the OpenAI API key?

A.Unity Catalog volumes with restricted permissions
B.Hardcoding the API key in a notebook and then deleting the cell
C.Databricks secrets with a secret scope
D.Environment variables set in the cluster's Spark configuration
AnswerC

Databricks secrets allow storing sensitive strings like API keys in a secret scope, which is backed by an encrypted store. Secrets are referenced using dbutils.secrets.get or in endpoint configurations without exposing the value in notebooks or logs. This is the standard, secure way to manage credentials for external services in Databricks.

Why this answer

Databricks secrets provide a secure, encrypted store for sensitive credentials such as API keys. They can be referenced in code and in external model endpoint configurations without revealing the secret value. This meets the requirement to keep the OpenAI API key secure and out of code and logs.

Exam trap

The trap here is assuming that environment variables or volumes provide sufficient security for secrets, when in fact they lack encryption and fine-grained access control.

2
MCQeasy

A team is designing a GenAI application that must call an external LLM provider from a Databricks notebook and from a Model Serving endpoint. The security team requires that the provider API key never appear in notebook source code, Git history, or the model artifact, and that the same governed credential be reused by both the notebook and the serving endpoint. Which design satisfies these requirements?

A.Store the API key in a Unity Catalog connection that supports external services, and reference that connection from both the notebook and the Model Serving endpoint.
B.Hard-code the API key as a widget default in the notebook and mirror it into the serving endpoint's environment variables so both components share one value.
C.Store the API key in a Databricks secret scope and reference it with dbutils.secrets.get in the notebook, then pass the resolved value into the serving endpoint's environment variables.
D.Place the API key in a cluster-scoped Spark configuration on the interactive cluster, and have the serving endpoint read the same configuration at request time.
AnswerA

A Unity Catalog connection centralizes credential storage and access control, so the API key never appears in code or artifacts. Both notebooks and Model Serving endpoints can reference the same governed connection, giving consistent secret rotation and auditability. This directly meets the requirement that one governed credential be reused without exposing the plaintext anywhere in source or configuration.

Why this answer

Centralizing the external provider credential in a Unity Catalog connection keeps plaintext out of source code, Git history, and artifacts while letting both notebooks and Model Serving endpoints reference the same governed object. Secret scopes with manual injection, widget defaults, and cluster Spark configurations either leak the value or fail to reach the serving runtime.

Exam trap

The trap here is treating any secret scope reference as sufficient, when the requirement also demands that the same governed credential be consumable by a Model Serving endpoint, which a manually resolved secret does not satisfy.

3
MCQmedium

When designing a production RAG application, which technique is most effective for preventing the LLM from hallucinating based on outdated information?

A.Increasing the temperature parameter of the model to maximum.
B.Enforcing a streaming data pipeline to keep the vector index updated.
C.Restricting the LLM to a specific list of keywords for its output.
D.Adding a long system prompt instructing the model not to lie.
AnswerB

Keeping the vector index synchronized with the source data via a streaming pipeline ensures that the context retrieved during RAG is always current. This minimizes the risk of the model using outdated information, which is a major source of hallucinations in production systems that rely on rapidly changing business data.

Why this answer

Implementing a robust data pipeline that refreshes the vector index in real-time or near-real-time ensures that the retrieved context is current. When coupled with source attribution, this allows the system to verify findings against the latest data. This approach is essential for maintaining accuracy, as stale information in the vector database directly leads to hallucinations that can damage user trust in the AI application.

Exam trap

Candidates often assume that simply increasing the LLM's context window or using a more advanced model will solve hallucinations, ignoring that the root cause is the outdated source data itself.

4
MCQhard

An engineer has a Databricks Vector Search index that stores chunk embeddings for a technical manual. Users report that queries containing exact part numbers return irrelevant chunks because the embedding model blurs numeric tokens. The engineer wants retrieval to consider both dense vector similarity and exact keyword matching on the same Delta table without building a separate search system. Which Databricks Vector Search feature should the engineer enable?

A.A larger embedding model dimension
B.Hybrid search on the index
C.A higher number of returned results per query
D.Delta table change data feed on the source table
AnswerB

Hybrid search in Databricks Vector Search combines dense vector similarity with keyword-based matching over the indexed text, so exact tokens such as part numbers contribute to ranking alongside semantic similarity. Enabling it on the existing index lets the engineer improve recall for exact identifiers without standing up a separate keyword search system or duplicating the source data.

Why this answer

Hybrid search augments dense vector similarity with keyword matching, which restores recall for exact tokens like part numbers while preserving semantic matching for natural-language queries. The other options affect index freshness, capacity, or result volume, but none of them introduce the exact-match signal that the scenario requires.

Exam trap

The trap here is treating retrieval quality as a function of embedding size or result count, when the missing capability is a keyword matching signal that only hybrid search provides.

5
MCQeasy

A GenAI engineer is using MLflow to track experiments for a fine-tuned language model. The engineer wants to ensure that model artifacts and parameters are governed by Unity Catalog, so that access can be controlled and audited. Which Unity Catalog object should the engineer use to register the model?

A.A Unity Catalog schema
B.A Unity Catalog registered model
C.A Unity Catalog volume
D.A Unity Catalog table
AnswerB

A Unity Catalog registered model is the object designed for governing machine learning models. It provides a three-level namespace, versioning, and the ability to grant privileges such as `EXECUTE` and `MANAGE`. By registering the model in Unity Catalog, the engineer ensures that access is controlled and audited, meeting the governance requirement.

Why this answer

Unity Catalog registered models are the correct object for governing machine learning models. They provide a three-level namespace and support fine-grained access control, versioning, and auditing. By registering the fine-tuned model as a Unity Catalog model, the engineer can manage access and track usage, satisfying the governance requirement.

Exam trap

The trap here is confusing volumes or tables with the model registration object, but only registered models provide model-specific governance.

6
MCQmedium

Which component of a Databricks Asset Bundle (DAB) allows you to define different configurations (e.g., instance sizes, variables) for development versus production environments?

A.The 'variables' block at the top level
B.The 'targets' block
C.The 'resources' block
D.The 'include' directive
AnswerB

Targets allow you to define environment-specific overrides for a deployment. You can specify a different workspace, job cluster configuration, or variable values for 'dev', 'staging', and 'prod'. This mechanism ensures the bundle remains portable while allowing the necessary infrastructure adjustments required for each distinct environment's performance and cost profile.

Why this answer

The 'targets' block in the 'databricks.yml' file allows developers to define environment-specific configurations. This enables a single bundle to adapt to different workspaces or compute needs. By using targets, you can ensure that development deployments are lightweight and cost-effective, while production deployments are scaled for high availability and performance, all while keeping the underlying application code identical across all deployment stages.

Exam trap

Candidates often look for 'environment' or 'config' blocks. They fail to identify 'targets' as the specific section in the DAB schema used for environment-specific infrastructure overrides.

7
MCQmedium

A generative AI team stores prompt/response logs in a Unity Catalog Delta table named `main.genai.interaction_logs`. Compliance requires that the raw `prompt_text` column be readable only by members of the `ai_compliance` group, while all other users querying the table must see the literal string `REDACTED`. Which Unity Catalog feature should be implemented directly on the table to satisfy this requirement without creating a separate view?

A.Row filter
B.Dynamic view with a CASE expression
C.Column mask
D.Attribute-based access control tag
AnswerC

A column mask is a Unity Catalog function bound to a column that evaluates the invoking user's identity and returns either the real value or a substituted value. Applying it to prompt_text with a CASE on is_account_group_member('ai_compliance') makes compliance users see raw text and everyone else see REDACTED, exactly as required, while keeping a single table.

Why this answer

Column masks are the Unity Catalog mechanism designed to transform individual column values based on the caller's group membership while leaving the table itself queryable by everyone. Binding a mask function to prompt_text that checks is_account_group_member('ai_compliance') gives compliance staff the raw text and returns REDACTED to all other users, meeting the requirement on the existing table.

Exam trap

The trap here is confusing row-level filtering, which removes rows, with column-level masking, which rewrites values for the same rows.

8
Multi-Selectmedium

A RAG assistant built with Mosaic AI Agent Framework is in production. Product managers complain that answers are sometimes plausible but unsupported by the retrieved documents, and separately that the assistant sometimes ignores documents that clearly contain the answer. The team wants automated, scheduled checks in MLflow LLM Evaluation that separate these two failure modes so each can be triaged independently. Which TWO evaluation metrics should they add to the evaluation suite? (Choose two.)

Select 2 answers
A.P95 response time measured across all endpoint requests in the evaluation window.
B.Latency, measured as the wall-clock time from request receipt to final token of the response.
C.Groundedness, which judges whether the claims in the generated response are supported by the retrieved context.
D.Context recall, which measures whether the retrieved context contains the information needed to answer the question.
E.Total token count, computed as prompt tokens plus completion tokens per request.
AnswersC, D

Groundedness isolates the failure mode where the answer sounds plausible but is not backed by the retrieved documents, because the judge compares each claim in the response against the supplied context. High groundedness means the generation stayed faithful to the evidence; low groundedness flags hallucinated or extrapolated content. This directly addresses the first complaint and can be scored automatically in a scheduled MLflow evaluation run.

Why this answer

The two complaints map to distinct pipeline stages. A groundedness or faithfulness metric detects responses that are plausible but unsupported by the retrieved context, catching generation-side hallucination. A context recall metric detects whether retrieval supplied the evidence at all, catching retrieval-side gaps that make the model answer without support.

Scoring both in the same scheduled evaluation run separates the failure modes so each can be fixed at its own stage, while latency and token metrics carry no information about either complaint.

Exam trap

The trap here is reaching for performance metrics such as latency or token counts to explain answer-quality complaints, when only faithfulness and retrieval-coverage metrics can separate hallucination from missing context.

9
MCQmedium

When designing a RAG application for a regulated industry, what is the most important Unity Catalog feature for ensuring data compliance during retrieval?

A.Catalog-level tagging for organization.
B.Row-level security and column masking.
C.Personal Access Tokens (PATs) for service principals.
D.Automatic data lineage tracking.
AnswerB

Row-level security and column masking are essential for ensuring that users only retrieve authorized information from the source Delta tables. When the index is queried, these Unity Catalog policies apply, ensuring that the RAG application respects organizational access policies, which is mandatory for compliance in regulated industries.

Why this answer

Unity Catalog's attribute-based access control (ABAC) or row-level security allows for granular control over which users can access specific pieces of data. In regulated industries, it is critical that the RAG model only retrieves information that the querying user is explicitly authorized to view, preventing sensitive data leakage and ensuring auditability at the catalog level.

Exam trap

Candidates often choose general encryption or workspace-level access. In regulated industries, security must be applied at the data level (rows/columns) to ensure users only see what they are authorized for.

10
MCQmedium

A GenAI engineer has a Mosaic AI Agent application packaged as a Databricks Asset Bundle with a serving endpoint defined in the bundle's resources. A teammate recently updated the agent's prompt template in the source files, and the engineer now needs to push that change to the existing production endpoint without recreating it. Which Databricks CLI command should the engineer run from the bundle root?

A.databricks bundle validate --target prod
B.databricks serving-endpoints update --name agent-endpoint
C.databricks bundle run agent_app --target prod
D.databricks bundle deploy --target prod
AnswerD

The Databricks CLI bundle deploy command reads the bundle's databricks.yml, resolves the specified target, and applies the current state of the resources to the workspace, updating the existing serving endpoint definition in place rather than recreating it. Running it with --target prod selects the production target's workspace host, variables, and permissions so the prompt change is deployed to the correct environment.

Why this answer

Deploying a modified Mosaic AI Agent that is defined as a bundle resource requires applying the bundle state to the target workspace, which the Databricks CLI does with bundle deploy. Selecting the production target ensures the correct workspace, variables, and permissions are used, and the existing serving endpoint is updated in place instead of being recreated.

Exam trap

The trap here is assuming a serving-endpoint-specific CLI command updates bundle-managed endpoints, when only the bundle deploy command reconciles the declarative bundle state.

11
MCQmedium

A data engineering team is using Databricks to prepare data for a RAG application. They want to ensure that document chunks are of consistent size and quality. Which tool should they use within the Databricks notebook environment to achieve this?

A.A standard Python for-loop on a single worker node.
B.Distributed Spark UDFs with LangChain text splitters.
C.Manual data extraction into a local CSV file.
D.Databricks SQL commands to manually truncate strings.
AnswerB

Using Spark UDFs allows developers to distribute text processing across the cluster, enabling high-throughput document chunking. Integrating LangChain within these UDFs provides standardized, industry-proven logic for splitting text, ensuring consistent chunk sizes that improve the quality and relevance of context retrieved during the RAG process for end users.

Why this answer

Databricks notebooks support a wide range of libraries, but for text processing, using Spark-native functions combined with libraries like LangChain allows for scalable and efficient chunking. By utilizing Spark's distributed processing capabilities, developers can process millions of documents in parallel. This is crucial for maintaining data quality in large-scale RAG systems, as it ensures that the context window of the LLM is not exceeded and that retrieval is focused and accurate.

Exam trap

Candidates tend to choose single-node Python libraries like standard LangChain text splitters without realizing that distributed Spark UDFs are required for scalable processing within notebooks.

12
MCQmedium

A GenAI engineer registers a vector search index in Unity Catalog that points to a Delta table containing customer support transcripts. The security team requires that when an end-user queries the index through a Databricks notebook, the underlying table's row filter and column mask policies are enforced. Which Unity Catalog feature should the engineer configure?

A.Apply tags to the Delta table columns and rely on tag-based access control to mask values at query time.
B.Create a separate vector search index for each user group and assign group-level permissions on those indexes.
C.Use a service principal with read access to the Delta table and query the index through that principal.
D.Enable row-level security and column masks directly on the Delta table, and query the vector search index with the user's identity propagated.
AnswerD

Unity Catalog row filters and column masks are enforced at query time against the caller's identity. When a vector search index is queried through a SQL warehouse or notebook session that carries the end-user's identity, the underlying Delta table policies apply, ensuring the security team's masking and filtering requirements are met without duplicating rules on the index itself.

Why this answer

Row filters and column masks in Unity Catalog are applied dynamically based on the identity executing the query. When a vector search index is queried with end-user identity propagation, the source table's policies are evaluated, so sensitive transcript fields are masked and rows are filtered according to group membership. This satisfies the security team's requirement without duplicating policies on the index.

Exam trap

The trap here is assuming that security policies must be redefined on the vector search index itself, when Unity Catalog enforces them on the underlying Delta table at query time.

13
MCQeasy

A team deploys a customer-support assistant on a Mosaic AI Model Serving endpoint and enables inference table logging to Unity Catalog. Compliance requires that every production response be traceable back to the exact request, the retrieved context, and the model version that produced it, and that reviewers can query this history with SQL months later. Which capability satisfies this requirement?

A.The endpoint's request and response payloads written to a Unity Catalog inference table, queried with SQL and joined to model version metadata.
B.Real-time dashboards built on the endpoint's built-in metrics such as request count and error rate.
C.MLflow experiment runs recorded during offline evaluation of the assistant before release.
D.Client-side application logs streamed to a log analytics workspace and retained for thirty days.
AnswerA

Inference tables persist the full request, retrieved context, response, and associated metadata as Delta tables governed by Unity Catalog, so reviewers can run SQL against historical traffic long after the fact. Because the logged rows include the served model version and timestamp, each response can be traced to its exact inputs and the model that generated it, which is precisely the auditability the compliance team requires.

Why this answer

Compliance-grade traceability demands durable, queryable records of each request, the context supplied to the model, the response, and the serving model version. Unity Catalog inference tables store exactly that as governed Delta tables, so SQL queries and joins to model version metadata answer audit questions months after the traffic occurred. Client logs, offline experiment runs, and aggregate dashboards all lack either the payload detail or the retention and governance needed.

Exam trap

The trap here is confusing operational dashboards or client-side logs with auditable records, when only governed payload-level logging preserves request, context, and model version together.

14
MCQhard

A GenAI engineer monitors a customer-facing RAG assistant hosted on Databricks. After a routine re-indexing job, groundedness scores from the MLflow LLM Evaluation job drop sharply while answer relevance stays flat. The application prompt and the LLM serving endpoint were untouched. Which conclusion is best supported by these signals?

A.The LLM judge model was silently upgraded and is scoring more harshly
B.Users are asking harder questions than before, lowering groundedness
C.The re-indexing job degraded retrieval, so the generator receives weaker supporting context
D.The generation prompt has drifted and needs to be rewritten
AnswerC

Groundedness depends on the retrieved context supporting the answer, so a sharp drop right after re-indexing implicates the retrieval pipeline. Flat relevance shows the generator still addresses the question, but now without adequate evidence, producing less grounded answers. Chunking changes, embedding mismatches, or partial index builds during re-indexing are the likely culprits to investigate first.

Why this answer

The drop begins exactly when the index was rebuilt, and relevance remains flat, which isolates the change to the evidence supplied to the generator rather than to the prompt, the LLM, or the judge. Degraded retrieval after re-indexing, from altered chunking, embedding mismatches, or an incomplete build, is the best-supported conclusion and the right place to investigate first.

Exam trap

The trap here is attributing a groundedness drop to the model or judge, when a sharp drop synchronized with re-indexing and stable relevance points to the retrieval context instead.

15
Multi-Selectmedium

A data engineer is preparing a large corpus of support tickets stored in a Unity Catalog volume for fine-tuning a Llama model on Databricks. They must remove personally identifiable information (PII) before the data reaches the training cluster. Which TWO approaches are appropriate for detecting and redacting PII at scale in this pipeline? (Choose two.)

Select 2 answers
A.Apply a Spark NLP or presidio-based pandas UDF that detects entities such as names, emails, and phone numbers and replaces them with placeholders.
B.Run a Databricks job that uses the ai_query() function with a foundation model endpoint to classify and rewrite each ticket, removing PII.
C.Set the table's retention policy to 0 days so that raw tickets are deleted immediately after ingestion.
D.Use Unity Catalog column masks on the raw text column so that anyone querying the table sees redacted values.
E.Enable server-side encryption on the Unity Catalog volume and rely on the storage layer to anonymize the text.
AnswersA, B

A pandas UDF running a PII detection library like Presidio or Spark NLP processes partitions in parallel and can redact entities before data is written to the training table. This keeps the redaction inside the lakehouse, scales with the cluster, and integrates with Unity Catalog governance. It is a common pattern for pre-training data sanitization.

Why this answer

PII must be removed from the content itself before training. Distributed detectors such as Presidio or Spark NLP in pandas UDFs, and model-based rewriting with ai_query(), both transform the text at scale inside Databricks. Column masks, encryption, and retention policies govern access or lifecycle but leave the underlying tokens intact, so they do not satisfy the preprocessing requirement.

Exam trap

The trap here is confusing access controls like column masks or encryption with actual content redaction, which must alter the text before training.

16
MCQhard

An engineer deployed a RAG agent to a Model Serving endpoint and enabled inference tables for payload logging. After a week, reviewers notice that logged requests contain customer email addresses and order identifiers. Compliance requires that raw prompts and responses not be stored in plain text. Which change should the engineer make?

A.Enable payload logging but write inference tables to a workspace-local path outside Unity Catalog.
B.Disable inference tables entirely so no request or response data is persisted.
C.Configure the endpoint to log only the model's token counts and latency, then reconstruct prompts from the application's own logs.
D.Keep inference tables enabled but restrict SELECT on the inference table to a compliance group and apply row filters or column masks on the sensitive columns.
AnswerD

Inference tables are Unity Catalog tables, so column masks and row filters can redact or hide sensitive fields, and GRANT controls who can query them. This keeps observability for authorized reviewers while ensuring raw personal data is not broadly readable, satisfying the compliance requirement without losing monitoring capability.

Why this answer

Because inference tables are Unity Catalog tables, the engineer can apply column masks and row filters to redact or restrict sensitive fields and grant SELECT only to authorized principals. This preserves the observability inference tables provide while preventing plain-text exposure of personal data, which is the precise compliance outcome required.

Exam trap

The trap here is treating inference table logging as all-or-nothing, when Unity Catalog masking and grants let you keep logs while redacting sensitive fields.

17
MCQmedium

An AI engineer is developing a real-time customer service chatbot application using Databricks Model Serving and needs to securely store API keys and database credentials without hardcoding them into the application code or notebook. Which approach should the engineer use?

A.Store the credentials in a plain-text JSON file located in the workspace root directory and read it programmatically at runtime.
B.Hardcode the credentials directly into the scoring script deployed to the Databricks Model Serving endpoint.
C.Utilize Databricks Secrets to create a secret scope and retrieve values securely using dbutils.secrets.get during application execution.
D.Pass the credentials as plain-text arguments in the command-line interface when triggering the job cluster.
AnswerC

Databricks Secrets encrypt credentials at rest and in transit, integrating seamlessly with Unity Catalog and workspace access controls. Using dbutils.secrets ensures that sensitive tokens are masked in logs and only exposed dynamically to authorized execution contexts, safeguarding production environments effectively.

Why this answer

Databricks Secrets provide a secure mechanism for managing credentials and sensitive configuration parameters. By leveraging secret scopes and the dbutils.secrets utility within notebooks or environment variables mapped to serving endpoints, developers ensure that sensitive tokens never leak into version control systems, adhering to strict enterprise security and compliance standards for production application deployments.

Exam trap

Candidates often suggest using environment variables directly or manual file uploads, overlooking that dbutils.secrets is the standard, secure, and platform-native way to retrieve sensitive values within Databricks.

18
MCQhard

An engineer built an agent using Mosaic AI Agent Framework and wants the agent to call a Unity Catalog function that returns customer order history. The function must be invoked by the LLM at runtime without exposing raw SQL to the model. Which approach should the engineer use?

A.Expose the function through a SQL warehouse and have the agent send natural-language queries to the warehouse endpoint.
B.Embed the SQL body of the function in the system prompt so the LLM can generate the query when needed.
C.Register the Unity Catalog function as a tool in the agent and let the LLM select it by name and arguments.
D.Create a Databricks job that runs the function on a schedule and writes results to a Delta table the agent reads.
AnswerC

Mosaic AI Agent Framework supports Unity Catalog functions as tools that the LLM can invoke by name with structured arguments. The function's signature and docstring become the tool schema, so the model never sees or writes SQL directly. This is the supported pattern for giving agents governed access to governed data while keeping the interface declarative.

Why this answer

Registering the Unity Catalog function as a tool gives the agent a typed, governed interface the LLM can invoke by name with structured arguments. The function's metadata supplies the schema, so no SQL is exposed to or generated by the model. Prompt-embedded SQL, scheduled jobs, and text-to-SQL all fail the requirement of runtime, governed invocation without model-authored queries.

Exam trap

The trap here is believing the LLM must see or generate SQL to query Unity Catalog data, when the agent framework exposes functions as typed tools instead.

19
MCQeasy

When deploying a Python-based application that interacts with Unity Catalog, which step is essential to ensure the code can authenticate securely to external services without hardcoding tokens?

A.Hardcoding tokens as environment variables in the notebook.
B.Using Databricks Secret Scopes to retrieve credentials.
C.Storing tokens in a plain text file inside the bundle.
D.Creating a public access policy for the external service.
AnswerB

Secret scopes are the secure way to store and access sensitive information in Databricks. By using the 'dbutils.secrets.get' function, the code retrieves credentials at runtime from a secure vault. This ensures that application logic remains decoupled from specific security credentials, facilitating safer code promotion across different environments.

Why this answer

Using secret scopes within Databricks is the recommended best practice for handling authentication credentials. Secrets allow code to retrieve sensitive information like API keys or database passwords at runtime, keeping these credentials out of the source code. This is a critical security requirement in any production application, as it prevents accidental exposure of sensitive keys in version control systems and allows for centralized management of authentication lifecycle.

Exam trap

Test-takers often select environment variables or configuration files, ignoring the Databricks security best practice of utilizing Secret Scopes to handle credentials safely.

20
MCQmedium

A generative AI engineer is building a RAG chain on Databricks using LangChain. The chain must call an external LLM provider via a Databricks Model Serving endpoint that proxies the provider, and the engineer wants to avoid hardcoding credentials in notebooks. Which approach should the engineer use to authenticate calls from the chain to the serving endpoint?

A.Pass the provider API key as a plaintext argument in the LangChain ChatOpenAI constructor and rely on notebook access controls to protect it.
B.Use a personal access token generated for a service principal and embed it in the chain's prompt template so the LLM can forward it to the provider.
C.Configure the serving endpoint with the provider credential stored as a secret in the endpoint's environment, and call the endpoint using the Databricks SDK or LangChain's Databricks LLM class with workspace authentication.
D.Store the provider API key in a Databricks secret scope and reference it with dbutils.secrets.get inside the notebook before instantiating the LLM client.
AnswerC

Model Serving endpoints can hold provider credentials as secrets in their environment, so the client only needs a Databricks token. LangChain's Databricks LLM integration and the Databricks SDK both use workspace authentication, keeping the provider key out of notebooks and enabling centralized governance and rotation.

Why this answer

The serving endpoint is the correct place to hold the external provider credential because it centralizes secret storage, rotation, and auditing. Clients then authenticate to the endpoint with workspace identity rather than the provider key. LangChain's Databricks LLM integration and the Databricks SDK both support this pattern, keeping credentials out of notebooks and enabling Unity Catalog governance over the endpoint.

Exam trap

The trap here is assuming the provider API key must be retrieved in the notebook, when the serving endpoint is designed to store and use that credential on the client's behalf.

21
MCQmedium

A media company's RAG assistant answers questions about its streaming catalog. Users report that the assistant often returns answers that ignore the retrieved documents and instead rely on outdated information from the base model. The team wants an MLflow LLM Evaluation metric that measures how much of the final answer is actually derived from the retrieved context rather than the model's prior knowledge. Which metric best fits this need?

A.relevance
B.groundedness
C.context_sufficiency
D.answer_correctness
AnswerB

Groundedness measures the degree to which the response's statements are supported by the retrieved context. When the assistant ignores retrieved documents and falls back on outdated knowledge, its claims will not be traceable to the context, producing low groundedness scores that expose exactly the behavior the team wants to quantify.

Why this answer

Groundedness directly assesses whether the response is supported by the retrieved context, so a pattern of ignoring documents and using stale base-model knowledge will show up as low groundedness. That makes it the metric that quantifies the behavior the team is investigating in their catalog assistant.

Exam trap

The trap here is assuming that answer correctness or relevance reveals whether the model used retrieved context, when only groundedness measures the link between response and context.

22
MCQhard

You are building a data preparation pipeline for a generative AI application. The raw data is stored in a Unity Catalog volume as JSON files with nested fields. You need to flatten the nested structure and extract specific fields into a Delta table for downstream embedding. The JSON schema may evolve, with new fields added occasionally. Which approach provides the most robust and maintainable solution?

A.Use the explode function on the nested arrays and then select the required fields.
B.Use the from_json function with a predefined schema and select the required fields.
C.Use the schema_of_json function to infer the schema at runtime and then apply from_json.
D.Use the variant data type to store the JSON and then use variant_get to extract fields.
AnswerD

The variant data type in Databricks can store semi-structured JSON without a predefined schema, and variant_get allows extracting fields by path. This handles schema evolution gracefully because new fields are automatically included in the variant. It is a robust, maintainable solution for evolving JSON, and it integrates with Delta tables for downstream processing.

Why this answer

The variant data type stores JSON without a rigid schema, and variant_get extracts fields dynamically, accommodating schema evolution. This avoids brittle predefined schemas and manual updates. Other options either require static schemas, rely on sampling, or only handle arrays, making them less robust for evolving nested JSON.

Exam trap

The trap here is assuming that inferring a schema at runtime solves schema evolution, when only a flexible type like variant truly accommodates new fields without pipeline changes.

23
MCQhard

A company is deploying a GenAI chatbot that uses a Databricks Model Serving endpoint hosting a fine-tuned Llama 2 model. The endpoint is registered in Unity Catalog. The security team wants to restrict which groups can invoke the endpoint and also log all inference requests for auditing. Which combination of Unity Catalog and Databricks features should the engineer use to meet these requirements?

A.Grant EXECUTE on the model version to the allowed groups and enable inference tables on the serving endpoint
B.Grant USE CATALOG and USE SCHEMA on the model's parent schema and enable verbose logging on the endpoint
C.Grant SELECT on the model version to the allowed groups and enable audit logs in the workspace admin console
D.Use a personal access token (PAT) with scoped permissions and configure a Delta table as a sink for endpoint logs
AnswerA

Unity Catalog privileges on model versions include EXECUTE, which controls who can use the model for inference. Granting EXECUTE to specific groups restricts invocation. Inference tables automatically capture request and response payloads for the serving endpoint, providing an audit trail. Together, they satisfy both access control and logging requirements without custom code.

Why this answer

To control who can invoke a model served by Databricks, you grant the EXECUTE privilege on the model version in Unity Catalog. Inference tables are a built-in feature of Model Serving that automatically log request and response data to a Delta table for auditing and monitoring. This combination directly addresses the security team's requirements for access restriction and request logging.

Exam trap

The trap here is confusing SELECT privilege on a model with EXECUTE, or assuming that workspace audit logs capture inference payloads.

24
MCQhard

A media company uses an LLM-as-a-judge evaluation pipeline on Databricks to score a summarization assistant nightly. The judge is the same model family as the assistant and was given a rubric that rewards stylistic fluency. Over two months, nightly scores climb steadily while editor spot-checks find summaries increasingly omit key facts. Which corrective action best restores the evaluation's ability to detect this regression?

A.Revise the judge rubric to weight factual coverage against the source document, add a small human-labeled calibration set, and periodically verify judge agreement with editors.
B.Increase the judge's temperature so its scores vary more between runs and regressions become statistically visible.
C.Run the judge more frequently, moving from nightly to hourly scoring, so trends are captured with finer granularity.
D.Switch the judge to the largest available frontier model and remove the rubric so the judge can apply its own judgment freely.
AnswerA

The judge is self-preferring a fluent sibling model and rewarding style over substance, so the rubric must explicitly score factual coverage and be validated against human labels. A calibration set and periodic agreement checks detect when the judge drifts away from editorial standards, which is the only way the nightly signal can be trusted again.

Why this answer

The rising scores reflect a judge that rewards fluency because it shares a model family with the assistant and was given a style-oriented rubric, so the instrument itself is blind to factual omission. Restoring detection requires redefining the rubric around factual coverage of the source, anchoring it with human-labeled examples, and monitoring judge-human agreement so drift in the judge is caught before it masks another regression.

Exam trap

The trap here is treating rising judge scores as improving quality when the judge and the assistant share a family and the rubric rewards fluency, so the metric drifts upward while factual omissions grow.

25
MCQhard

A team uses an LLM judge to score their RAG agent on Databricks and sees scores that fluctuate by several points between identical evaluation runs. They need more stable, reproducible quality signals for release gating. Which action best addresses the root cause?

A.Increase the number of evaluation examples until the average stabilizes across runs.
B.Average the scores from three different judge models and use the mean for gating.
C.Switch from an LLM judge to BLEU or ROUGE string-overlap metrics.
D.Pin the judge model version and set its temperature to zero, then validate the judge against a human-labeled sample.
AnswerD

Judge variance often comes from sampling temperature and model version changes. Setting temperature to zero and pinning the judge version makes scoring deterministic and reproducible, while human-labeled validation confirms the judge's scores are trustworthy before they gate releases.

Why this answer

Fluctuating judge scores usually stem from nondeterministic sampling and drifting judge model versions. Pinning the judge version and setting temperature to zero makes each scoring pass reproducible, and validating the judge against human labels confirms that the stabilized scores are also accurate enough to gate releases reliably.

Exam trap

The trap here is chasing aggregate stability by enlarging the dataset or blending judges, when the per-example variance originates from judge sampling settings and version drift.

26
MCQhard

A Generative AI engineer is preparing a Delta table of support tickets for embedding generation. The pipeline computes embeddings with ai_query and stores them in a separate embeddings table keyed by ticket_id. Tickets are frequently updated by agents, and the engineer wants the embeddings table to reflect the latest ticket text without recomputing embeddings for tickets whose text has not changed. Which design best achieves this?

A.Add a hash column of the ticket text and use a QUALIFY clause to filter unchanged rows within the same query.
B.Use change data capture with APPLY CHANGES to merge ticket updates into the embeddings table, recomputing embeddings only for changed rows.
C.Enable Auto Optimize on the embeddings table so only modified files are rewritten during updates.
D.Run a full refresh of the embeddings table on every pipeline update so all ticket embeddings stay current.
AnswerB

APPLY CHANGES processes the change feed and applies inserts, updates, and deletes to the embeddings table. By deriving embeddings from the changed rows only, unchanged tickets keep their existing vectors and are not re-embedded. This matches the requirement to reflect updates while avoiding redundant embedding computation.

Why this answer

To update embeddings only for changed tickets, the pipeline must consume a change feed and apply changes to the embeddings table. APPLY CHANGES merges inserts, updates, and deletes while allowing embedding derivation to run only on changed rows. Full refreshes recompute everything, and file-level or query-level filters cannot persist decisions across updates.

Exam trap

The trap here is assuming a hash or file-optimization feature can skip unchanged rows across runs, when only a change data capture flow persists that decision.

27
MCQmedium

A GenAI engineer is building a customer-support assistant on Databricks. The assistant must invoke an external LLM endpoint that requires an API token, and the token must not appear in notebook source or logs. The engineer wants to store the token once and let notebooks and Databricks Jobs reference it by name. Which Databricks capability should the engineer use to satisfy this requirement?

A.A Unity Catalog volume mounted at /Volumes/secure/token.txt
B.A workspace environment variable set in the cluster Spark configuration
C.A personal access token stored in the Databricks CLI configuration file
D.Databricks secrets backed by a secret scope
AnswerD

Databricks secrets store sensitive values in a secret scope and expose them only through dbutils.secrets.get or secret references, so the token never appears in notebook source, logs, or job definitions. This satisfies the requirement of storing the token once and referencing it by name from notebooks and jobs while keeping it out of version control.

Why this answer

Databricks secrets provide a governed, auditable store for sensitive values and expose them through dbutils.secrets.get or secret references, so the external LLM token is never embedded in code or logs. Other storage locations either expose plaintext to readers or serve a different purpose, such as authenticating to the workspace rather than to an external provider.

Exam trap

The trap here is assuming any file or variable accessible from a notebook is safe for credentials, when only a purpose-built secret store redacts values and enforces access control and auditing.

28
MCQmedium

A GenAI engineer is preparing a large text corpus stored in a Unity Catalog volume for fine-tuning a chat model. The raw files are JSON Lines, each containing a 'conversation' array with alternating 'user' and 'assistant' turns, but many records have malformed turns or missing roles. The engineer needs to convert this into a Delta table with a schema of (conversation_id STRING, messages ARRAY<STRUCT<role:STRING, content:STRING>>) while filtering out records where any turn has a null role or empty content. Which approach uses the appropriate Databricks-native capability for this transformation?

A.Load the files as a text RDD, apply a regular expression to extract role and content pairs, and then convert the RDD to a DataFrame with the desired schema using createDataFrame.
B.Use the from_json function with a defined StructType schema to parse each JSON line into the conversation structure, then apply a filter using a higher-order function such as filter on the messages array to retain only records where all turns have non-null role and non-empty content.
C.Use Databricks Auto Loader with schema evolution to ingest the files into a Bronze table, then use Delta Live Tables expectations to drop rows where the role column is null or the content column is empty.
D.Use the spark.read.json method with inferSchema enabled to automatically detect the schema, then use a UDF written in Python to iterate over each turn and remove records with null roles or empty content.
AnswerB

from_json with an explicit StructType accurately parses nested JSON into the required ARRAY<STRUCT<role:STRING, content:STRING>> schema. Applying a higher-order filter (e.g., array filtering) or an exists/forall expression on the parsed array lets you discard records with malformed turns. This leverages Spark SQL's native JSON and array functions, which are designed for scalable, schema-enforced transformations on large corpora, exactly matching the need to enforce structure and filter invalid records.

Why this answer

The correct approach uses from_json with an explicit schema to parse the nested JSON into the required array of structs, then applies a higher-order function to filter out records with malformed turns. This leverages Spark SQL's native, optimized JSON parsing and array operations, which scale well and enforce the target schema. It avoids the performance pitfalls of UDFs, the fragility of regex, and the indirectness of ingestion frameworks not designed for immediate nested transformation.

Exam trap

The trap here is assuming that schema inference or UDF-based cleanup is necessary when native from_json and higher-order functions can both parse and validate the nested structure in a single, optimized pass.

29
MCQhard

Refer to the exhibit. A developer is registering a model. Why is the model signature, as shown in the exhibit, considered a best practice for model registration?

A.It enables automatic data masking for sensitive strings.
B.It provides automated input validation to improve service reliability.
C.It automatically scales the number of serving nodes based on the prompt size.
D.It forces the model to use a specific version of Python for execution.
AnswerB

By defining the input schema, the model signature allows the serving infrastructure to validate every incoming request against the expected format. If the request is malformed, the system rejects it, preventing potential failures inside the model logic and ensuring that the serving endpoint maintains its stability under various load conditions.

Why this answer

Model signatures define the schema for inputs and outputs, allowing the serving layer to perform automated validation. When an incoming request doesn't match the signature, the serving infrastructure can reject it immediately, providing useful error messages. This prevents invalid data from reaching the model, reducing runtime errors and debugging time, and ensures that the application behaves predictably in production environments, which is critical for maintaining high service quality and reliability.

Exam trap

Candidates often assume model signatures are primarily used for logging metrics or tracking lineage, overlooking their critical runtime role in automated input validation.

30
MCQhard

In a RAG application, which architectural component ensures that the system handles changes in the source data effectively?

A.The LLM's internal knowledge base.
B.The Delta Lake Change Data Feed (CDF).
C.The prompt engineering template.
D.The user session cache.
AnswerB

CDF provides a structured way to track changes in Delta tables. By leveraging CDF, the vector search index can incrementally ingest only the delta between the old and new states of the data. This is significantly more efficient than full re-indexing and ensures high data consistency with minimal compute cost.

Why this answer

The Delta Lake Change Data Feed (CDF) or the automatic synchronization of Vector Search indexes is designed specifically to track and apply changes (inserts, updates, deletes) to downstream systems. This ensures that the vector representation of the data is always in sync with the source of truth, avoiding the need for expensive full re-indexes and ensuring that retrieved information is always up-to-date.

Exam trap

Candidates often select manual scheduling tools or general caching mechanisms instead of recognizing Delta Lake Change Data Feed as the native engine tracking incremental table changes.

31
MCQhard

A platform team is serving a retrieval-augmented generation application. The vector index is built from a Delta table in Unity Catalog, and the team wants every similarity search executed by an end user to be automatically restricted to documents that user is allowed to see, without maintaining a separate index per department. Which approach best enforces this at query time?

A.Apply a column mask to the embedding column so unauthorized users receive null vectors
B.Build one vector index per department and route queries based on the caller's group
C.Grant SELECT only to a service principal and have the application connect as that principal for all users
D.Create a row filter function on the source Delta table that filters rows by the invoking user's group membership
AnswerD

Row filters on the underlying Delta table are evaluated against the invoking principal, so a retrieval query that reads through the table returns only rows that caller may see. Because the vector index is refreshed from that table, embedding the filter there enforces per-user visibility automatically without duplicating indexes per department, which matches the requirement precisely.

Why this answer

Row filters evaluate the invoking user's identity as part of the query, so placing the filter on the source Delta table means every read, including retrieval reads that feed similarity search, is automatically scoped to allowed rows. This keeps a single governed index and avoids per-department duplication while still enforcing per-user document visibility at query time.

Exam trap

The trap here is assuming vector search bypasses Unity Catalog authorization, when reads through governed tables still honor row filters.

32
MCQmedium

A team is deploying a fine-tuned LLM using Mosaic AI Model Serving. To reduce the cost of serving the model while maintaining acceptable performance, which strategy should be prioritized?

A.Use the largest available instance type to ensure zero request failures.
B.Implement model quantization to reduce memory footprint and hardware requirements.
C.Switch to a Multi-Model Serving endpoint regardless of throughput.
D.Disable all logging to save on storage costs.
AnswerB

Quantization reduces the precision of the model weights, which leads to smaller memory usage and faster inference times. This allows the model to run on smaller, cheaper instances without a significant drop in accuracy, providing a highly effective way to balance performance with operational costs in production environments.

Why this answer

Using smaller, optimized model architectures or quantization techniques significantly reduces the compute and memory footprint required for serving. By shrinking the model size, you can effectively use smaller instance types for serving, leading to direct cost savings. Furthermore, optimizing the underlying serving configuration by choosing the right hardware type ensures you are not paying for expensive GPU capacity when CPU-based serving would suffice for your throughput needs.

Exam trap

Candidates incorrectly suggest increasing GPU instance sizes to handle LLM costs, overlooking model optimization techniques that reduce hardware requirements altogether.

33
Multi-Selectmedium

Which TWO factors should be considered when choosing an embedding model for a RAG application?

Select 2 answers
A.The model's native dimensionality and its effect on vector index storage size.
B.The model's ability to generate creative, hallucinatory content.
C.The compatibility of the model with the target languages of the documents.
D.The maximum number of characters allowed in a single prompt.
E.The color profile of the documents being ingested.
AnswersA, C

Higher dimensionality embeddings provide better semantic resolution but significantly increase memory usage and storage costs. Choosing the right dimensionality is a trade-off between the precision of the similarity search and the infrastructure overhead required to maintain the index, which directly impacts the scalability and cost of the RAG application.

Why this answer

The choice of embedding model determines the quality of semantic retrieval and the computational requirements of the system. By aligning the model's language support, dimensionality, and performance with the application's needs, you ensure a balanced design. It is also vital to consider the model's licensing and whether it needs to be fine-tuned to capture domain-specific terminology that general models might misinterpret.

Exam trap

Candidates often focus solely on model accuracy. They neglect the practical constraints of vector index storage size and language compatibility, which are critical for production scalability and performance.

34
MCQhard

A Generative AI engineer is using MLflow Tracing to monitor a RAG application deployed on Mosaic AI Model Serving. They want to capture the retrieved documents, the final prompt, and the model's response for each request to debug a quality issue. Which approach should they use to ensure all three are logged in a single trace?

A.Instrument the application with the MLflow tracing SDK, creating spans for retrieval, prompt construction, and generation within the same trace context.
B.Use the `mlflow.evaluate()` API with a custom evaluator that logs retrieved documents, prompt, and response as metrics.
C.Enable inference table logging on the serving endpoint and query the payload column for the retrieved documents and response.
D.Configure the serving endpoint to log all requests to a Delta table and then join with a separate table of retrieved documents using a request ID.
AnswerA

MLflow Tracing allows manual instrumentation with spans that share a trace context. By wrapping retrieval, prompt construction, and generation in spans, all three are captured in one trace. This provides end-to-end visibility, which is essential for debugging RAG quality issues. It works with Mosaic AI Model Serving and integrates with the MLflow UI for inspection.

Why this answer

To capture retrieved documents, the final prompt, and the model response in a single trace, the engineer should use MLflow Tracing with manual instrumentation. Spans for each step share a trace context, providing a unified view. This is the native Databricks approach for debugging RAG applications and integrates with Mosaic AI Model Serving for production monitoring.

Exam trap

The trap here is assuming that inference table logging automatically includes intermediate retrieval steps and the exact prompt, when it only captures request and response payloads.

35
MCQmedium

A developer is deploying a custom LangChain agent application as a custom Python model to Databricks Model Serving. The application depends on a specific third-party library that is not included in the standard Databricks runtime environment. How should the developer ensure the dependency is installed when the model is loaded into the serving container?

A.Manually SSH into the serving container nodes after deployment and run pip install using terminal access.
B.Include the external library installation command inside a standard print statement within the model scoring function.
C.Specify the required packages in the extra_pip_requirements argument when logging the MLflow model artifact.
D.Rely on the serving endpoint to automatically discover and install any imported Python module dynamically from the public internet.
AnswerC

extra_pip_requirements records the third-party library in the MLflow model's dependency metadata, so Model Serving installs it into the container at load time. This satisfies the constraint that the package is absent from the standard Databricks runtime.

Why this answer

When logging custom PyFunc models to MLflow, developers must explicitly specify Python dependencies using the extra_pip_requirements parameter or a conda.yaml file. This ensures that the Databricks Model Serving environment automatically provisions and installs all necessary external libraries during container builds, preventing runtime import errors and ensuring reliable model execution.

Exam trap

Candidates often suggest manually installing packages on the cluster or using init scripts, failing to use the MLflow-native 'extra_pip_requirements' parameter that ensures environment reproducibility.

36
Multi-Selecthard

A financial services company has registered a Mosaic AI Agent application in Unity Catalog and wants to serve it through a Databricks Model Serving endpoint that queries a Delta table containing sensitive customer records. The security team requires that the endpoint access the table using a dedicated service principal with least privilege, and that the agent's LLM calls go through a governed gateway that logs usage. Which TWO configurations should the team apply to meet these requirements? (Choose two.)

Select 2 answers
A.Store the table's credentials in the agent's environment variables and have the agent open a direct JDBC connection to the Delta storage account.
B.Embed the customer records directly into the agent's MLflow model artifact so the endpoint never queries Unity Catalog at runtime.
C.Disable Unity Catalog enforcement on the endpoint so the agent can read the table with the workspace owner's inherited permissions.
D.Grant the serving endpoint's service principal SELECT on the specific customer table and its parent schema, and no broader catalog privileges.
E.Configure a Databricks AI Gateway on the serving endpoint so that requests to the foundation model are routed, authenticated, and logged centrally.
AnswersD, E

Model Serving endpoints execute under a service principal, and Unity Catalog enforces that principal's grants at query time. Granting SELECT only on the specific table and its parent schema satisfies least privilege while still allowing the agent's retriever to read the records it needs. Broader catalog-wide grants would violate the security team's requirement and expose unrelated tables to the endpoint.

Why this answer

The endpoint runs as its own service principal, so Unity Catalog grants on the specific table and schema give the agent exactly the read access it needs without broader exposure. Databricks AI Gateway on the serving endpoint provides governed routing, authentication, and logging of foundation-model calls. Together these two controls satisfy the least-privilege and governed-gateway requirements while keeping the agent's data access auditable.

Exam trap

The trap here is thinking that embedding data in the model artifact or opening a direct connection avoids governance problems, when in fact both bypass Unity Catalog and break the audit and least-privilege requirements.

37
MCQhard

A GenAI engineer is developing an agent using Databricks Mosaic AI Agent Framework. The agent must call an external API to fetch real-time stock prices. The engineer wants to ensure the agent can be evaluated and deployed with minimal changes between development and production. Which approach should the engineer take to define the tool-calling logic?

A.Create a separate Databricks job that periodically fetches stock prices and writes them to a Delta table, then have the agent query that table.
B.Implement the API call as a Python function and register it as a tool in the agent's LangChain or similar framework, then log the agent with MLflow.
C.Use a Databricks SQL function to call the external API via a UDF, and invoke it from the agent.
D.Hardcode the API call directly in the agent's prompt as a description of the endpoint.
AnswerB

Defining tools as Python functions and registering them within the agent framework allows the logic to be logged as part of the MLflow model. This ensures the same code runs in evaluation and deployment, meeting the requirement of minimal changes between environments.

Why this answer

The Mosaic AI Agent Framework supports defining tools as Python functions that the agent can invoke. By registering the function as a tool and logging the agent with MLflow, the same code is used for evaluation and deployment, ensuring consistency. Other approaches either do not execute the call, introduce delay, or add architectural complexity that hinders portability.

Exam trap

The trap here is thinking that describing an API in the prompt enables the LLM to call it; in reality, tool execution requires code that the agent framework can invoke.

38
MCQmedium

Which approach is most effective for managing the dependencies of a custom ML model when deploying it to Mosaic AI Model Serving?

A.Install all dependencies manually on the serving cluster after deployment.
B.Include a 'requirements.txt' or 'conda.yaml' file in the model artifact.
C.Assume the serving environment has all standard ML libraries pre-installed.
D.Package the entire Python environment inside the model artifact as a zip.
AnswerB

Including dependency files in the model artifact ensures the model carries its environment definition with it. Mosaic AI Model Serving detects these files and builds the necessary environment for the model, ensuring that the production serving environment is identical to the one in which the model was validated.

Why this answer

Defining a custom environment using Conda or a 'requirements.txt' file ensures that the exact library versions required by the model are captured and recreated in the serving environment. This eliminates 'dependency hell' where a model works in development but fails in production due to library version mismatches. Ensuring consistent environments is a foundational step in robust ML engineering that guarantees reproducible performance in the serving infrastructure.

Exam trap

Candidates often assume that standard notebooks automatically bundle local Python environments into deployed endpoints, forgetting that Mosaic AI Model Serving requires explicit dependency files like requirements.txt or conda.yaml to recreate the correct libraries.

39
MCQeasy

A data science team has registered a new version of a retrieval-augmented generation (RAG) agent in Unity Catalog as a model named main.genai.support_agent. They need to make this version available as a low-latency REST endpoint with automatic scaling and no cluster management. Which Databricks capability should they use?

A.Databricks SQL warehouses, by registering the agent as a SQL function and querying it from a dashboard.
B.All-purpose interactive clusters, by starting a cluster and calling the model through a notebook.
C.Jobs compute, by scheduling the agent as a nightly batch task and writing predictions to a Delta table.
D.Mosaic AI Model Serving, by creating a serving endpoint that serves the registered Unity Catalog model.
AnswerD

Mosaic AI Model Serving hosts registered Unity Catalog models and agent models behind a REST endpoint with managed, automatic scaling and no cluster management, which is exactly what the team needs for low-latency inference. Creating a serving endpoint that references the registered model version deploys it directly, preserving governance and lineage from Unity Catalog.

Why this answer

Deploying a registered Unity Catalog agent model as a Mosaic AI Model Serving endpoint gives the team a managed REST inference API with automatic scaling and no cluster administration. It preserves Unity Catalog governance and lineage, supports low-latency online serving, and is the intended deployment path for GenAI agents on Databricks. Batch, interactive, and SQL compute cannot provide the same managed serving contract.

Exam trap

The trap here is assuming any Databricks compute (clusters, SQL warehouses, or jobs) can serve a production REST inference endpoint, when only Mosaic AI Model Serving provides that managed contract.

40
MCQmedium

A team has deployed a RAG chatbot on Databricks and enabled inference table logging. They want to set up automated monitoring to detect when the average response length increases significantly compared to the baseline. Which Databricks feature should they use to create this monitor?

A.Mosaic AI Model Serving endpoint logs analyzed with the `mlflow.evaluate()` API on a schedule.
B.Databricks SQL alerts on a scheduled query that calculates the average response length from the inference table.
C.Lakehouse Monitoring for the inference table, with a custom metric for response length.
D.MLflow Model Registry webhooks to trigger a retraining pipeline when response length changes.
AnswerC

Lakehouse Monitoring can track statistical properties of Delta tables, including inference tables. By defining a custom metric for response length, the team can monitor drift and receive alerts when the average deviates from the baseline. This is the native Databricks solution for monitoring data quality and model performance over time. It integrates with the inference table without additional ETL.

Why this answer

Lakehouse Monitoring is designed to monitor Delta tables, including inference tables, for data quality and drift. By creating a monitor with a custom metric for response length, the team can automatically detect significant changes from the baseline and receive alerts. This is the most direct and integrated approach on Databricks for this monitoring requirement.

Exam trap

The trap here is confusing batch evaluation tools like `mlflow.evaluate()` with continuous monitoring features like Lakehouse Monitoring, which are purpose-built for tracking drift in production tables.

41
MCQmedium

An engineer is designing a GenAI application where a Databricks workflow must call an external LLM provider that enforces a strict rate limit and occasionally returns transient errors. The engineer wants the application to degrade gracefully instead of failing the whole job. Which design choice best addresses this requirement?

A.Increase the driver node type on the job cluster
B.Cache all prompts and responses in a Delta table before calling the provider
C.Set the job to retry the entire task on failure
D.Wrap the external call in retry logic with exponential backoff and a fallback response path
AnswerD

Retries with exponential backoff absorb transient rate-limit and server errors without immediately failing the job, and a fallback path lets the application return a degraded but useful response when retries are exhausted. Together they prevent a single external hiccup from aborting the entire workflow.

Why this answer

Resilience to external rate limits and transient failures comes from handling the call itself: bounded retries with exponential backoff smooth over temporary rejections, and a fallback response keeps the application functional when the provider stays unavailable. Infrastructure sizing, task-level retries, and response caching do not address the failure at its source.

Exam trap

The trap here is reaching for cluster or job-level scaling to solve an external dependency problem, when the failure originates at the API boundary and must be handled in the calling code.

42
MCQmedium

An AI engineer is building a RAG application and wants to log the retrieval context and generated response for each request to MLflow for evaluation. They are using the `mlflow.langchain` flavor. Which method should they use to log the model so that MLflow automatically captures the necessary artifacts?

A.mlflow.sklearn.log_model
B.mlflow.tensorflow.log_model
C.mlflow.langchain.log_model
D.mlflow.pyfunc.log_model
AnswerC

`mlflow.langchain.log_model` is designed specifically for LangChain models. It automatically logs the chain's structure, prompts, and other artifacts, enabling MLflow to capture inputs and outputs for evaluation. This method simplifies logging and ensures compatibility with MLflow's evaluation tools, making it the correct choice.

Why this answer

For LangChain models, `mlflow.langchain.log_model` is the specialized logging function that captures the chain's configuration and enables MLflow to track inputs and outputs. This allows the engineer to later evaluate retrieval context and responses using MLflow's evaluation capabilities. Other logging methods do not provide this automatic integration.

Exam trap

The trap here is assuming that any logging function works, but only the flavor-specific logger for LangChain automatically captures the necessary artifacts for RAG evaluation.

43
MCQmedium

A data engineer wants to track who accessed a specific table in Unity Catalog. Which tool should they use?

A.Databricks cluster logs
B.Unity Catalog system tables
C.Cloud provider logs
D.Workspace activity console
AnswerB

Unity Catalog system tables (like 'system.access.audit') store comprehensive logs of all data access events. These tables are the standardized way to query historical audit information across the entire metastore, enabling engineers to perform detailed analysis of who accessed specific tables and when.

Why this answer

The System Tables (specifically the access_logs table) in Unity Catalog provide detailed audit logs of all interactions with data objects. By querying these tables using SQL, administrators can monitor data access patterns, identify unauthorized attempts to view sensitive data, and fulfill audit requirements. This centralized logging is a cornerstone of the governance transparency provided by Unity Catalog.

Exam trap

Candidates commonly suggest standard Spark logs or cluster event logs instead of utilizing the specialized audit capabilities found in Unity Catalog system tables.

44
MCQmedium

A developer is configuring a model serving endpoint as shown in the exhibit. They observe that the endpoint fails to respond quickly to the first request after a period of inactivity. What is the cause of this behavior?

A.The model version is incompatible with the CPU workload type.
B.The endpoint is entering a cold-start phase due to scale-to-zero.
C.The model name is not registered in the Unity Catalog.
D.The workload type 'CPU' is incorrectly specified as a string.
AnswerB

Setting 'scale_to_zero_enabled' to true triggers the termination of compute resources during idle periods to minimize costs. The subsequent latency is caused by the time required to provision new compute and load the model back into memory, which is the intended mechanism for this serverless configuration setting.

Why this answer

The 'scale_to_zero_enabled' flag is set to true, which instructs the Databricks infrastructure to shut down the compute resources entirely when no traffic is detected to save costs. When a new request arrives, the infrastructure must perform a 'cold start,' initializing the containers and loading the model into memory. This introduces latency, which is expected behavior when optimizing for infrastructure cost over instant availability in a serverless environment.

Exam trap

Candidates often mistake this for a network latency issue or an API rate limit. The 'scale-to-zero' configuration is a cost-saving feature that inherently introduces a cold-start delay.

45
MCQeasy

A team has developed a GenAI chatbot using the Databricks Mosaic AI Agent Framework and logged it with MLflow. They now want to expose it as a REST API that their customer support portal can call, with autoscaling and built-in monitoring. Which Databricks capability should they use to host the agent?

A.A Databricks SQL warehouse that exposes the agent through a SQL function the portal calls.
B.An all-purpose interactive cluster that the portal connects to over JDBC to invoke the agent.
C.Databricks Model Serving, deploying the MLflow-logged agent to a serving endpoint.
D.A Databricks job that runs the agent on a schedule and writes responses to a Delta table the portal polls.
AnswerC

Model Serving is the Databricks capability that turns an MLflow-logged model or agent into a REST endpoint with autoscaling, concurrency controls, and integrated monitoring. Deploying the agent there gives the support portal a stable URL and handles capacity automatically. It is the intended path for serving Mosaic AI Agent Framework applications in production.

Why this answer

Mosaic AI Agent Framework applications logged with MLflow are served on Databricks Model Serving, which provides a REST endpoint, autoscaling, concurrency management, and built-in monitoring. That makes it the correct host for a customer support portal that needs to call the chatbot over HTTP. The other options are batch, development, or analytics compute surfaces that cannot serve interactive inference.

Exam trap

The trap here is assuming any Databricks compute can host the agent, when only Model Serving provides the REST endpoint, autoscaling, and serving-grade monitoring required for interactive traffic.

46
Multi-Selectmedium

A team is developing a generative AI application that uses an LLM to answer questions based on internal documents. They want to ensure the application is robust and provides accurate responses. Which TWO practices should they implement to improve the reliability of the application? (Choose two.)

Select 2 answers
A.Implement retrieval-augmented generation (RAG) to ground the LLM's responses in the internal documents.
B.Fine-tune the LLM on the internal documents to embed knowledge directly into the model weights.
C.Implement guardrails to validate and filter the LLM's outputs before returning them to the user.
D.Use a larger LLM with more parameters to increase the likelihood of correct answers.
E.Set the LLM's temperature to a high value to encourage more creative and diverse responses.
AnswersA, C

RAG retrieves relevant documents and includes them in the prompt, so the LLM generates answers based on factual, up-to-date internal content. This reduces hallucinations and improves accuracy. It is a core practice for building reliable GenAI applications that need to answer domain-specific questions.

Why this answer

The two practices are RAG and guardrails. RAG grounds responses in internal documents, reducing hallucinations and providing accurate, up-to-date information. Guardrails validate outputs to filter out harmful or incorrect content.

Together, they significantly improve the reliability and safety of a GenAI application, especially for enterprise use cases.

Exam trap

The trap here is assuming that fine-tuning or using a larger model is the best way to improve reliability, but without retrieval and output validation, these approaches can still produce hallucinations and lack source attribution.

47
MCQmedium

A team is deploying a Databricks RAG chatbot that must serve interactive traffic with low latency while also allowing the data science team to test prompt variations safely. The engineer wants the production endpoint to keep serving stable traffic even while a new prompt template is being evaluated. Which approach best meets these requirements?

A.Lower the production endpoint's concurrency limit and run the experiment on the same endpoint during off-peak hours
B.Edit the prompt template in the production endpoint's registered model version and restart the endpoint
C.Add the new prompt as a branch inside the production chain and select it at runtime using a random number generator
D.Create a separate Model Serving endpoint for the experimental prompt and route a small percentage of traffic to it for comparison
AnswerD

Isolating the experimental prompt on its own serving endpoint lets the team evaluate it with real or synthetic traffic without risking the production endpoint's stability. Traffic can be split at the application layer for A/B comparison, and if the experiment degrades, production remains unaffected because the two endpoints scale and version independently.

Why this answer

Separating the experimental prompt onto its own Model Serving endpoint preserves production stability while enabling controlled comparison. The team can direct a small share of traffic to the experimental endpoint, measure quality and latency, and promote the prompt only after it proves better, all without redeploying or interrupting the production endpoint.

Exam trap

The trap here is assuming a single endpoint can host both stable and experimental prompts safely, when shared capacity and shared code paths let an experiment degrade production traffic.

48
MCQmedium

When monitoring a RAG application, you notice a high discrepancy between the retrieved context and the generated answer. Which metric would specifically help identify if the model is ignoring the provided context?

A.Context Precision
B.Faithfulness
C.Retrieval Recall
D.Semantic Similarity
AnswerB

Faithfulness specifically assesses whether the answer is logically derived from the provided context. High faithfulness indicates the model is respecting the context; low faithfulness suggests the model is generating responses based on its own training data, which leads to hallucinations and incorrect information in RAG systems.

Why this answer

Faithfulness measures whether the generated answer is derived exclusively from the retrieved context. If a model generates information not supported by the source text, it is 'hallucinating' or ignoring the context. Monitoring faithfulness is crucial for RAG systems because it directly detects when the model drifts away from the ground truth provided by the internal knowledge base, which is the primary value proposition of a RAG architecture.

Exam trap

Candidates often confuse faithfulness with context relevance, failing to realize that faithfulness specifically evaluates whether the answer is derived directly from the retrieved context.

49
Multi-Selecthard

A Databricks team is building a retrieval corpus from mixed-format documents stored in a Unity Catalog volume. They need a preparation pipeline that preserves document structure for later chunking and that records which source file each chunk came from so retrieval results can cite evidence. Which two design choices best meet these requirements? (Choose two.)

Select 2 answers
A.Concatenate all documents into a single large text field before chunking to reduce the number of rows.
B.Carry the source file path as a metadata column on each chunk and store it in the Delta table alongside the chunk text and embedding.
C.Convert every document to plain text and discard page and heading markers to normalize the corpus.
D.Generate embeddings for entire documents rather than chunks to avoid storing multiple rows per file.
E.Use the ai_parse_document function to extract text and layout elements from PDFs and images, retaining the document structure for downstream chunking.
AnswersB, E

Attaching the source file path as a metadata column lets retrieval results reference the originating document, enabling citations and traceability. Delta tables support arbitrary metadata columns, so the path can travel with each chunk through embedding generation and into the vector index without additional joins.

Why this answer

Preserving document structure requires a parser that extracts layout elements rather than flattening everything to plain text, and ai_parse_document provides that for PDFs and images. Provenance requires carrying the source file path on each chunk so retrieval results can cite the originating document. Together these choices keep structure intact and make every chunk traceable.

Exam trap

The trap here is treating normalization to plain text as harmless, when discarding layout markers actually removes the structure needed for coherent chunking and provenance.

50
Multi-Selecthard

You are preparing a large text corpus for fine-tuning a generative AI model. The corpus is stored in a Delta table with columns: doc_id, raw_text, and metadata. You need to create a cleaned dataset that removes personally identifiable information (PII) and normalizes whitespace, while preserving document boundaries for training. Which two actions should you perform to achieve this in a scalable and maintainable way? (Choose two.)

Select 2 answers
A.Convert the Delta table to Parquet files and manually inspect each file for PII.
B.Use the ai_classify function to label each document as 'clean' or 'dirty' and drop dirty ones.
C.Use the ai_analyze_sentiment function to filter out negative documents before training.
D.Apply a regular expression to replace all whitespace sequences with a single space in the raw_text column.
E.Use the ai_mask function to redact PII entities from the raw_text column.
AnswersD, E

Normalizing whitespace with a regular expression such as regexp_replace(raw_text, '\\s+', ' ') collapses multiple spaces, tabs, and newlines into single spaces. This standardizes the text for tokenization and reduces noise without altering document boundaries. It is a scalable and deterministic step that can be applied in a SELECT or withColumn transformation.

Why this answer

ai_mask redacts PII in place, and regexp_replace normalizes whitespace, together achieving the cleaning goals at scale. Both are declarative, repeatable, and preserve document boundaries. Sentiment filtering, manual inspection, and classification are either irrelevant or destructive to the dataset, and they do not address PII removal and whitespace normalization.

Exam trap

The trap here is thinking that any AI function can clean text, when only ai_mask targets PII and regex handles whitespace; other AI functions like sentiment or classification do not perform the required transformations.

51
Multi-Selectmedium

Which THREE strategies improve the quality of retrieval in a Databricks Vector Search-based RAG application?

Select 3 answers
A.Implementing hybrid search using both semantic vectors and keyword-based filtering.
B.Increasing the chunk size to include the entire dataset in a single vector.
C.Adding relevant metadata tags to documents to enable targeted filtering.
D.Optimizing the chunking strategy to maintain context boundaries.
E.Removing all stop words from the documents during the ingestion phase.
AnswersA, C, D

Hybrid search combines the strengths of semantic understanding with the precision of keyword matching. This ensures that specific technical terms or IDs are found accurately while still capturing the intent behind user queries, significantly improving retrieval quality for complex domain-specific datasets where semantic similarity alone might lead to irrelevant results.

Why this answer

Improving retrieval quality requires a combination of high-quality data chunking, metadata filtering, and optimized search parameters. These strategies directly influence the relevance and accuracy of retrieved context, which is the foundation of a successful RAG application. Mastering these techniques allows engineers to fine-tune system behavior to meet specific user needs and reduce hallucinations during the generation phase.

Exam trap

Candidates often assume that only vector embeddings matter, ignoring how hybrid search and metadata filtering are essential for combining semantic relevance with precise keyword matches to boost retrieval quality.

52
MCQhard

An organization wants to implement 'Guardrails' on their model outputs. Which deployment strategy best facilitates this?

A.Hardcode guardrails inside the base model training script.
B.Deploy a wrapper model that processes the output for safety.
C.Instruct users to manually check the output for safety.
D.Disable all logging to prevent guardrails from slowing down.
AnswerB

A wrapper or chain approach allows the model to generate text, which is then passed through an inspection layer for validation. If the output violates safety guidelines, it can be blocked or replaced. This modularity is the standard approach for applying consistent guardrails without impacting core model performance.

Why this answer

The best way to implement guardrails is to build a secondary validation model or use an API-based gatekeeper that intercepts and inspects model outputs before they are returned to the user. Integrating this logic into the inference pipeline within the Databricks environment ensures that all content is screened. This pattern is essential for enterprise deployments where safety, compliance, and preventing hallucinations are top priorities for generative AI applications.

Exam trap

Test-takers often think prompt engineering alone is a sufficient guardrail, failing to realize enterprise apps require programmatic output filtering.

53
Multi-Selectmedium

A fintech company runs a customer-support RAG assistant on Databricks. Before promoting a new prompt template from staging to production, the ML team must demonstrate that the change does not regress answer quality. Which TWO evaluation practices should they apply in Mosaic AI Agent Evaluation to make this promotion decision defensible? (Choose two.)

Select 2 answers
A.Pin the evaluation dataset, judge model version, and retrieval configuration so both templates are scored under identical conditions and results are reproducible.
B.Deploy the candidate template to a small percentage of production traffic and rely on user thumbs feedback alone as the promotion signal.
C.Run the evaluation dataset against both the current production template and the candidate template, then compare per-question scores rather than only aggregate averages.
D.Increase the judge model size until its scores on the candidate template exceed the production template's scores.
E.Evaluate only the questions where the candidate template produced different answers, skipping the rest to save judge tokens.
AnswersA, C

Holding dataset, judge version, and retrieval settings constant ensures any score delta is attributable to the prompt template alone. Without this control, a judge upgrade or index refresh between runs confounds the comparison. Pinning also makes the promotion decision reproducible for auditors and lets the team re-run the gate after future infrastructure changes.

Why this answer

A defensible promotion gate requires a controlled comparison where the prompt template is the only variable, which means running both templates over the same dataset with the judge model and retrieval configuration pinned, and inspecting per-question deltas instead of relying on averages that can mask localized regressions. Together these practices attribute any score movement to the change under test and make the decision reproducible.

Exam trap

The trap here is chasing a passing aggregate score by enlarging or re-tuning the judge, instead of holding evaluation conditions constant and comparing the same questions under both prompt templates.

54
MCQmedium

You are preparing a large corpus of customer support emails stored as Parquet files in a Unity Catalog volume. The emails must be cleaned by removing boilerplate signatures and disclaimers before tokenization for a fine-tuning dataset. Your team wants to enforce this transformation as a declarative, testable pipeline stage that fails fast if any cleaned record still contains a known boilerplate marker. Which Databricks capability should you use to implement this cleaning step with built-in data quality expectations?

A.Use Spark Structured Streaming with a foreachBatch function that calls a Python UDF to strip signatures and raises an exception if a marker remains.
B.Create a Databricks SQL query that applies regexp_replace to the email body and schedule it as a SQL warehouse alert when a marker is detected.
C.Write a notebook that reads the Parquet files, applies a pandas UDF to clean the text, and writes the result back to the volume, then manually review the output.
D.Define the cleaning logic in a Delta Live Tables pipeline and attach an EXPECT constraint that asserts no boilerplate marker remains in the cleaned column.
AnswerD

Delta Live Tables pipelines support declarative expectations that can drop, fail, or quarantine records based on a condition. By attaching an EXPECT constraint to the cleaned table, the pipeline fails fast if any record still contains a boilerplate marker, and the expectation metrics are automatically captured in the event log for auditing and testing.

Why this answer

Delta Live Tables expectations are the declarative, testable mechanism for enforcing data quality during preparation. Attaching an EXPECT constraint to the cleaned table ensures the pipeline fails fast if any boilerplate marker remains, while capturing metrics for auditing. This approach provides lineage, observability, and reproducibility, unlike imperative notebooks or asynchronous alerts.

Exam trap

The trap here is assuming that any code that raises an exception on bad data is equivalent to a declarative data quality expectation, when only Delta Live Tables expectations provide built-in fail-fast semantics and metrics.

55
MCQhard

A team is preparing a pretraining corpus and must filter out near-duplicate documents before tokenization. The corpus contains 500 million short text records in a Delta table. They want a scalable, deterministic deduplication signal that can be computed per record and compared across the dataset. Which approach best fits this requirement?

A.Sort the corpus by document length and drop every record whose length falls within one standard deviation of another record.
B.Compute an MD5 hash of each full document and drop rows sharing the same hash value.
C.Compute a MinHash signature per document with a fixed number of hash permutations and compare signatures with a locality-sensitive hashing banding scheme.
D.Train a sentence embedding model and cluster documents with k-means, then drop all but one document per cluster.
AnswerC

MinHash with LSH banding produces a fixed-length signature per document and a scalable candidate-generation step, so near-duplicates are found without all-pairs comparison. It is deterministic given fixed seeds and permutations, distributes cleanly across Spark partitions, and directly targets near-duplicate detection rather than exact matches, fitting the 500 million record corpus.

Why this answer

Near-duplicate detection at scale needs a compact per-record signature plus a candidate-generation strategy, which MinHash with locality-sensitive hashing banding provides. Exact hashing misses near-duplicates, embedding clustering is costly and nondeterministic, and length-based filtering has no content signal, so none of those satisfy the deterministic, scalable requirement.

Exam trap

The trap here is equating deduplication with exact-match hashing, when near-duplicate removal requires a similarity-preserving signature rather than a cryptographic digest.

56
Multi-Selecthard

A company is scaling its RAG applications. Which THREE of the following are benefits of using Databricks Vector Search over managing a standalone vector database outside of the platform?

Select 3 answers
A.Automatic, continuous synchronization with Delta tables.
B.Ability to use SQL as the only way to perform embedding calculations.
C.Inheritance of Unity Catalog security and governance policies.
D.Reduction in operational overhead by using managed infrastructure.
E.Support for non-Databricks proprietary vector file formats.
AnswersA, C, D

Native integration with Delta tables allows the vector index to update automatically whenever the source table changes. This removes the need for manual ETL or complex cron jobs, drastically reducing the maintenance overhead and ensuring that the retrieved information is always fresh for the RAG application.

Why this answer

Databricks Vector Search offers benefits through tight integration, including automatic synchronization with Delta tables, which removes the need for custom ETL. Because it resides in the same security boundary, it inherits Unity Catalog's governance and access controls, ensuring data security. Finally, it uses the platform's managed infrastructure, eliminating the overhead of maintaining external database clusters, which allows developers to focus on application logic rather than infrastructure maintenance.

Exam trap

Test-takers sometimes pick features related to external model training or manual ETL pipelines, missing the native governance and synchronization benefits of Databricks Vector Search.

57
MCQmedium

A GenAI engineer has built a Retrieval Augmented Generation (RAG) application using Databricks Vector Search. The index is created on a Delta table that contains sensitive customer support transcripts. Company policy requires that end users can only retrieve chunks from documents they are authorized to view. The engineer wants to enforce this at query time without duplicating the index. Which approach should the engineer use?

A.Create a separate Vector Search index for each user group and route queries based on the user's group membership.
B.Use Databricks Vector Search with the 'filter' parameter in the query, based on user attributes passed from the application.
C.Apply Unity Catalog column masks on the embedding column of the source Delta table to redact sensitive text before vectorization.
D.Enable Attribute-Based Access Control (ABAC) on the source Delta table using row filter and column mask functions.
AnswerB

Databricks Vector Search supports query-time filtering on metadata columns. The engineer can include user authorization attributes as metadata during index creation and then pass a filter expression at query time (e.g., 'allowed_groups' containing the user's group). This enforces document-level security without duplicating the index, directly satisfying the policy requirement while keeping a single index for all users.

Why this answer

Databricks Vector Search enables document-level security by allowing metadata filtering at query time. The engineer can include authorization attributes (such as group memberships) as metadata fields when creating the index, then pass a filter expression in the query that matches the user's attributes. This enforces access control without duplicating the index, aligning with the requirement to keep a single index while restricting retrieval to authorized documents.

Exam trap

The trap here is assuming that Unity Catalog row filters and column masks on the source table automatically apply to Vector Search index queries, when in fact the index is a separate serving layer that requires its own query-time filtering.

58
MCQeasy

Which Databricks component is recommended for orchestrating the end-to-end RAG pipeline, including data ingestion, transformation, and vector indexing?

A.Databricks SQL.
B.Databricks Workflows.
C.The Databricks File System (DBFS).
D.Unity Catalog.
AnswerB

Databricks Workflows provides the orchestration layer needed to schedule and manage tasks like data ingestion, transformation, and vector index updates. Its ability to manage dependencies between tasks and provide monitoring is critical for keeping the RAG pipeline functional, up-to-date, and reliable in a production environment.

Why this answer

Databricks Workflows allows for the scheduled and triggered execution of tasks, such as notebook runs or Delta Live Tables pipelines. This is the recommended way to orchestrate complex RAG pipelines where data must be cleaned, chunked, and indexed systematically. By using Workflows, engineers ensure that the entire pipeline is reproducible, monitorable, and resilient to failures, which is essential for maintaining a production-grade AI application.

Exam trap

Candidates tend to confuse cluster management tools with orchestration tools, selecting cluster settings instead of Databricks Workflows for running multi-step ETL and RAG pipelines.

59
Multi-Selecthard

A team is designing a Databricks GenAI application that must return grounded answers with citations to source documents. The application uses Databricks Vector Search for retrieval and a Foundation Model API for generation. Which TWO design choices are required to return accurate citations alongside each answer? (Choose two.)

Select 2 answers
A.Instruct the model in the prompt to reference the provided chunks by their identifiers and include those identifiers in the response.
B.Set the model's temperature to zero to guarantee that citations are factually correct.
C.Increase the embedding model's dimensionality to the maximum supported value so retrieval is more precise.
D.Enable Foundation Model API pay-per-token billing so the model has access to citation formatting.
E.Store source metadata such as document ID, title, and chunk position alongside each embedding in the Vector Search index.
AnswersA, E

Even with metadata available, the model must be told to attribute its statements to specific chunks. Prompting the model to cite chunk identifiers, and structuring the prompt so each chunk carries a visible ID, gives the model the information it needs to ground its answer and lets the application resolve those IDs into human-readable citations.

Why this answer

Citations require two things working together: source metadata attached to each retrieved chunk so the application knows where content came from, and prompt instructions that tell the model to attribute statements to specific chunk identifiers. Precision tuning, billing mode, and temperature settings do not supply the provenance information or the attribution behavior needed to render citations.

Exam trap

The trap here is assuming that better retrieval quality automatically yields citations, when provenance metadata and explicit attribution instructions are what actually enable them.

60
MCQmedium

Which of the following is a primary benefit of using Unity Catalog to manage models for a RAG application?

A.It automatically generates the optimal prompt for the model.
B.It provides centralized lineage and access control.
C.It eliminates the need for vector embeddings.
D.It forces all data to be stored in the cloud root.
AnswerB

Unity Catalog enables organizations to track data lineage from the source to the model, which is critical for compliance and debugging. Furthermore, it applies consistent access controls across workspaces, ensuring that sensitive data used in RAG applications is protected by the same security policies as the rest of the enterprise.

Why this answer

Unity Catalog provides a unified governance framework that manages permissions, lineage, and discovery across all data and AI assets. In a RAG context, this ensures that the entire pipeline—from the source Delta tables to the vector indexes and the final models—is traceable and governed. This consistency is essential for enterprise compliance and simplifies the operational management of the AI lifecycle by providing a centralized point for auditing and security policy application.

Exam trap

Candidates often choose 'Model Registry' or 'MLflow' without Unity Catalog. While those manage model artifacts, Unity Catalog specifically provides the enterprise-grade governance, lineage, and permissioning required for RAG pipelines.

61
MCQhard

A team deploys a RAG agent to a Mosaic AI Model Serving endpoint. During load testing, requests intermittently return HTTP 429 responses even though the endpoint shows healthy replicas. The agent calls an external vector search index and a foundation model endpoint. Which action most directly addresses the 429 responses?

A.Increase the endpoint's provisioned concurrency and configure appropriate rate limits so the agent can handle the incoming request rate.
B.Move the vector search index into the same Unity Catalog schema as the agent model to reduce cross-service latency.
C.Add retry logic with exponential backoff and jitter to the client calling the endpoint.
D.Switch the agent's foundation model calls to a smaller context window to reduce token usage per request.
AnswerA

HTTP 429 indicates the endpoint is rate-limiting requests because the incoming rate exceeds provisioned concurrency or configured limits. Raising provisioned concurrency and tuning rate limits lets the agent absorb the load, which directly addresses the 429s rather than masking them. This aligns capacity with the observed traffic during load testing.

Why this answer

A 429 from a Mosaic AI Model Serving endpoint signals that inbound request rate exceeds the endpoint's provisioned capacity or configured rate limits. Raising provisioned concurrency and setting realistic rate limits aligns capacity with the load-test traffic and directly removes the rate-limiting responses. Client retries, smaller prompts, and index relocation do not increase the endpoint's ability to accept concurrent requests, so they leave the root cause unaddressed.

Exam trap

The trap here is reaching for client-side retries as the fix for 429s, when a sustained 429 pattern points to insufficient endpoint capacity rather than transient failure.

62
MCQeasy

A data engineer is preparing a dataset of product descriptions for embedding generation. The source table in Unity Catalog contains a column 'description' with mixed languages, and the team wants to filter to English-only text before vectorization. They need a scalable, built-in Databricks function that can detect the language of each description without external API calls. Which function should they use?

A.regexp_extract
B.ai_detect_language
C.ai_classify
D.ai_analyze_sentiment
AnswerB

ai_detect_language is a built-in Databricks SQL function that returns the detected language code for a given text column. It is designed for scalable language identification directly in SQL or PySpark, requiring no external API calls. Filtering on the returned language code allows the team to keep only English descriptions before embedding, exactly matching the requirement.

Why this answer

The built-in ai_detect_language function is purpose-built for language identification at scale within Databricks. It returns a language code that can be used to filter the dataset to English-only records before embedding generation. Other AI functions like ai_analyze_sentiment or ai_classify serve different purposes, and regex cannot reliably determine language.

Exam trap

The trap here is confusing general-purpose AI functions such as ai_classify with a dedicated language detection function, when only ai_detect_language directly returns a language code for filtering.

63
MCQmedium

An organization needs to build a RAG application on Databricks that minimizes data egress and maximizes security by keeping all data within the workspace perimeter. Which architectural pattern best satisfies this requirement?

A.Call external LLM APIs from a standard Python notebook without VPC constraints.
B.Export data to an external vector database and use a cloud-hosted LLM.
C.Deploy an embedding model on Mosaic AI Model Serving and utilize Databricks Vector Search.
D.Use a public LLM endpoint with a public bucket to store the vector index.
AnswerC

This approach keeps all data processing, storage, and inference within the Databricks workspace perimeter. By using internal serving for embeddings and native vector search capabilities, the architecture minimizes egress traffic and simplifies security policy enforcement, ensuring that data never leaves the protected environment during the RAG retrieval process.

Why this answer

Utilizing Mosaic AI Model Serving with private endpoints and leveraging Vector Search indexes ensures that both the embedding model and the retrieval process occur within the Databricks control plane. By avoiding external API calls to third-party providers, the organization maintains strict governance, data residency compliance, and lower latency for inference, which is critical for enterprise-grade generative AI applications handling sensitive corporate documents.

Exam trap

Candidates often suggest external API-based embedding models, ignoring the requirement to keep data within the workspace perimeter to minimize egress and satisfy strict residency requirements.

64
MCQhard

A team runs a RAG application on Mosaic AI Model Serving and logs all requests to an inference table. Reviewers report that answers are sometimes fluent but contradict the retrieved documents. The team wants a recurring, automated check that quantifies this contradiction on production traffic and alerts when it exceeds a threshold. Which approach best fits?

A.Enable inference table payload logging and inspect raw prompt and response text manually each week.
B.Increase the number of retrieved chunks per query so the model has more context to draw from.
C.Schedule a Databricks job that runs mlflow.evaluate() with the groundedness judge over recent inference-table records and writes results to a Delta table for alerting.
D.Monitor the endpoint's p95 latency and error rate in the serving endpoint metrics and alert when either degrades.
AnswerC

A scheduled job that applies the groundedness judge to recent inference-table records turns production traffic into a repeatable quality measurement. Writing scores to Delta enables threshold-based alerts and trend analysis. This directly targets the contradiction symptom because groundedness measures whether the answer is supported by retrieved context.

Why this answer

Groundedness evaluation compares generated answers against retrieved context, which is precisely the failure mode described. Running it as a scheduled job over recent inference-table records produces a quantified, recurring metric, and persisting results to Delta supports alerting on threshold breaches. Operational metrics and manual review cannot detect semantically unsupported but fluent responses.

Exam trap

The trap here is treating operational telemetry such as latency or error rate as a proxy for answer quality, when semantically wrong answers still return successful responses.

65
Multi-Selectmedium

Which TWO factors should be prioritized when selecting an embedding model for a domain-specific RAG application on Databricks?

Select 2 answers
A.The total number of parameters in the model regardless of the domain.
B.The semantic relevance of the model to the target domain's terminology.
C.The availability of the model on the public Hugging Face repository.
D.The computational resource requirements for inference latency.
E.The color scheme of the model's documentation page.
AnswersB, D

Embedding models must map domain-specific terms to accurate vector representations to ensure relevant document retrieval. A model that has not been trained or fine-tuned on the specific domain’s jargon will produce poor vector alignments, leading to inaccurate RAG responses, regardless of the model's performance on general-purpose benchmarks.

Why this answer

Selecting the right embedding model requires balancing semantic accuracy within the specific domain and the operational cost of maintaining the model. By prioritizing domain-specific performance and resource efficiency, teams ensure that the RAG pipeline provides relevant search results without incurring excessive latency or compute costs, which are foundational for sustaining long-term generative AI production workloads.

Exam trap

Candidates often prioritize model popularity or parameter count, ignoring the critical balance between domain-specific semantic relevance and the operational latency costs of the model.

66
Multi-Selecthard

A GenAI engineer is developing an agent on Databricks that must call a custom Python function to query an inventory database. The agent must decide when to invoke the function and must receive the result back into its reasoning loop. Which TWO actions are required to expose the function as a tool the agent can call? (Choose two.)

Select 2 answers
A.Deploy the function as a separate Mosaic AI Model Serving endpoint and hard-code its URL in the system prompt.
B.Add the function's source code to the retrieval index so the model can retrieve it as context at inference time.
C.Register the function with the agent using the framework's tool decorator or tool registration API so it appears in the tool list passed to the model.
D.Define the function with a clear docstring and type hints so the agent can infer its name, description, and parameters.
E.Convert the function's return value into an embedding and store it in Vector Search for similarity lookups.
AnswersC, D

Even a well-documented function is invisible to the agent unless it is registered. The framework's tool decorator or registration call adds the function to the set of tools advertised to the LLM, enabling the model to emit a tool call. This registration step is what connects the Python callable to the agent's reasoning loop.

Why this answer

Exposing a Python function as an agent tool requires two things: a well-formed schema the model can understand, derived from the function's name, docstring, and type hints, and explicit registration so the function is included in the tool list the agent passes to the LLM. Together these let the model decide when to call the tool and let the agent execute it and return results into the reasoning loop. The remaining options neither register the function nor execute it.

Exam trap

The trap here is thinking that documenting a function or making it retrievable is enough, when the agent also needs explicit tool registration to actually invoke it.

67
MCQeasy

A team wants every MLflow evaluation run of their RAG chain to be reproducible months later, including the exact prompt template, model parameters, and retrieved context used. Which practice best ensures this reproducibility?

A.Log the evaluation as an MLflow run with the model parameters, prompt template, and input dataset version recorded as run metadata and artifacts.
B.Export the evaluation results to a CSV file stored on a personal laptop so the engineer can compare future runs manually.
C.Rely on the built-in default judge prompts so that no prompt configuration needs to be tracked between runs.
D.Store only the final aggregate metric values in a shared spreadsheet, since the underlying outputs can always be regenerated on demand.
AnswerA

An MLflow run records parameters, metrics, tags, and artifacts tied to a run ID. Capturing the prompt template, decoding parameters, and the versioned evaluation dataset as artifacts and metadata makes it possible to reconstruct exactly what was scored later, which is the core requirement for reproducibility.

Why this answer

Reproducibility depends on capturing inputs and configuration, not just scores. Logging the run with parameters, the prompt template, and a versioned dataset as artifacts and metadata ties every result to the exact conditions that produced it. That lets anyone rerun or audit the evaluation later, even after library upgrades or dataset changes.

Exam trap

The trap here is equating saved metric numbers with reproducibility, when replaying a run actually requires the inputs and configuration to be captured too.

68
MCQeasy

A team is deploying a RAG chatbot that answers questions from internal policy documents. The documents change frequently, and the team wants retrieval to reflect updates within minutes without re-running a full embedding job. The source Delta table already has change data feed enabled. Which approach should the team use to keep the Databricks Vector Search index current?

A.Increase the cluster size of the vector search endpoint
B.Configure the index to sync from the Delta table so changes are propagated incrementally
C.Drop and recreate the vector search index on a schedule
D.Add a metadata filter for the document modification timestamp
AnswerB

A vector search index that syncs from a Delta table with change data feed enabled picks up inserts, updates, and deletes incrementally, so newly modified policy chunks become searchable shortly after they land. This meets the requirement of reflecting updates within minutes without recomputing embeddings for the entire corpus.

Why this answer

A Delta-synced vector search index consumes change data feed events to apply inserts, updates, and deletes incrementally, keeping embeddings aligned with the source table. Rebuilding the index or tuning the endpoint addresses cost and throughput rather than freshness, and metadata filters only narrow an already-indexed set.

Exam trap

The trap here is confusing index freshness with endpoint capacity, when freshness depends on how changes flow from the Delta table into the index rather than on how much compute serves queries.

69
Multi-Selecthard

An organization is deploying a GenAI application using Databricks Model Serving. Which TWO steps are required to ensure the deployment environment handles model governance and observability effectively?

Select 2 answers
A.Enable Unity Catalog for the registered model and its versions.
B.Use MLflow to enable inference logging for model endpoints.
C.Hardcode API credentials within the model inference script.
D.Disable automatic scaling to maintain consistent latency.
E.Manually deploy the model via the standard cluster user interface.
AnswersA, B

Unity Catalog acts as the central governance layer for all data and AI assets. Enabling it for registered models provides unified access control, lineage tracking, and auditability, ensuring that only authorized users can deploy or modify models, which is a foundational requirement for robust enterprise AI security and governance.

Why this answer

Effective model governance and observability require integrating Unity Catalog for centralized access control and MLflow for tracking inference payloads. These components allow organizations to monitor model performance, lineage, and bias over time. Mastering these tools ensures that production deployments remain compliant, transparent, and auditable, which is essential for regulated industries using generative AI to make data-driven decisions while minimizing operational risks associated with model drift and unauthorized access.

Exam trap

Candidates often select only one of the two options, missing that governance (Unity Catalog) and observability (MLflow inference logging) are distinct, mandatory requirements for a production-grade deployment.

70
MCQhard

Refer to the exhibit. A developer wants to update this serving endpoint configuration to ensure it handles high-concurrency requests with consistent latency. Which change should be applied to the configuration?

A.Change 'workload_type' to 'GPU' and remove 'scale_to_zero_enabled'.
B.Set 'scale_to_zero_enabled' to 'false' and define 'min_provisioned_concurrency'.
C.Increase the 'model_version' to 'latest' to trigger automatic load balancing.
D.Add a 'timeout' parameter to the JSON configuration block.
AnswerB

Disabling scale-to-zero prevents the endpoint from shutting down, which eliminates cold-start latency. By defining 'min_provisioned_concurrency', the developer reserves a set number of replicas that are always running, ensuring that the system can handle concurrent requests immediately without waiting for infrastructure initialization, thus stabilizing latency under heavy load.

Why this answer

To handle high-concurrency with consistent latency, the configuration must move away from 'scale_to_zero_enabled: true' and utilize 'min_provisioned_concurrency'. Scaling to zero introduces cold-start latency that disrupts consistent response times required for high-concurrency production applications. By pinning minimum resources, the model remains active and ready to handle incoming traffic immediately, ensuring that performance remains predictable even during intermittent bursts of user activity within the production environment.

Exam trap

Candidates frequently select scale-to-zero options when high concurrency and consistent low latency are required, confusing cost savings with strict performance SLO requirements.

71
Multi-Selecthard

A team deploys a Mosaic AI Agent application through Databricks Apps and must let the app call a Model Serving endpoint that enforces Unity Catalog permissions on the underlying model. Which TWO configurations are required for the app to authenticate and be authorized to query that endpoint? (Choose two.)

Select 2 answers
A.Grant the app's service principal the appropriate privileges on the serving endpoint and the model it serves.
B.Configure the app to obtain an OAuth token for its service principal and pass it as a bearer token when calling the endpoint.
C.Register the app as a Delta Sharing recipient and share the model with it.
D.Enable public network access on the serving endpoint so the app can reach it without a token.
E.Embed a personal access token for a workspace admin in the app source so the endpoint always resolves an admin identity.
AnswersA, B

Databricks Apps run under a service principal, and Unity Catalog-governed serving endpoints check that identity's privileges. Granting the app's service principal access to the endpoint and the served model is what authorizes the call. Without those grants the app authenticates successfully but receives a permission error, so this is a required authorization step rather than an optional hardening measure.

Why this answer

Querying a Unity Catalog-governed serving endpoint requires both a valid identity and the right privileges. The app must present an OAuth token for its service principal so the request is authenticated, and that service principal must be granted access to the endpoint and the served model so authorization succeeds. Together these let the app call the endpoint without embedding human credentials.

Exam trap

The trap here is treating network exposure or data-sharing features as substitutes for authenticating the app's service principal and granting it endpoint privileges.

72
MCQhard

An engineer is designing a Databricks RAG application where the same retrieved context must be reused across several prompt variants during evaluation. They want to reduce token cost and improve consistency between variants. Which design choice best supports this?

A.Use a higher temperature during evaluation so the model explores more responses and the best one can be selected.
B.Compute the retrieval step once, store the retrieved documents in a Delta table keyed by query, and have each prompt variant read from that table.
C.Configure the model serving endpoint to cache identical requests so repeated prompts with the same context are served from cache.
D.Increase the chunk size in the Vector Search index so that each retrieval returns more context, reducing the need to call the retriever multiple times.
AnswerB

Caching retrieval results in a Delta table decouples retrieval from generation, so multiple prompt variants consume the same context without repeating embedding lookups or Vector Search calls. This reduces token and compute cost and ensures every variant is evaluated against identical context, which makes comparisons fair and reproducible.

Why this answer

Separating retrieval from generation and persisting the retrieved context lets several prompt variants share one retrieval result. This lowers cost because embeddings and Vector Search are invoked once, and it improves consistency because all variants see the same evidence. The alternatives change generation or chunking behavior but do not eliminate redundant retrieval.

Exam trap

The trap here is confusing generation-side caching with retrieval reuse, when the cost and consistency problem originates in repeated retriever calls.

73
Multi-Selectmedium

Which THREE practices are recommended when using MLflow for managing LLM experiments in Databricks?

Select 3 answers
A.Log model parameters such as temperature, top_p, and system prompts.
B.Store all model weights directly in the MLflow run object.
C.Define model signatures for input and output data validation.
D.Always delete old MLflow experiments to save storage space.
E.Use MLflow tags to categorize experiments based on project or model version.
AnswersA, C, E

Logging hyperparameters and prompt configurations is essential for experiment reproducibility. It allows developers to compare how different settings impact model output, ensuring that the best configuration can be identified and replicated consistently across different environments, which is vital for maintaining model performance quality over time.

Why this answer

Effective experiment management requires traceability, reproducibility, and structured documentation. Logging parameters like temperature and system prompts ensures results can be replicated. Using signature definitions allows the system to validate inputs, reducing integration errors.

Finally, tagging models aids in lifecycle management, allowing teams to distinguish between prototypes and production-ready versions. These practices ensure that teams can maintain high standards of rigor, enabling scalable AI operations within the enterprise environment.

Exam trap

Candidates often focus only on logging metrics while forgetting to log generative hyper-parameters, system prompts, and model signatures which are essential for LLM reproducibility.

74
MCQhard

Refer to the exhibit. What is the most likely cause for this error in a production RAG application?

A.The model is too large for the GPU memory.
B.Network connectivity between Model Serving and the vector store is misconfigured.
C.The model weights are corrupted.
D.The input prompt is too long for the model.
AnswerB

The timeout error explicitly points to a failure in establishing a connection to the vector store. This suggests that the network routing, VPC peering, or security group rules are blocking the communication path, which is a common deployment issue that must be addressed to restore RAG service functionality.

Why this answer

The error indicates a network timeout when connecting to the vector store. This is a common connectivity issue between the serving endpoint and the data source. Investigating network security groups, firewall rules, or DNS resolution issues within the Databricks environment is necessary to resolve it.

Ensuring robust, low-latency connectivity to the retrieval source is fundamental for reliable RAG performance and minimizing application downtime in production.

Exam trap

Candidates often blame the model or the code logic itself. They fail to recognize that RAG-specific errors are usually infrastructure-related, specifically network timeouts when the model attempts to query a vector store.

75
MCQmedium

Refer to the exhibit. A developer wants to enable monitoring for their deployed LLM endpoint. Given the current configuration, what must the developer change to ensure that request/response logs are captured for analysis?

A.Change the task to 'llm/v1/completions'.
B.Update 'auto_capture_request_payload' to true.
C.Add a new route with 0% traffic to the 'traffic_config'.
D.Increase the 'traffic_percentage' to 200%.
AnswerB

Setting 'auto_capture_request_payload' to true instructs the inference service to log incoming requests and outgoing responses to a managed Delta table. This is the direct configuration setting required to enable observability, ensuring that all data passed to the model is stored for later quality assessment and monitoring purposes.

Why this answer

The current configuration has 'auto_capture_request_payload' set to false, which prevents the logging of inference traffic. By updating this flag to true, the system will start capturing payloads into a Delta table. This is critical for monitoring model performance, drift, and quality.

Enabling this feature allows data teams to perform retrospective analysis, which is essential for iterating on model prompts and fine-tuning configurations based on real user interactions.

Exam trap

Candidates frequently attempt to create custom logging logic or external monitoring scripts, missing the built-in configuration flag 'auto_capture_request_payload' which is specifically designed for this purpose.

Page 1 of 5

Page 2

All pages