Courseiva

Databricks Certified Generative AI Engineer Associate (Databricks-GenAI-Assoc) — Questions 226–300

330 questions total · 5pages · All types, answers revealed

Page 3

Page 4 of 5

Page 5
226
MCQhard

Refer to the exhibit. An engineer is configuring a canary deployment for a churn prediction model. Based on the provided traffic configuration, what is the expected behavior of the endpoint?

A.The endpoint will error because total traffic must equal 100% for each model version individually.
B.90% of requests are routed to churn_v1 and 10% to churn_v2.
C.The endpoint will only route traffic to churn_v1 after churn_v2 reaches capacity.
D.The endpoint will round-robin requests between v1 and v2 regardless of the weights.
AnswerB

The traffic_config explicitly defines the routing weights for the served models. This split enables a canary release where the new version (v2) receives a small fraction of real-world traffic. This allows for performance benchmarking against the baseline (v1) before committing to a full deployment transition.

Why this answer

The configuration implements a traffic-splitting strategy, directing 90% of requests to 'churn_v1' and 10% to 'churn_v2'. This is a standard pattern for A/B testing or canary releases, allowing teams to validate new models in production with minimal risk. By monitoring the performance of the 10% traffic slice, engineers can decide whether to promote the model, ensuring stability before a full rollout.

Exam trap

Candidates often misread the traffic percentages or assume the configuration implies a different strategy, failing to parse the standard canary deployment logic defined in the endpoint traffic policy.

227
MCQmedium

Which command is used within the Databricks CLI to deploy a project defined by a Databricks Asset Bundle?

A.databricks workspace sync
B.databricks bundle deploy
C.databricks job create
D.databricks deploy project
AnswerB

This is the correct command to deploy resources specified in a Databricks Asset Bundle. It reads the local configuration, prepares the assets, and pushes them to the Databricks workspace defined in the 'target' section of the YAML file, completing the deployment process as specified in the configuration.

Why this answer

The 'databricks bundle deploy' command is the primary interface for deploying resources defined in a Databricks Asset Bundle. It interprets the 'databricks.yml' configuration, validates the resources, and then synchronizes the code and configuration with the targeted Databricks workspace. This CLI-first approach is essential for automating CI/CD pipelines, as it allows developers to trigger deployments programmatically without relying on manual browser-based interactions or fragile API scripts.

Exam trap

Test-takers frequently mix up CLI subcommands like 'bundle validate', 'bundle run', and 'bundle deploy' when asked how to push code to a workspace.

228
MCQhard

A retail company runs a customer-facing RAG assistant on a Mosaic AI Model Serving endpoint. The team wants every production request, response, and retrieved context to be captured automatically into a Unity Catalog Delta table so they can monitor quality and latency trends over time. Which action should they take?

A.Configure MLflow autologging on the notebook that calls the serving endpoint.
B.Register the model in Unity Catalog and enable lineage tracking on the catalog.
C.Enable inference tables on the Model Serving endpoint.
D.Enable Model Serving request logging to a cloud storage bucket.
AnswerC

Inference tables are the Databricks feature that automatically logs the request payload, response, and metadata from a Model Serving endpoint into a Unity Catalog Delta table. Enabling them on the endpoint requires no application code changes and captures retrieved context when the agent passes it through the payload, giving the team the monitoring substrate they want for trend analysis.

Why this answer

Inference tables are the native Databricks mechanism for capturing production traffic from a Model Serving endpoint into a governed Unity Catalog Delta table. They log payloads, responses, and metadata automatically, which is exactly what the team needs to analyze quality and latency over time. Other options either log only development activity or record lineage instead of runtime traffic, so they cannot satisfy the monitoring requirement.

Exam trap

The trap here is assuming that MLflow autologging or Unity Catalog lineage will capture live endpoint traffic, when only inference tables persist request and response payloads into a Delta table.

229
MCQmedium

Your organization requires an auditable record of all LLM evaluation results for compliance. Which Databricks feature provides the best centralized storage for these evaluation runs?

A.Databricks File System (DBFS) text logs
B.MLflow Experiments
C.Unity Catalog volume for raw CSV storage
D.Git commit history of the inference code
AnswerB

MLflow Experiments provide a structured environment to log parameters, metrics, and artifacts. This creates a highly auditable, searchable, and version-controlled record of every evaluation, which is ideal for compliance and tracking the history of model performance.

Why this answer

MLflow Experiments act as the centralized repository for all model development and evaluation runs. By logging evaluation results to MLflow, teams create a permanent, version-controlled audit trail. This allows stakeholders to compare different iterations of models, review evaluation metrics, and verify that quality gates were met before deployment, ensuring compliance with internal AI governance policies and external regulatory frameworks for model transparency.

230
MCQmedium

A team is preparing a fine-tuning dataset from customer reviews stored in a Delta table. They need to filter out reviews shorter than 20 tokens and reviews flagged as spam by a classifier, then write the result to a Unity Catalog table for training. Which approach best fits Databricks best practices?

A.Load the table into a pandas DataFrame on the driver, apply filtering, and write it back using the Databricks SQL connector row by row.
B.Create a view that filters short and spam reviews, then point the fine-tuning job at the view without materializing a new table.
C.Use Delta Live Tables with a streaming table that ingests all reviews and applies expectations to drop short and spam rows at write time.
D.Use a Spark DataFrame with a token-count UDF and a filter on the spam flag, then write the result to a Unity Catalog table with .write.mode("overwrite").saveAsTable().
AnswerD

Filtering with Spark transformations keeps the work distributed and lets Catalyst optimize the plan. A token-count UDF computes length per row, and a simple filter removes spam-flagged rows. Writing with saveAsTable into Unity Catalog creates a governed training table that downstream fine-tuning jobs can read with lineage intact.

Why this answer

Batch preprocessing of an existing Delta table is best done with Spark DataFrame transformations, which distribute token counting and filtering across executors. Writing the cleaned result to a Unity Catalog table creates a governed, reproducible training artifact. Driver-side pandas and row-by-row writes do not scale, views recompute on every read, and streaming pipelines are overkill for a one-time preparation step.

Exam trap

The trap here is choosing a view or a streaming pipeline for a one-time batch cleaning task, when a materialized Spark write is the appropriate pattern.

231
MCQeasy

A developer needs to store prompt templates, model parameters, and evaluation results for a GenAI application so that each iteration can be compared and reproduced later. Which Databricks capability should they use?

A.Delta Live Tables pipelines
B.MLflow Tracking with experiments and runs
C.Databricks SQL dashboards
D.Unity Catalog volumes
AnswerB

MLflow Tracking records parameters, metrics, artifacts, and tags per run within an experiment, which is exactly what is needed to compare prompt templates, model settings, and evaluation metrics across iterations. It is the standard mechanism in Databricks for reproducible GenAI development and integrates with Mosaic AI evaluation outputs.

Why this answer

MLflow Tracking provides experiments and runs that capture parameters, metrics, and artifacts, making it the right tool to version prompts, record model settings, and store evaluation results for comparison and reproduction. Data pipelines, volumes, and dashboards serve different purposes and do not offer run-based experiment tracking.

Exam trap

The trap here is assuming that any storage location for prompts, such as a Unity Catalog volume, also provides experiment tracking and run comparison.

232
Multi-Selectmedium

A Generative AI engineer is preparing a Delta table of product reviews for a retrieval-augmented generation application. The reviews contain HTML tags, inconsistent casing, and occasional very long paragraphs that exceed the embedding model's context window. The engineer wants to clean and normalize the text before chunking and embedding. Which two actions should the engineer take to directly address the stated quality issues? (Choose two.)

Select 2 answers
A.Drop every review shorter than 100 characters to reduce noise in the corpus.
B.Store the raw HTML in a separate column so it remains available for future parsing needs.
C.Split long paragraphs into overlapping chunks sized to the embedding model's token limit.
D.Strip HTML tags and normalize whitespace using a Spark transformation before chunking.
E.Convert all review text to uppercase to standardize casing across the corpus.
AnswersC, D

The reviews contain paragraphs that exceed the model context window, so chunking with overlap ensures no content is silently truncated and context is preserved across boundaries. Sizing chunks to the token limit directly addresses the stated overflow problem. Overlap helps retrieval quality by preventing answers that straddle a boundary from being split apart.

Why this answer

The reviews need markup removed and length controlled before embedding. Stripping HTML and normalizing whitespace cleans the text, while chunking with overlap sized to the model token limit prevents truncation of long paragraphs. Uppercasing, dropping short reviews, and archiving raw HTML do not resolve the stated noise and overflow problems.

Exam trap

The trap here is treating generic data hygiene steps like casing changes or row filtering as fixes, when the stated defects are markup noise and context-window overflow.

233
MCQmedium

A GenAI engineer is building a Databricks RAG application that answers questions over a Delta table containing 40 million support tickets. Users report that simple keyword lookups return irrelevant results because the tickets use inconsistent terminology. The engineer needs semantic retrieval that stays synchronized as new tickets stream in every few minutes. Which Databricks component should be used to serve this retrieval layer?

A.A Feature Store table that materializes ticket embeddings refreshed by a nightly job
B.A Databricks Vector Search index with a Delta Sync index over the tickets table
C.A Delta Live Tables pipeline that computes embeddings and writes them to a Parquet directory
D.A Databricks SQL warehouse running a LIKE query over the ticket text column
AnswerB

A Delta Sync index continuously tracks the source Delta table and automatically updates embeddings as new tickets arrive, so semantic similarity search stays current without manual reindexing. This directly addresses the inconsistent terminology because retrieval is based on embedding proximity rather than literal keyword matching, and it scales to millions of rows on Databricks-managed infrastructure.

Why this answer

Semantic retrieval over frequently changing data requires both embedding-based similarity and automatic synchronization with the source table. A Databricks Vector Search Delta Sync index provides exactly this: it computes and maintains embeddings as the Delta table changes and exposes a low-latency endpoint for similarity queries, so terminology variations are handled by vector proximity rather than literal string matching.

Exam trap

The trap here is assuming that a scheduled embedding job plus a plain table is equivalent to a managed vector index, when only a Delta Sync index keeps retrieval current within minutes without manual reindexing.

234
MCQmedium

When preparing data for a fine-tuning task, you realize the dataset is severely imbalanced. Which Databricks technique should you use to create a more balanced dataset?

A.Run a standard 'SELECT * FROM table' query.
B.Use the sampleBy() method to perform stratified sampling.
C.Increase the number of epochs during the training process.
D.Use a UDF to delete all rows that belong to majority classes.
AnswerB

Stratified sampling allows you to select specific proportions from each category, enabling effective oversampling of minority classes or undersampling of majority classes. This technique is standard in Spark for preparing balanced training datasets, ensuring that the model learns effectively from all classes, regardless of their original prevalence in the data.

Why this answer

Using Spark's 'sample()' or 'sampleBy()' transformation allows you to perform stratified oversampling or undersampling to balance classes. This ensures that the fine-tuned model doesn't become biased toward the most frequent categories in the dataset. Proper balancing is essential for ensuring robust model performance across all target classes, preventing the model from underperforming on rare but critical edge cases in real-world applications.

Exam trap

Candidates often suggest manual filtering or simple random sampling. Random sampling does not solve class imbalance, whereas stratified sampling specifically ensures minority classes are adequately represented in the training set.

235
MCQmedium

You are preparing a large dataset for fine-tuning a model using Databricks Delta Live Tables (DLT). Which configuration is best for ensuring data quality and lineage in this pipeline?

A.Use standard Spark SQL jobs without DLT.
B.Implement DLT Expectations to flag or drop invalid records.
C.Write all data to a temporary folder and clean it after training.
D.Rely on the LLM to automatically filter out low-quality data during training.
AnswerB

Expectations provide a declarative way to enforce quality in DLT pipelines. They ensure that the training data meets specific criteria, such as length or content validity. This reduces noise in the model training process, leading to better results and more reliable behavior, which is essential for enterprise GenAI deployments.

Why this answer

Delta Live Tables (DLT) with Expectations allows you to define constraints on data quality. By tagging data that fails these constraints, engineers ensure only clean data reaches the final training table. This provides built-in lineage and automatic quality monitoring.

In the context of GenAI, clean data is paramount, as noise in the training set leads to poor model performance and unpredictable behavior in downstream applications.

Exam trap

Candidates often overlook 'Expectations' as a quality tool, viewing them as optional. In a GenAI pipeline, they are essential for ensuring that only high-quality, validated data enters the model training.

236
MCQmedium

A GenAI engineer is building a RAG application on Databricks. They have registered a foundation model endpoint in Unity Catalog as a model. The application needs to query the endpoint, and the engineer wants to ensure that only members of the group 'genai_team' can invoke it. Which Unity Catalog privilege must be granted on the model object to allow invocation?

A.USE
B.SELECT
C.EXECUTE
D.CREATE
AnswerC

EXECUTE is the privilege required to invoke a model registered in Unity Catalog, including foundation model endpoints. Granting EXECUTE to the genai_team group on the model object allows its members to call the model endpoint from their applications, ensuring only authorized users can consume the model.

Why this answer

To invoke a model registered in Unity Catalog, a user must have the EXECUTE privilege on that model. Granting EXECUTE to the genai_team group ensures that only its members can call the model endpoint, aligning with the requirement to restrict access. Other privileges like SELECT, USE, or CREATE do not authorize model invocation.

Exam trap

The trap here is assuming that SELECT or USE grants invocation rights on a model, when actually EXECUTE is the specific privilege required.

237
MCQhard

A generative AI engineer is designing an agent on Databricks that uses a LangChain agent with tool-calling capabilities. The agent must call a Databricks SQL warehouse to run queries and a Vector Search index for retrieval. Which design consideration is most important for controlling agent behavior in production?

A.Define clear tool descriptions and limit the agent's available tools to only those required for the task.
B.Use the largest available foundation model for the agent to maximize reasoning quality.
C.Enable verbose logging of every intermediate step and store the logs in a Delta table for later review.
D.Increase the agent's maximum iterations to a very high number so it can always complete complex tasks.
AnswerA

The agent selects tools based on their descriptions, so precise descriptions and a minimal toolset reduce incorrect tool selection and unintended actions. Limiting tools also narrows the blast radius if the agent misbehaves. This is a core design principle for production agents that must be predictable and governable.

Why this answer

Agent behavior is driven by tool descriptions and the set of tools available. Clear descriptions help the model choose the right tool, and a minimal toolset reduces the chance of unintended calls. Iteration limits, logging, and model size affect performance or observability but do not directly control which tools the agent invokes.

Exam trap

The trap here is assuming that a more capable model or more iterations will make an agent reliable, when reliability comes from precise tool definitions and a constrained toolset.

238
MCQmedium

A Generative AI engineer is building a Delta Live Tables pipeline that ingests raw JSON event logs into a bronze table, then uses ai_query to classify each event's free-text field. The classification call is expensive, so the engineer wants to avoid re-running it on events that have already been processed in previous pipeline updates. The source table is append-only and new events arrive continuously. Which Delta Live Tables feature should the engineer configure on the bronze table to prevent reprocessing of previously ingested rows?

A.Configure the bronze table as a streaming table so it processes only new data on each pipeline update.
B.Add a QUALIFY clause that filters out rows where the classification column is already populated.
C.Enable change data capture by setting pipelines.cdcEnabled to true on the bronze table.
D.Set the table property pipelines.autoOptimize.managed to true on the bronze table.
AnswerA

Streaming tables in Delta Live Tables read incrementally, so on each pipeline update only newly arrived source rows flow through the ai_query classification. Previously ingested events are not re-read, which is exactly what avoids paying for repeated classification calls. This is the standard pattern for append-only ingestion with expensive downstream transformations.

Why this answer

An append-only bronze ingestion table should be defined as a streaming table so each pipeline update consumes only new source records. This makes expensive operations like ai_query run once per new event rather than on the full history. Materialized views and batch-style definitions re-read all source data, and storage properties like Auto Optimize only affect file layout, not incremental read behavior.

Exam trap

The trap here is assuming that a table property or filter clause can prevent reprocessing, when incremental read semantics in Delta Live Tables come from declaring the dataset as a streaming table.

239
MCQeasy

What is the primary benefit of using MLflow Model Evaluation for generative AI applications compared to manual evaluation methods?

A.It automatically rewrites the prompt templates based on user feedback.
B.It enables consistent, repeatable evaluation metrics across different model versions.
C.It removes the need for ground truth datasets during the testing phase.
D.It creates the vector index automatically from raw unstructured data.
AnswerB

MLflow offers a consistent framework to calculate standard metrics like faithfulness and relevance. This ensures that every model update is evaluated using the same methodology, providing a reliable baseline for comparing performance improvements or regressions across versions.

Why this answer

MLflow Model Evaluation provides a standardized framework to calculate metrics like RAGAS and perform side-by-side comparisons of model versions. By automating this process, organizations can ensure consistency across releases, track performance history over time, and reduce human bias. This systematic approach allows teams to make data-driven decisions when deciding whether to promote a model to production, ensuring reliability and auditability in the AI lifecycle.

Exam trap

Candidates often assume MLflow Model Evaluation automatically tunes hyper-parameters, overlooking its core role in providing consistent, repeatable evaluation metrics across model versions.

240
MCQeasy

Which component in the Databricks GenAI stack is responsible for orchestrating the flow between data retrieval, prompt construction, and model invocation?

A.Unity Catalog.
B.MLflow Tracking.
C.Mosaic AI Agent Framework.
D.Databricks SQL.
AnswerC

The Mosaic AI Agent Framework provides the tools and abstractions needed to build, evaluate, and deploy agentic AI applications. It acts as the orchestration layer that connects data retrieval tools with model endpoints and manages the prompt engineering lifecycle, making it the correct choice for defining complex AI application logic.

Why this answer

Mosaic AI Agent Framework is designed to manage the complex orchestration required for RAG and agentic workflows. By providing a structured way to define tools, chains, and prompt strategies, it allows developers to build sophisticated applications that dynamically retrieve data and interact with LLMs. This orchestration layer is vital for building robust, maintainable AI applications where the logic needs to be clearly separated from the underlying model serving infrastructure.

Exam trap

Candidates incorrectly attribute prompt construction and retrieval orchestration to basic model serving or MLflow, ignoring the specialized role of the Mosaic AI Agent Framework.

241
MCQmedium

A developer is building a RAG application using Mosaic AI Model Serving. They need to ensure that the model endpoint logs inference requests and responses for audit purposes. Which configuration parameter should they enable?

A.Enable 'auto_capture_logs' in the endpoint environment variables.
B.Configure the 'inference_table_config' block in the endpoint request JSON.
C.Set the 'log_level' to 'DEBUG' in the Serving Endpoint settings.
D.Enable 'external_logging' in the Unity Catalog schema settings.
AnswerB

The 'inference_table_config' parameter is the specific configuration key required to enable automated request and response logging in Databricks Model Serving. This maps incoming requests directly to a specified Delta table, providing a structured, queryable record of all model interactions, which is essential for ongoing performance monitoring.

Why this answer

Enabling 'inference_table_config' in the model serving endpoint configuration is the standard Databricks approach for capturing telemetry data. This feature automatically writes request and response payloads to a Delta table, enabling compliance auditing, model monitoring, and drift detection. Understanding this integration is critical for production-grade AI deployments where transparency, debugging, and regulatory logging are mandatory requirements for enterprise-scale machine learning operations.

Exam trap

Candidates often assume logging is automatic or handled by the workspace. You must explicitly configure the 'inference_table_config' to capture request/response data into a Delta table for auditing.

242
MCQmedium

Which Unity Catalog feature is best suited for sharing data assets with users outside of your Databricks account?

A.Unity Catalog cross-account roles
B.Delta Sharing
C.External locations
D.Workspace federation
AnswerB

Delta Sharing is designed specifically for secure data sharing with entities outside your organization. It supports sharing data from your Unity Catalog metastore to any Delta Sharing recipient, even if they do not use Databricks, providing a platform-agnostic way to collaborate on large datasets.

Why this answer

Delta Sharing is an open-protocol standard that allows organizations to share data with users regardless of whether they have a Databricks account. It enables secure, read-only data sharing across different platforms and cloud environments. This is essential for organizations that need to collaborate with partners, vendors, or external clients while maintaining centralized control and auditability within Unity Catalog.

Exam trap

Candidates often suggest 'Unity Catalog' or 'Workspace Sharing', not realizing that Delta Sharing is the specific protocol designed for cross-account and cross-platform data distribution.

243
Multi-Selectmedium

When evaluating LLM outputs, which TWO metrics are most appropriate for measuring the 'quality' of a response in a RAG system? (Choose two)

Select 2 answers
A.Faithfulness
B.Inference Latency
C.Answer Relevance
D.GPU Utilization
E.Training Iterations
AnswersA, C

Faithfulness checks if the generated answer is derived from the retrieved context. This is the primary metric for preventing hallucinations, which are the biggest risk in RAG deployments. A faithful model stays within the bounds of its provided information, ensuring reliability and accuracy for the end user.

Why this answer

Answering quality in RAG requires looking at two distinct dimensions: the correctness of the content (faithfulness) and the relevance of the answer to the user's intent. Faithfulness ensures the model remains grounded in the provided evidence, while answer relevance captures whether the model actually addressed the specific user question. These metrics provide a balanced view, ensuring the system is both accurate and useful, which is essential for end-user satisfaction.

244
MCQmedium

A developer needs to deploy a custom Python model that requires non-standard library dependencies. Which MLflow feature should the developer use to specify these environment requirements during model logging?

A.Model Schema definition
B.MLflow environment requirements (pip_requirements)
C.Global workspace library settings
D.Custom model signatures
AnswerB

By explicitly providing a list of required libraries via the `pip_requirements` argument during the logging process, you guarantee that the Model Serving environment will install them before loading the model. This is the industry-standard way to ensure that complex, custom model code runs successfully in production.

Why this answer

When logging a custom model with MLflow, the developer should use the `pip_requirements` or `conda_env` parameter in the `mlflow.pyfunc.log_model` function. This ensures that the environment is correctly packaged and reproducible. During deployment, the Databricks Model Serving service reads these requirements to recreate the identical software environment, preventing runtime errors caused by missing dependencies when the model is loaded and served in the production environment.

Exam trap

Candidates often assume the environment is automatically captured from the local machine. They fail to explicitly define 'pip_requirements' or 'conda_env', causing deployment failures when the server lacks local dependencies.

245
MCQmedium

A data scientist is using Databricks to fine-tune a large language model. The training data is stored in a Unity Catalog table that contains sensitive customer information. The data scientist needs to read the table but should not be able to see the raw values of certain columns. Which Unity Catalog feature should be used to dynamically mask the sensitive columns based on the user's group membership?

A.Dynamic column masking
B.Column-level encryption
C.Attribute-based access control
D.Row-level security
AnswerA

Dynamic column masking allows you to define a masking function that returns different values depending on the user's group membership. This enables sensitive columns to be masked for unauthorized users while remaining visible to authorized ones, exactly matching the scenario.

Why this answer

Dynamic column masking in Unity Catalog enables the data scientist to see masked values for sensitive columns based on their group membership. By applying a column mask that checks group membership, the raw values are hidden from unauthorized users while still allowing the data scientist to read the table for fine-tuning.

Exam trap

The trap here is confusing row-level security with column-level masking, or assuming that encryption provides dynamic masking.

246
Multi-Selectmedium

A data engineer is preparing a large Delta table of conversation logs for embedding generation. The table has frequent small appends, and the engineer needs to reduce file fragmentation and improve read throughput before the embedding job runs. Which two actions should the engineer take? (Choose two.)

Select 2 answers
A.Run VACUUM with a retention of zero hours to delete old files and free space before the embedding job.
B.Increase the number of shuffle partitions to the maximum supported value so that each task writes a single small file.
C.Enable auto compaction and optimized writes on the Delta table so small writes are coalesced automatically.
D.Run OPTIMIZE on the table to compact small files, optionally with ZORDER on columns frequently used in filters.
E.Convert the table to a Parquet directory and rely on the file system to merge small files during reads.
AnswersC, D

Auto compaction merges small files after a write, and optimized writes shuffle data so fewer, larger files are produced in the first place. Together they keep fragmentation low between manual maintenance runs, which suits a table receiving frequent small appends. This directly addresses the file-count problem before the embedding job reads the data.

Why this answer

File fragmentation from frequent appends is addressed by consolidating small files. OPTIMIZE with optional ZORDER compacts and clusters existing data, while auto compaction and optimized writes prevent new fragmentation from accumulating. Together they reduce the number of files the embedding job must open and can improve predicate pushdown, directly improving read throughput.

Exam trap

The trap here is confusing VACUUM, which deletes unreferenced files, with OPTIMIZE, which rewrites live data into fewer files.

247
MCQmedium

When evaluating the performance of a RAG application, which metric is most useful for measuring the quality of the retrieved context?

A.Total number of documents retrieved.
B.The total latency of the RAG pipeline.
C.Context Precision.
D.The number of parameters in the LLM.
AnswerC

Context Precision evaluates the ratio of relevant documents within the retrieved set. High precision ensures that the LLM is provided with high-quality, actionable context, which is the cornerstone of effective RAG. Monitoring this metric helps identify when the retrieval strategy needs adjustment to improve the application's overall accuracy.

Why this answer

Context Precision measures how much of the retrieved information is actually relevant to the user query. This is a critical metric for RAG systems because high retrieval precision directly reduces the noise in the LLM input, leading to more accurate and focused answers. Measuring this allows engineers to iteratively improve the retrieval pipeline, which is essential for maximizing the utility of the application.

Exam trap

Candidates frequently confuse evaluation metrics, mixing up context precision (what was retrieved vs. relevant) with generation metrics like answer correctness or faithfulness.

248
MCQmedium

A Data Engineer needs to ensure that PII data in a Delta table is masked before serving it to non-privileged users. Which Databricks feature provides the most efficient, centralized control for this requirement?

A.Apply Row-Level Security filters using Delta table constraints.
B.Create materialized views with hard-coded redacted values.
C.Utilize Unity Catalog column-level masking functions.
D.Implement Spark UDFs to perform in-memory data masking.
AnswerC

Unity Catalog enables defining dynamic masks on columns using SQL functions. This centralizes security policy enforcement, applying masking rules consistently across all users and compute clusters. It is the most robust method for PII protection because it prevents unauthorized visibility without altering the underlying raw data storage files.

Why this answer

Unity Catalog's dynamic views and column-level masking policies are the standard for securing PII. By applying functions like MASK or current_user() within a view definition, administrators decouple security logic from physical table storage. This approach is essential for compliance, ensuring that sensitive data is hidden at query time without duplicating datasets, thus maintaining a single source of truth while enforcing granular access control across all Databricks workspaces and compute resources.

Exam trap

Candidates often suggest creating separate tables or filtering via application code. Unity Catalog masking is the centralized, efficient way to handle this without duplicating data or creating security gaps.

249
MCQmedium

A GenAI engineer has a RAG application whose retrieval step uses Databricks Vector Search. Users report that answers are sometimes irrelevant because the retriever pulls chunks from documents the user is not authorized to see. The engineer must enforce per-user document ACLs at query time without re-indexing the corpus. Which approach should the engineer take?

A.Post-filter the retrieved chunks in the application layer by querying Unity Catalog for the user's group memberships and dropping chunks that fail the check.
B.Create a separate Vector Search index per user and have the application route each request to the user-specific index.
C.Add a metadata column that stores the document's ACL principals, then pass a filter string with the caller's identity to the Vector Search query.
D.Enable Delta Sharing on the source Delta table and rely on the workspace's existing table ACLs to filter Vector Search results.
AnswerC

Databricks Vector Search supports metadata columns on the index and accepts a filters argument at query time. Storing allowed principals (users or groups) in a metadata column and passing the caller's identity as a filter ensures only authorized chunks are returned, with no re-indexing. This is the documented pattern for per-user authorization in RAG retrieval on Databricks.

Why this answer

Vector Search indexes can carry metadata columns that are queryable through the filters parameter, letting the application pass the caller's identity so only permitted chunks are returned. This satisfies per-user ACL enforcement at query time without rebuilding the index when permissions change, and it avoids leaking unauthorized content into the application tier. The other approaches either do not affect index results or are operationally impractical at scale.

Exam trap

The trap here is assuming that Unity Catalog or Delta Sharing ACLs automatically propagate to a Vector Search index's query results, when in fact the index must be filtered explicitly via metadata.

250
MCQhard

Which CI/CD approach for model deployment best minimizes downtime during a model update?

A.Delete the old model and immediately deploy the new version.
B.Use a blue-green deployment pattern.
C.Deploy all updates directly to the production endpoint.
D.Schedule deployments during peak traffic hours.
AnswerB

Blue-green deployment allows for seamless traffic shifting between model versions. By maintaining two separate environments and routing traffic only when the new version is verified as stable, teams can ensure zero downtime during upgrades, which is essential for maintaining a robust, professional production AI service environment.

Why this answer

A blue-green deployment strategy is the gold standard for minimizing downtime. By spinning up a new version of the model (green) alongside the existing one (blue) and switching traffic only after verifying the health of the green endpoint, engineers ensure zero downtime and an immediate rollback path if issues occur. This approach is critical for high-availability AI services where any interruption would negatively impact users.

Exam trap

Candidates often suggest 'redeploying' or 'updating in place', which causes temporary downtime. They overlook that blue-green deployment is the specific pattern designed to avoid this service interruption.

251
MCQeasy

When evaluating a generative AI model using the Mosaic AI Model Evaluation tool, what is the primary purpose of providing a 'baseline' dataset?

A.To increase the number of tokens the model can generate.
B.To provide a reference for comparative performance analysis.
C.To act as a secondary training set for fine-tuning.
D.To optimize the vector database search latency.
AnswerB

A baseline allows developers to quantitatively compare current model performance against known good results. This comparison is vital for validating that new model versions or updated RAG configurations maintain or exceed performance standards, facilitating data-driven decisions about whether to promote a new version to production.

Why this answer

A baseline dataset serves as a ground-truth reference point. By comparing the model's current outputs against this known high-quality set, engineers can objectively measure performance degradation or improvement. This benchmarking approach is fundamental to scientific model evaluation, ensuring that updates to the model or data retrieval logic do not silently introduce regressions in the quality of the generated answers during the deployment process.

Exam trap

Candidates incorrectly assume a baseline dataset is used for fine-tuning the model, rather than understanding its purpose as a reference point for comparative evaluation.

252
Multi-Selecthard

An engineer is designing a Databricks RAG application that must support multi-turn conversations where follow-up questions refer to earlier turns. They want the retrieval step to remain accurate as the conversation progresses. Which TWO design elements should they include? (Choose two.)

Select 2 answers
A.Maintain a conversation history buffer and condense prior turns into a standalone query before invoking Vector Search.
B.Increase the number of retrieved chunks for every turn to compensate for conversational ambiguity.
C.Include the previous assistant response as additional context in the prompt sent to the generation model.
D.Store the conversation history in the prompt and pass it unchanged to the retriever as the search query.
E.Disable metadata filtering during multi-turn conversations so more historical documents are eligible.
AnswersA, C

Follow-up questions often contain pronouns or omitted subjects that are ambiguous in isolation. Condensing the history into a self-contained query gives the retriever enough context to find relevant chunks, improving recall for multi-turn interactions. This is a standard pattern in conversational RAG design on Databricks.

Why this answer

Conversational RAG needs two coordinated elements: a rewritten, standalone query for retrieval, and conversation history for generation. Rewriting resolves references so Vector Search returns on-topic chunks, while including prior turns helps the model produce coherent follow-ups. Simply enlarging the retrieval window or removing filters does not address the ambiguity that multi-turn questions introduce.

Exam trap

The trap here is assuming that adding more conversation text to the retriever query improves multi-turn accuracy, when it actually dilutes the embedding and harms relevance.

253
Multi-Selecthard

Which TWO factors are most important when selecting a chunking strategy for text data prior to vectorization?

Select 2 answers
A.The maximum input token length of the target embedding model.
B.The file format of the source documents (e.g., PDF vs. Word).
C.The desired level of semantic granularity for retrieval.
D.The total number of documents in the corpus.
E.The storage cost of the resulting vector database.
AnswersA, C

Embedding models have strict limits on the number of tokens they can process in a single sequence. Exceeding this limit leads to truncation, which destroys semantic context and makes the resulting vectors inaccurate. Aligning chunk size with model token limits is a critical technical requirement for successful data preparation.

Why this answer

Selecting a chunking strategy requires balancing semantic context and technical limitations like context window size. If chunks are too small, they lack meaning; if too large, they exceed the limits of the embedding model and introduce irrelevant noise. Proper chunking is vital for maximizing the accuracy of RAG systems, as it determines the granularity of the information available for the retriever to present to the LLM.

Exam trap

Candidates often focus solely on the embedding model's limits while forgetting that the chunking strategy must also align with the business goal of retrieving semantically meaningful information.

254
MCQmedium

A data engineer needs to ensure that sensitive PII columns are masked for specific groups while remaining visible to analysts. Which Unity Catalog feature should be used to implement this requirement?

A.Row-level security filters
B.Dynamic data masking
C.Credential passthrough
D.Attribute-based access control (ABAC)
AnswerB

Dynamic data masking is specifically designed to redact or transform sensitive column data in real-time based on the user's identity or group. By attaching a masking function to a column in Unity Catalog, data engineers can ensure that analysts see masked results without modifying the underlying data files.

Why this answer

Unity Catalog dynamic data masking allows administrators to apply functions to columns that redact or obfuscate data based on the user's role or group membership. By using SQL functions like mask_email or custom UDFs within a masking policy, the data remains consistent at the physical layer while presenting transformed values at query time. This ensures compliance with privacy regulations without creating multiple copies of datasets.

Exam trap

Candidates often suggest creating multiple views or tables with filtered data. This is inefficient and prone to errors; dynamic data masking is the correct, centralized feature for this.

255
MCQhard

A GenAI engineer is designing an agentic application on Databricks that uses a foundation model to decide which external tools to call. The team wants the agent to be able to invoke a Databricks SQL warehouse query and a Python function registered as a Unity Catalog function, and they need the model to select tools based on natural language requests. Which design element is required for the model to select and invoke these tools correctly?

A.Enable inference tables on the model serving endpoint to capture tool calls.
B.Store the tool list in a Unity Catalog volume and have the agent read it before each request.
C.Provide tool definitions with names, descriptions, and parameter schemas to the model in the request.
D.Fine-tune the foundation model on the team's historical tool-call logs.
AnswerC

Function-calling capable models select tools based on the tool definitions supplied in the request, which include the tool name, a natural language description, and a JSON schema for parameters. Without these definitions, the model has no knowledge of available tools or how to structure arguments. Supplying accurate definitions for the SQL warehouse query and the Unity Catalog function enables correct tool selection and argument generation.

Why this answer

Function-calling models rely on tool definitions passed in the request, including name, description, and parameter schema, to decide which tool to invoke and how to format arguments. Providing these definitions for the SQL warehouse query and Unity Catalog function gives the model the runtime contract it needs. Fine-tuning, volume storage, and inference tables do not supply that contract and therefore cannot enable correct tool selection.

Exam trap

The trap here is assuming that fine-tuning or logging replaces the runtime tool definitions, when the model needs explicit function schemas in every request to invoke tools correctly.

256
MCQhard

An engineer is deploying a RAG application whose retrieval step calls an external vector database over the public internet. The serving endpoint must reach that database, but corporate policy forbids exposing credentials in the model artifact and forbids public egress from the serving environment. Which configuration satisfies both constraints?

A.Have the retrieval function read the credential from a Delta table at query time and rely on the default public internet route for database access.
B.Embed the database credential as an environment variable inside the logged MLflow model so the serving container can read it at startup.
C.Store the database credential in a Databricks secret scope, reference it from the endpoint configuration, and enable network connectivity through a private link or VPC configuration for the serving endpoint.
D.Package the credential into a Python wheel installed as a model dependency and configure the endpoint with serverless compute only.
AnswerC

Secret scopes keep credentials out of the artifact and inject them at runtime, while private connectivity settings for Model Serving allow egress to the external database without traversing the public internet. Together they meet both the credential and network policy requirements.

Why this answer

Credentials belong in a Databricks secret scope so they are injected at runtime rather than stored in the artifact, and reaching a private external service from Model Serving requires explicit private connectivity rather than default public routing. Satisfying both constraints means combining secret injection with a private network path.

Exam trap

The trap here is solving only the credential problem and assuming network egress is automatically handled, when private connectivity to the external database must be configured separately.

257
MCQmedium

Which approach is recommended for managing secrets, such as API keys for external LLM providers, in Databricks?

A.Hardcoding keys in a configuration file within a Repo.
B.Storing keys in Databricks Secrets.
C.Passing keys as plain text arguments in a job parameter.
D.Storing keys in a public S3 bucket with restricted access.
AnswerB

Databricks Secrets provides an encrypted, centralized way to manage sensitive keys. By using 'dbutils.secrets.get', you can inject these values into your code at runtime without them ever appearing in cleartext within your scripts, notebooks, or version control, ensuring a robust security posture for your production application.

Why this answer

Databricks Secrets provides a secure, encrypted storage mechanism for sensitive information, preventing credentials from being hardcoded in notebooks or scripts. This is a foundational security best practice. By using scope-based secret management, you ensure that only authorized users or service principals can access the credentials, keeping your RAG application secure and compliant with enterprise security standards.

Exam trap

Candidates often suggest environment variables or hardcoding secrets in notebooks, mistakenly believing these are acceptable for internal development, which violates basic Databricks security compliance standards.

258
MCQmedium

Refer to the exhibit. A developer is deploying a model using the provided JSON configuration. What is the primary benefit of setting 'scale_to_zero_enabled' to true in this production RAG application?

A.It increases the throughput of the model by optimizing memory allocation during inference.
B.It ensures that the model always stays warm to minimize latency for every user request.
C.It reduces operational costs by shutting down compute resources when the endpoint is not serving traffic.
D.It enables high availability by automatically replicating the model across multiple availability zones.
AnswerC

The primary purpose of enabling scale-to-zero is cost reduction. In environments where request patterns are sporadic or non-continuous, this feature ensures that the company is only billed for the compute resources actually consumed during active inference periods, rather than paying for constant uptime when no requests are present.

Why this answer

Setting 'scale_to_zero_enabled' to true allows the Databricks Model Serving endpoint to automatically shut down compute resources when no requests are being processed. This is highly beneficial for cost optimization, as it eliminates idle runtime costs during periods of inactivity. When a new request arrives, the service automatically initializes the model, ensuring cost efficiency without requiring manual intervention to scale the underlying infrastructure up or down.

Exam trap

Candidates often confuse 'scale_to_zero_enabled' with model accuracy improvements or latency reduction, forgetting that its primary purpose is exclusively cost optimization by eliminating idle compute charges during inactive periods.

259
MCQmedium

When deploying an application using Databricks Asset Bundles (DABs), which file is the primary entry point to define the project structure and configuration?

A.requirements.txt
B.databricks.yml
C.setup.py
D.bundle.json
AnswerB

The databricks.yml file is the core configuration manifest for Databricks Asset Bundles. It contains the definitions for resources such as jobs, pipelines, and models, as well as deployment targets. It drives the 'bundle deploy' command, ensuring the environment is orchestrated exactly as specified in the configuration.

Why this answer

Databricks Asset Bundles use a YAML-based configuration file, typically named 'databricks.yml', to define the project's resources, targets, and settings. This file acts as the single source of truth for the project's lifecycle, enabling consistent deployments across development, staging, and production environments. By using a declarative configuration, teams can version-control their entire infrastructure, ensuring that deployment patterns are standardized and repeatable across the organization.

Exam trap

Candidates occasionally confuse the project structure with individual notebook settings, failing to identify 'databricks.yml' as the central, mandatory file for defining the entire bundle's lifecycle and configuration.

260
MCQhard

An engineer registers a GenAI agent to Unity Catalog and enables inference tables on its Mosaic AI Model Serving endpoint. They want to automatically detect when response quality degrades in production without waiting for human review. Which capability should they configure to achieve continuous automated quality monitoring?

A.Configure autoscaling on the serving endpoint so additional model replicas absorb traffic spikes and stabilize response quality.
B.Set up a scheduled job that calls the endpoint with synthetic prompts and records latency percentiles to a Delta table for trending.
C.A Databricks Lakehouse Monitoring monitor over the inference table, joined with a labeled evaluation dataset to compute quality metrics on a schedule.
D.Enable model signature enforcement on the endpoint so requests that violate the expected schema are rejected before scoring.
AnswerC

Inference tables capture request and response payloads in a Delta table, and Lakehouse Monitoring can build a monitor over that table to compute profile and drift metrics on a schedule. Joining with labeled ground-truth data lets the monitor compute quality metrics continuously, providing automated degradation detection without manual review of each response.

Why this answer

Inference tables persist production request and response payloads as Delta, which makes them queryable by Lakehouse Monitoring. Building a monitor over that table, especially when joined with labeled data, yields scheduled quality and drift metrics that surface degradation automatically. This combination is the supported path for continuous, automated quality monitoring on Databricks rather than ad hoc manual inspection.

Exam trap

The trap here is confusing operational telemetry such as latency and autoscaling with semantic quality monitoring, which requires payload capture plus a monitor.

261
MCQmedium

A GenAI team at a retail company has built a RAG chatbot on Databricks that answers customer questions from a product catalog stored in Delta Lake. The catalog is updated nightly, and the team wants the chatbot to reflect those updates without manual intervention. They are deciding how to keep the Vector Search index synchronized. Which approach should they use?

A.Use a Databricks SQL scheduled query to copy the catalog table into a Parquet file and point the retriever at that file each night.
B.Configure a Vector Search Delta Sync index on the catalog table so the index is refreshed automatically when the source table changes.
C.Create a standard Vector Search index and schedule a daily notebook that deletes the index and recreates it from the catalog table.
D.Enable Change Data Feed on the catalog table and rely on it to update the Vector Search index automatically.
AnswerB

A Delta Sync index in Databricks Vector Search tracks a Delta table as its source and can automatically sync new or changed rows, so nightly catalog updates flow into the index without a separate pipeline. This matches the requirement for unattended freshness and keeps retrieval aligned with the Delta Lake table.

Why this answer

A Delta Sync index is the Databricks Vector Search construct that keeps an index aligned with a Delta table, including incremental changes. Because the catalog is a Delta table that changes nightly, configuring Delta Sync lets the chatbot retrieve current product information without a custom rebuild job. The other choices either fail to update the index or introduce unnecessary full recomputation.

Exam trap

The trap here is assuming that enabling Change Data Feed on the source table is sufficient to refresh a Vector Search index, when the index itself must be configured for Delta Sync.

262
MCQmedium

A GenAI engineer is building an agent with LangChain on Databricks. The agent must call a Unity Catalog function `catalog.schema.get_weather` to fetch current weather. They want the LLM to decide when to invoke this function. Which LangChain component should they use to expose the Unity Catalog function to the LLM?

A.Databricks Vector Search retriever
B.Databricks Unity Catalog function as a LangChain tool
C.MLflow pyfunc model wrapper
D.Databricks SQL Connector
AnswerB

LangChain provides a `UCFunctionToolkit` that wraps Unity Catalog functions as LangChain tools. This allows the LLM to see the function's metadata and invoke it when needed. The toolkit handles authentication and execution via the Databricks SDK, making it the correct choice for integrating Unity Catalog functions into an agent.

Why this answer

To let an LLM decide when to call a Unity Catalog function, the function must be presented as a LangChain tool. The `UCFunctionToolkit` from Databricks integrates with LangChain, automatically generating tool definitions from Unity Catalog functions. This enables the agent to invoke the function dynamically based on the conversation, which is essential for building responsive GenAI applications.

Exam trap

The trap here is confusing data retrieval components like Vector Search with tool-calling mechanisms, overlooking that Unity Catalog functions require a specific toolkit to be exposed as tools.

263
MCQmedium

An engineer is preparing to deploy a Mosaic AI Agent application with Databricks Asset Bundles. The bundle defines the agent, the serving endpoint, and a job that refreshes the vector index. The engineer wants the deployment to target a staging workspace and a production workspace with different endpoint names and different Unity Catalog catalog names, without editing files between deployments. Which approach should the engineer use?

A.Hard-code the production endpoint and catalog names and use a post-deployment notebook to rename resources in staging.
B.Define bundle variables for the endpoint name and catalog name, and set their values per target in the bundle's databricks.yml targets section.
C.Maintain two separate bundle projects, one per workspace, and keep them synchronized manually.
D.Store the environment-specific values in a secrets scope and read them at runtime inside the agent code.
AnswerB

Databricks Asset Bundles support variables that can be overridden per target, so the same bundle definition can deploy to staging and production with different endpoint and catalog names. This keeps a single source of truth while allowing environment-specific values, which is exactly what the scenario requires. The engineer selects a target at deploy time and no files are edited.

Why this answer

Databricks Asset Bundles let a single declarative project target multiple workspaces through targets, and variables can be given different values per target. Defining variables for the endpoint and catalog names and overriding them in each target lets the same bundle deploy correctly to staging and production without editing files. This preserves one source of truth and matches the scenario's requirement.

Exam trap

The trap here is reaching for secrets or duplicate projects to handle environment differences, when bundle variables overridden per target are the built-in mechanism for non-sensitive, deployment-time values.

264
MCQhard

A GenAI engineer is building a multi-step agent that uses Databricks Foundation Model APIs. The agent must decide when to call a weather tool and when to answer directly. The engineer wants to ensure the agent's decision-making is reliable and that failures in tool calls are handled gracefully. Which design approach should the engineer use?

A.Chain multiple LLM calls where the first call always invokes the weather tool and the second call formats the answer.
B.Implement a ReAct-style loop where the LLM outputs a thought, action, and action input, and the agent executes the tool and feeds back the observation.
C.Fine-tune the LLM on examples of weather queries and answers, then deploy it without tools.
D.Use a single prompt that instructs the LLM to output a JSON with either an answer or a tool call, and parse it once.
AnswerB

A ReAct-style loop structures the agent's reasoning and tool use, allowing it to decide when to call the weather tool and incorporate the result. It also provides a clear place to handle tool errors by catching exceptions and feeding error messages back to the LLM for recovery.

Why this answer

A ReAct-style loop enables the agent to reason about whether a tool is needed, execute it, and observe the result before deciding the next step. This iterative process supports graceful error handling because exceptions can be captured and returned as observations, allowing the LLM to adjust. Single-pass or fixed-chain approaches lack this adaptability and feedback.

Exam trap

The trap here is assuming that a single LLM call can both decide to use a tool and produce a final answer incorporating the tool's output, which is not possible without a loop.

265
MCQhard

Refer to the exhibit. Why did the analyst group lose access after the table was recreated?

A.The catalog owner needs to refresh the table metadata.
B.The analyst group needs to be re-added to the schema permissions.
C.The new table is a different object, so the GRANT statement must be repeated.
D.The user who recreated the table is not the catalog owner.
AnswerC

Unity Catalog treats the newly created table as a entirely new resource with a fresh identity. Any permissions applied to the previous version of the table do not carry over to the replacement, requiring administrators to re-issue GRANT commands for the new object.

Why this answer

In Unity Catalog, privileges are bound to the specific object ID. When a table is dropped and re-created, it becomes a new object with a new unique identifier. The previous GRANT statements, which were associated with the original object ID, do not automatically apply to the new table.

This mechanism protects against security drift by ensuring that every new object must have its permissions explicitly managed.

Exam trap

Candidates mistakenly believe that object permissions persist even if the underlying table is dropped and recreated, forgetting that Unity Catalog treats a new table as a distinct object ID.

266
MCQhard

Refer to the exhibit. The user is a member of the 'finance_team'. Why might the user encounter an access error when executing this join query?

A.The user lacks the SELECT privilege on the schema object.
B.The user lacks the USAGE privilege on the 'sales' schema.
C.The user requires the MODIFY privilege to perform joins.
D.The user needs to be an owner of the tables to perform joins.
AnswerB

Accessing any object in Unity Catalog requires the USAGE privilege on all containing objects, including the schema. Even if a user has explicit SELECT rights on the tables, the query will fail if they have not been granted USAGE on the parent schema container.

Why this answer

In Unity Catalog, users require the USAGE privilege on all parent objects—the catalog and the schema—to access a table. While the user has USAGE on the catalog and SELECT on the tables, they lack the USAGE privilege on the 'sales' schema. Without explicit USAGE permission on the schema, the user cannot traverse the hierarchy to reference the tables, causing a permission denied error.

Exam trap

Candidates often assume that having 'SELECT' on a table is enough to query it, overlooking the mandatory 'USAGE' permission required at the schema and catalog levels.

267
MCQmedium

A GenAI engineer has registered a RAG chain in Unity Catalog as a model and now needs to deploy it for real-time inference with per-request token usage and latency captured automatically. Which Databricks capability should they enable on the serving endpoint?

A.Attach an MLflow experiment to the endpoint and rely on the run history for telemetry.
B.Enable Unity Catalog lineage on the registered model version.
C.Configure the endpoint with autoscaling enabled to log per-request metrics.
D.Enable inference tables on the Mosaic AI Model Serving endpoint.
AnswerD

Inference tables on Mosaic AI Model Serving automatically capture the request payload, response, and metadata for each served request, enabling monitoring of token usage and latency without custom logging code. For a RAG chain registered in Unity Catalog, this gives immediate observability into production traffic and is the supported mechanism for capturing real-time inference telemetry.

Why this answer

Inference tables are the native Mosaic AI Model Serving feature that persists request and response payloads plus metadata to a Delta table, giving automatic capture of token usage and latency for real-time endpoints. Autoscaling, MLflow experiments, and Unity Catalog lineage address capacity, development tracking, and governance respectively, but none record per-request serving telemetry.

Exam trap

The trap here is assuming that any monitoring-adjacent feature, such as autoscaling or lineage, will automatically capture request-level telemetry for a served model.

268
Multi-Selecthard

Which TWO of the following are mandatory requirements for developing an AI application using the Databricks Mosaic AI Model Serving environment?

Select 2 answers
A.The model must be stored in a legacy Databricks Workspace folder.
B.The model must be registered as a model version within a Unity Catalog schema.
C.The model must be served using a shared-access mode interactive cluster.
D.The serving endpoint must be configured with a defined compute resource.
E.The application code must perform manual model sharding across nodes.
AnswersB, D

Unity Catalog acts as the central repository for model artifacts and versions in Databricks. Registering the model here provides the necessary metadata, lineage, and access control required by the serving infrastructure to deploy the model securely and ensure it remains reachable by authorized internal or external applications.

Why this answer

Mosaic AI Model Serving requires specific configurations to ensure secure and performant access. Firstly, the model must be registered in Unity Catalog to maintain governance and lineage. Secondly, the serving endpoint requires defined compute resources, typically GPU-accelerated for LLMs, to handle inference requests.

These requirements are essential for productionizing models, as they ensure that models are discoverable, governed, and have the necessary hardware to meet low-latency performance targets in real-time scenarios.

Exam trap

Candidates often focus on model training parameters or API authentication methods, missing the foundational requirements of Unity Catalog registration and defined compute resources for serving.

269
MCQmedium

An engineer is packaging a GenAI agent application with Databricks Asset Bundles so that the same bundle deploys to a development and a production workspace. The agent's serving endpoint name must differ per target, and the production endpoint needs more concurrent capacity. Which mechanism in the bundle configuration should the engineer use?

A.Use a post-deployment notebook that calls the serving endpoints API to rename the endpoint and adjust replicas after each bundle deployment.
B.Store the endpoint name and replica count as secrets in a Databricks secret scope and read them at deployment time with the Databricks CLI.
C.Define target-specific variables in databricks.yml and reference them with ${var.endpoint_name} and ${var.min_replicas} in the resource definition.
D.Hard-code the production endpoint name and replica count, then maintain a separate copy of the bundle for development.
AnswerC

Databricks Asset Bundles support variables declared at the bundle level and overridden per target, which is the supported way to vary values such as endpoint names and replica counts across development and production. Referencing them with the ${var.} syntax keeps one resource definition while letting each target supply different values, matching the requirement precisely.

Why this answer

Databricks Asset Bundles let you declare variables and override them per target, so a single resource definition can deploy the agent with different endpoint names and replica counts in development and production. Referencing variables with the ${var.} syntax keeps the bundle DRY while honoring environment differences. Hard-coded values, secret scopes, and post-deployment mutation scripts all fail to provide the clean, declarative per-target configuration the scenario requires.

Exam trap

The trap here is treating secrets or post-deploy scripts as the way to vary resource attributes, when bundle variables with per-target overrides are the purpose-built mechanism.

270
Multi-Selectmedium

Which TWO of the following are benefits of using Databricks Asset Bundles for deploying AI applications?

Select 2 answers
A.Enables automated testing through CI/CD pipeline integration
B.Provides a graphical drag-and-drop interface for deployment
C.Ensures environment consistency across development and production
D.Automatically handles data labeling and cleaning tasks
E.Allows users to bypass Unity Catalog governance
AnswersA, C

Because DABs are command-line driven, they integrate seamlessly with CI/CD tools like GitHub Actions or GitLab CI. This allows teams to automate unit tests, integration tests, and deployment steps, ensuring that only validated code is deployed to production, thereby significantly reducing the likelihood of runtime failures.

Why this answer

Databricks Asset Bundles standardize the deployment process, making it repeatable, version-controlled, and easier to integrate into CI/CD pipelines. This reduces the risk of configuration drift, where the production environment deviates from the development environment due to manual changes in the UI. By treating infrastructure as code, teams can maintain a clear history of changes and improve the reliability and auditability of their AI application releases.

Exam trap

Candidates often select options related to 'data transformation' or 'model training performance' instead of focusing on the DevOps-centric benefits of Bundles like consistency and CI/CD integration.

271
MCQeasy

Which of the following is a primary reason to prefer Databricks Asset Bundles (DABs) over manual workspace deployment?

A.DABs provide better performance for Spark queries.
B.DABs allow for version-controlled and automated deployments.
C.DABs automatically convert your code to SQL.
D.DABs require less memory to run than the standard workspace.
AnswerB

DABs enable infrastructure-as-code, meaning configurations are stored in version control (like Git). This allows for automated deployments via CI/CD, enabling reproducible environments, easier rollbacks, and a clear audit trail of all changes made to the infrastructure, which is a major improvement over manual, ad-hoc workspace changes.

Why this answer

The primary advantage of DABs is the elimination of manual, error-prone configuration. By defining infrastructure as code, teams gain repeatability, version control, and automation. This leads to significantly fewer deployment failures and allows for consistent, auditable environments across development and production.

It is a fundamental shift toward mature MLOps practices, where infrastructure is as controlled and tested as the application code itself.

Exam trap

Candidates tend to think DABs are just a specialized Python library or cluster type, forgetting their core value is automated, version-controlled infrastructure deployment.

272
MCQeasy

A developer is creating a Databricks notebook to prototype a GenAI application. They need to install the `databricks-langchain` library to use LangChain integrations with Databricks. Which command should they use in the notebook?

A.%conda install databricks-langchain
B.dbutils.library.installPyPI("databricks-langchain")
C.!pip install databricks-langchain
D.%pip install databricks-langchain
AnswerD

The `%pip install` magic command is the standard way to install Python packages in Databricks notebooks. It ensures the package is installed in the notebook's environment and available for import. This is the correct command to install the `databricks-langchain` library, which provides LangChain integrations for Databricks.

Why this answer

In Databricks notebooks, the `%pip install` magic command is the recommended way to install Python libraries. It ensures the package is installed in the notebook's isolated environment and is immediately available for import. Using `%pip` with `databricks-langchain` correctly sets up the LangChain integrations needed for the GenAI application.

Exam trap

The trap here is using outdated installation methods like `dbutils.library.installPyPI` or shell commands, which are not recommended in current Databricks runtimes.

273
Multi-Selectmedium

When designing a production-ready Databricks notebook for model inference, which TWO practices improve maintainability and performance?

Select 2 answers
A.Embed all logic including data preprocessing in a single, large notebook.
B.Use Databricks Widgets for parameterizing inputs like model paths.
C.Hardcode all file paths and configuration settings for consistency.
D.Refactor reusable code into libraries imported in the notebook.
E.Always run inference on the driver node to avoid network latency.
AnswersB, D

Widgets allow developers to pass parameters to notebooks at runtime, which is essential for making notebooks reusable. By parameterizing critical values like model paths, input table names, or configuration flags, developers can trigger the same notebook in different workflows without modifying the code, significantly increasing flexibility and maintainability in production.

Why this answer

Modular code design allows for easier testing and unit-level debugging, which is crucial for complex inference logic. Using parameterization enables the notebook to be reused across different environments and models without changing the core code. These practices reduce technical debt and simplify the lifecycle of ML models, making the transition from development to production much smoother and ensuring consistent performance across different stages of the CI/CD pipeline.

Exam trap

Test-takers mistakenly believe that hardcoding file paths or keeping all logic inside a single monolithic notebook improves execution speed, ignoring best practices for code maintainability.

274
MCQeasy

A team stores raw documents in a Unity Catalog volume and needs to build a training set for fine-tuning a large language model. The raw files are in mixed formats including PDF, DOCX, and plain text. The team wants a single Delta table where every row is one document with its extracted text and source path, and wants the extraction to run in parallel across the cluster. Which approach should the team use?

A.Mount the volume as a local filesystem path and use a Python for-loop to read and extract each file on the driver.
B.Use binaryFile data source to read the volume, then apply a UDF that extracts text per file and writes results to a Delta table.
C.Use spark.read.text on the volume, which automatically parses PDF and DOCX content into text columns.
D.Create an external table over the volume with a schema that includes columns for each document format's fields.
AnswerB

The binaryFile data source reads each file as a row with path and binary content, which lets Spark distribute file processing across executors. Applying an extraction UDF per row produces one output row per document with text and source path. Writing to Delta gives the unified table the team wants for fine-tuning.

Why this answer

The binaryFile data source is designed to read arbitrary files as rows with path and content, enabling distributed extraction with a UDF. This yields one row per document with text and source path in a Delta table. Driver loops and plain text readers cannot handle mixed binary formats at scale or in parallel.

Exam trap

The trap here is assuming a generic text reader or external table can parse PDF and DOCX, when binary formats require the binaryFile data source plus explicit extraction logic.

275
MCQeasy

What is the primary function of the 'Prompt' object in a Databricks RAG architecture?

A.To store the user's chat history for session persistence.
B.To format instructions and retrieved context for LLM consumption.
C.To index documents for faster similarity search.
D.To calculate the latency of the RAG pipeline.
AnswerB

The prompt is responsible for assembling retrieved documents and user queries into a format the model understands. This ensures that the context is properly delimited, the role of the model is defined, and the task is clearly stated, which is critical for generating high-quality, relevant, and accurate answers.

Why this answer

The prompt acts as the interface between the data retrieved from the vector index and the LLM's generative capability. It organizes the context, instructions, and user request into a coherent message structure. Proper prompt design is essential to focus the LLM's attention, reduce extraneous output, and ensure that the generative output remains grounded in the provided evidence.

Exam trap

Candidates think the prompt executes generation or runs vector searches, missing its actual role as a structured template organizing context and instructions for the LLM.

276
MCQeasy

A developer is using MLflow to track experiments for a RAG application. They want to log the retrieval step's parameters, such as the number of documents retrieved (k) and the embedding model used. Which MLflow API should they use?

A.mlflow.log_artifact()
B.mlflow.set_tag()
C.mlflow.log_param()
D.mlflow.log_metric()
AnswerC

mlflow.log_param() is used to log a single parameter (key-value pair) for a run. Parameters are typically configuration settings like k or model name. This is the correct API for logging retrieval parameters such as the number of documents and the embedding model, as they are scalar values that define the run's configuration.

Why this answer

MLflow parameters are meant for logging configuration settings that are constant for a run. The number of documents retrieved and the embedding model are such settings. mlflow.log_param() records them as key-value pairs, enabling easy comparison across runs in the MLflow UI. Metrics, artifacts, and tags serve different purposes and are not suitable for this use case.

Exam trap

The trap here is confusing parameters with metrics; parameters are configuration inputs, while metrics are output measurements that can vary during training.

277
MCQhard

A developer deploys a new model version as shown in the exhibit. What is the purpose of this configuration?

A.To increase compute capacity for the primary model.
B.To conduct A/B testing or a canary rollout.
C.To load balance between two different cloud regions.
D.To force all traffic to model-v1 when model-v2 errors.
AnswerB

Configuring traffic percentages allows for controlled, incremental rollouts, which are essential for A/B testing or canary deployments. This allows developers to validate the new model's performance on live traffic while minimizing the impact if the new model version exhibits unexpected behavior or suboptimal accuracy metrics.

Why this answer

This configuration implements a canary deployment strategy by splitting traffic between two model versions. By routing 90% of requests to the stable version and 10% to the new version, the team can monitor performance and accuracy in a real-world scenario with minimal risk. This is a standard practice for safely rolling out model improvements, allowing teams to collect data on the new version's behavior before a full-scale transition.

Exam trap

Candidates often misinterpret traffic splitting as a load balancing or scaling technique, rather than identifying its primary purpose: testing new models against live traffic to minimize risk.

278
Multi-Selectmedium

A company is preparing a generative AI application for production and must demonstrate that model inputs and outputs are traceable and that access to sensitive prompt data is controlled. Which TWO Unity Catalog capabilities should the team rely on to meet these governance objectives? (Choose two.)

Select 2 answers
A.Fine-grained privileges on catalogs, schemas, tables, and models
B.Embedding caching to reduce inference latency
C.Vector index partitioning by document topic
D.Audit logs that record which principals accessed governed tables and models
E.Automatic hyperparameter tuning of registered models
AnswersA, D

Unity Catalog's privilege model lets administrators grant narrowly scoped rights such as SELECT on a table or EXECUTE on a model to specific groups. This enforces least privilege on sensitive prompt data and model invocation, which is the second governance objective in the scenario and complements the audit trail.

Why this answer

Meeting the governance objectives requires both an enforcement mechanism and an evidence trail. Fine-grained privileges on the relevant catalogs, schemas, tables, and models enforce least-privilege access to sensitive prompt data, while Unity Catalog audit logs record which principals accessed those governed objects, providing the traceability that production compliance reviews demand.

Exam trap

The trap here is selecting performance features like caching or tuning as governance controls, when governance needs enforcement plus audit evidence.

279
MCQeasy

A team has deployed a RAG application on Databricks and wants to monitor the quality of responses in production. They have enabled inference table logging. Which built-in Databricks capability allows them to periodically evaluate the logged requests and responses for quality metrics like groundedness?

A.Databricks SQL dashboards that display the count of requests per hour.
B.Delta Live Tables expectations that validate response length.
C.Mosaic AI Agent Evaluation, which can be run on the inference table to compute quality metrics.
D.MLflow Tracking with custom metrics logged during model serving.
AnswerC

Mosaic AI Agent Evaluation is designed to evaluate agent and RAG applications. It can be run on logged inference tables to compute metrics such as groundedness and relevance. This allows continuous monitoring of production quality without manual intervention. It integrates with MLflow and Databricks workflows, making it the appropriate built-in capability for this scenario.

Why this answer

Mosaic AI Agent Evaluation is the built-in Databricks capability for evaluating RAG and agent applications. It can process inference tables and compute quality metrics such as groundedness, relevance, and answer correctness. This enables periodic, automated monitoring of production quality without custom code.

It is the correct tool for the described scenario.

Exam trap

The trap here is assuming that any monitoring tool, like SQL dashboards or Delta Live Tables expectations, can compute semantic quality metrics, when only Agent Evaluation provides LLM-specific judges.

280
MCQhard

Refer to the exhibit. What is the current permission state for the 'analyst_group' after the execution of the REVOKE statement?

A.The group retains access because they are still mentioned in the grant history.
B.The group loses access to the table.
C.The group retains access due to the order of operations in the metastore.
D.The group can still access the table if they are members of the manager_group.
AnswerB

The REVOKE command removes the SELECT privilege that was previously granted to the group. Unity Catalog follows the most recent explicit command for a privilege, so once the revoke command is successfully applied, the group's access rights are removed, and they can no longer query the table.

Why this answer

Unity Catalog employs a standard additive-subtractive permission model. When a user or group is explicitly denied or revoked a privilege, the revocation takes precedence. After the REVOKE command is executed, the SELECT privilege is removed from 'analyst_group'.

Even if the group was previously granted access, the REVOKE command effectively clears that privilege, ensuring the group can no longer access the specified table.

Exam trap

Candidates often assume that permissions are additive only or that a group might retain access through another role, ignoring that a specific REVOKE command overrides previous grants.

281
MCQmedium

When developing a feature engineering pipeline using Feature Store, which practice ensures maximum code reusability across training and inference?

A.Calculate features in the application code and pass them as raw inputs.
B.Hardcode feature calculation logic inside the training notebook.
C.Use the Feature Store client to log feature definitions and retrieve them.
D.Export features to a static CSV file after every training run.
AnswerC

Utilizing the Feature Store API allows developers to define features once and publish them to a Feature Table. This centralized repository acts as a single source of truth, allowing both training pipelines and online serving endpoints to fetch consistent, pre-computed feature values using the same lookup key.

Why this answer

Defining features as code using the Databricks Feature Store ensures that the same logic is applied during both batch training and real-time inference. By encapsulating feature calculations in Feature Tables, developers avoid the 'training-serving skew' where features are calculated differently in production. This practice is critical in GenAI development to ensure model performance consistency and reduce technical debt caused by disjointed preprocessing pipelines across different stages of the ML lifecycle.

Exam trap

Candidates often suggest using local Python functions or manual SQL scripts. These approaches do not track lineage or ensure consistency, leading to 'training-serving skew' in production.

282
Multi-Selectmedium

Which TWO actions are necessary to ensure that a model serving endpoint in Databricks remains available and performant during peak traffic hours?

Select 2 answers
A.Disable the serving endpoint's logging to save on compute cycles.
B.Enable auto-scaling for the serving endpoint.
C.Configure the endpoint to run on a single, fixed-size node to minimize latency.
D.Select appropriate GPU-accelerated compute types for the workload.
E.Manually restart the endpoint every hour to clear the cache.
AnswersB, D

Auto-scaling allows the serving endpoint to dynamically adjust its instance count based on current request volume. This ensures that the system handles spikes in traffic effectively without manual intervention, maintaining consistent performance and avoiding service degradation when user demand increases suddenly during peak hours or application usage cycles.

Why this answer

Maintaining performance requires both vertical and horizontal scaling strategies. First, configuring auto-scaling on the endpoint allows the infrastructure to adjust to fluctuating demand automatically. Second, ensuring the endpoint uses appropriate instance types—such as GPU-accelerated instances for LLMs—ensures the compute capacity is sufficient for inference tasks.

These two actions are foundational for achieving high availability and low latency, preventing request timeouts and ensuring a smooth experience during heavy usage periods.

Exam trap

Students often select only one correct action or mistakenly choose manual scaling scripts, forgetting that both auto-scaling and GPU acceleration are required for high performance.

283
MCQhard

A team is preparing a Delta table of product descriptions for a RAG application. The table receives continuous upserts from a streaming pipeline, and the embedding job reads the table every hour. Engineers notice the embedding job reprocesses every row on each run even though only a few rows change. Which change should be made to the source table to let the embedding job process only new or updated rows?

A.Partition the Delta table by the product category column so the embedding job reads fewer files.
B.Enable Change Data Feed on the Delta table and have the embedding job read the change feed since the last processed version.
C.Convert the table to a view that filters rows by the current timestamp so only recent rows appear.
D.Run OPTIMIZE with Z-ORDER on the product identifier column to compact small files before each embedding run.
AnswerB

Change Data Feed records row-level inserts, updates, and deletes with commit versions, so the embedding job can read only changes since its last checkpoint. This avoids reprocessing unchanged rows on every run. It is the intended mechanism for incremental downstream consumption of a Delta table that receives continuous upserts.

Why this answer

Change Data Feed captures row-level changes with commit versions, allowing a downstream job to read only the inserts, updates, and deletes that occurred since its last checkpoint. This directly solves the problem of reprocessing unchanged rows. Partitioning, compaction, and timestamp filtering improve read characteristics but do not expose which rows changed, so they cannot deliver incremental processing.

Exam trap

The trap here is confusing read-performance optimizations such as partitioning or OPTIMIZE with change tracking, when only Change Data Feed exposes which rows were inserted, updated, or deleted.

284
Multi-Selecthard

Which THREE components are critical to include in a comprehensive evaluation strategy for a RAG-based Generative AI application?

Select 3 answers
A.Retrieval metrics (e.g., Hit Rate, MRR)
B.Generation metrics (e.g., Faithfulness, Answer Relevance)
C.End-to-end performance metrics (e.g., Answer Correctness)
D.Hardware temperature monitoring for GPUs
E.Network packet loss percentage
AnswersA, B, C

Retrieval metrics are essential because they confirm whether the system is successfully finding the correct documents. If retrieval fails, no amount of prompt engineering or fine-tuning will lead the generator to provide a correct, factual answer.

Why this answer

A robust RAG evaluation requires a holistic view that covers the entire pipeline. Evaluating just the generation or just the retrieval is insufficient because the quality of the final response is dependent on both stages. By measuring retrieval performance, generation faithfulness, and end-to-end answer quality, engineers obtain a full picture of where the application is succeeding and where it needs tuning to meet user requirements.

285
MCQeasy

A GenAI engineer has written a notebook that calls the Databricks Foundation Model APIs to summarize documents. Before the notebook can run in production, the security team wants to confirm exactly which workspace users and service principals are permitted to invoke the pay-per-token foundation model endpoints. Where should the engineer point them to review and manage those permissions?

A.The Unity Catalog metastore's default storage location settings
B.The Unity Catalog privileges on the system.ai schema and its model functions
C.The workspace admin settings page for cluster policies
D.The Azure Databricks access connector assigned to the workspace
AnswerB

Databricks exposes foundation models as Unity Catalog functions under the system.ai schema, and invocation is governed by Unity Catalog privileges such as EXECUTE on each model function. Reviewing and granting these privileges shows precisely which users and service principals may call the pay-per-token endpoints, which is exactly what the security team requested.

Why this answer

Foundation Model APIs are surfaced as Unity Catalog functions in the system.ai schema, so access is controlled with Unity Catalog privileges, chiefly EXECUTE on the relevant model function, plus the usual USE CATALOG and USE SCHEMA on system.ai. Directing the security team to those grants gives an auditable, centralized view of exactly which users and service principals can invoke each pay-per-token model.

Exam trap

The trap here is assuming foundation model access is configured through workspace-level settings or compute policies rather than Unity Catalog privileges on the system.ai model functions.

286
MCQeasy

Which Databricks feature is specifically designed to facilitate the rapid development and deployment of LLM applications by providing a managed environment for hosting and testing prompts?

A.Databricks SQL Warehouse
B.Databricks Mosaic AI Playground
C.Unity Catalog Volumes
D.Delta Live Tables
AnswerB

Mosaic AI Playground is the dedicated UI for interacting with models, testing prompts, and comparing responses in real-time. It is the primary tool for early-stage development, allowing developers to see how different model configurations and prompts affect output quality without writing code, effectively accelerating the initial application development cycle.

Why this answer

Databricks Mosaic AI Playground provides a low-code interface for developers to experiment with different LLMs and system prompts. It allows for quick iteration and testing of model behavior before moving into full production deployment. This environment is crucial because it bridges the gap between experimentation and application development, ensuring that developers can validate prompt engineering strategies within the same governed environment where their data and production models reside.

Exam trap

Test-takers frequently guess generic cloud tools or external third-party playgrounds instead of the native Databricks-specific environment explicitly designed for rapid prompt engineering and LLM testing.

287
MCQeasy

A Generative AI engineer is deploying a new version of a RAG chain to a Mosaic AI Model Serving endpoint. Before promoting it to production, they want to run an evaluation that checks whether the generated answers are faithful to the retrieved documents. Which Mosaic AI Agent Evaluation metric should they examine?

A.chunk_relevance
B.latency
C.request_count
D.groundedness
AnswerD

groundedness measures whether the generated answer is supported by the retrieved context, directly assessing faithfulness to the documents. This is exactly the quality dimension the engineer needs to verify before promotion. It helps catch hallucinations where the model adds information not present in the retrieved chunks.

Why this answer

To verify that generated answers are faithful to the retrieved documents, the evaluation must compare the answer against the context. groundedness is the Agent Evaluation metric that performs this check, flagging responses that include unsupported claims. It is the appropriate pre-promotion quality gate for a RAG chain.

Exam trap

The trap here is confusing retrieval-stage metrics like chunk_relevance with answer-stage metrics, when faithfulness of the final answer is measured by groundedness.

288
MCQmedium

An engineer is designing a RAG application that uses Databricks Vector Search to retrieve documents and a foundation model endpoint to generate answers. The team wants to log all user queries, retrieved documents, and generated responses for auditing and continuous improvement. They also need to monitor for drift in retrieval quality over time. Which Databricks capability should they integrate into the application design?

A.Delta Live Tables with expectations to validate the retrieved documents.
B.Databricks SQL dashboards querying the source Delta table for document updates.
C.Unity Catalog audit logs and lineage for the Vector Search index.
D.MLflow Tracing with inference tables on the model serving endpoint.
AnswerD

MLflow Tracing captures detailed spans for each step of a generative AI application, including retrieval and LLM calls, enabling end-to-end observability. Inference tables on the serving endpoint automatically log request and response payloads for auditing. Together they provide the query, retrieved documents, and generated response logging, plus monitoring for drift. This is the recommended Databricks approach for tracing and logging generative AI applications.

Why this answer

MLflow Tracing provides detailed spans across the RAG pipeline, capturing queries, retrieved documents, and LLM responses. Inference tables on the serving endpoint automatically log request and response payloads for auditing. This combination enables both observability and drift monitoring.

Other options focus on data pipelines, dashboards, or governance, which do not capture the runtime application behavior needed for continuous improvement.

Exam trap

The trap here is confusing governance auditing (Unity Catalog) or data pipeline quality (Delta Live Tables) with application-level tracing and payload logging, which are needed to monitor retrieval quality and improve the RAG application.

289
MCQhard

An engineer uses MLflow LLM Evaluation with mlflow.evaluate() to score a RAG application. The judge model is configured with a temperature of 0.9 and no explicit metric thresholds are set. Reruns of the identical evaluation dataset produce relevance scores that swing by up to 20 percentage points, and the team cannot tell whether a prompt change helped. Which change most directly improves the reliability of the evaluation comparison?

A.Increase the evaluation dataset size by duplicating each question three times so variance averages out across more judge calls.
B.Set the judge model's temperature to 0 and define fixed pass/fail thresholds for each metric before comparing runs.
C.Switch the evaluation to a larger judge model with more parameters while keeping the existing sampling settings.
D.Log only the aggregate mean relevance score per run and discard the per-row judgments from the MLflow run artifacts.
AnswerB

The instability originates from stochastic sampling in the judge, and temperature 0 makes the judge's verdicts deterministic for a given input, removing run-to-run score drift. Pairing that with predefined metric thresholds converts continuous, noisy scores into stable pass/fail decisions, so a prompt change can be attributed to real quality movement instead of sampling noise. This combination directly addresses reproducibility and comparability of evaluation runs.

Why this answer

Score swings across identical runs are caused by nondeterministic judge sampling, so the fix must remove that randomness and make the decision rule stable. Setting the judge temperature to zero makes each judgment reproducible, and predefined thresholds turn noisy continuous scores into consistent pass/fail outcomes. Enlarging the judge, duplicating rows, or hiding per-row results leaves the stochasticity in place and keeps run-to-run comparisons unreliable.

Exam trap

The trap here is treating evaluation variance as a dataset-size problem when the root cause is a high-temperature judge producing different verdicts on identical inputs.

290
MCQmedium

You are building an application that uses Model Serving to host a fine-tuned LLM. Which configuration is required to optimize for high-concurrency request throughput?

A.Set the serving endpoint to use CPU-only instances for all LLM inference tasks.
B.Enable 'Provisioned Throughput' on the serving endpoint configuration.
C.Disable auto-scaling and fix the number of replicas to one.
D.Configure the endpoint to use an external API key for every request.
AnswerB

Provisioned Throughput provides dedicated capacity for foundation models, which is essential for high-concurrency scenarios. It ensures that the model has the necessary resources to handle concurrent requests without performance degradation, making it the correct choice for scaling LLM applications beyond simple development or low-traffic testing environments.

Why this answer

Enabling Provisioned Throughput for Model Serving allows for dedicated resources that can handle high-concurrency workloads efficiently. This is crucial for applications serving many users simultaneously. By optimizing resource allocation, you prevent bottlenecks during peak usage, ensuring consistent performance and minimizing request latency, which directly impacts the user experience and the scalability of the overall application architecture.

Exam trap

Candidates often select standard auto-scaling for high-concurrency needs, failing to realize that Provisioned Throughput is specifically required to guarantee performance for high-load production LLM workloads.

291
MCQmedium

Refer to the exhibit. An engineer observes an unexpected drop in "relevance" for the latest deployment. What is the most likely cause related to the evaluation process itself?

A.The model was trained on too much data.
B.The evaluation dataset is misaligned with current production query trends.
C.The inference cluster has too much compute capacity.
D.The faithfulness score increased too much.
AnswerB

If production traffic evolves but the evaluation set remains static, the metrics will reflect outdated user needs. A drop in relevance usually signifies that the model is no longer performing well on the types of questions users are currently asking.

Why this answer

A drop in relevance often suggests that the evaluation dataset (eval_set_v4) may have become outdated or misaligned with the current production use cases. If the user base or query patterns shift, the original evaluation set may no longer accurately reflect the actual performance requirements. Regularly refreshing the evaluation dataset to match current production trends is essential to ensure that metrics remain valid indicators of system performance.

Exam trap

Candidates often blame model degradation first, overlooking that an outdated or misaligned evaluation dataset itself can cause an unexpected drop in metric scores.

292
MCQhard

A healthcare company uses Databricks to build a RAG application over clinical notes stored in a Unity Catalog table. An auditor requires proof that only authorized personnel can view raw note text, while the RAG application must use embeddings generated from those notes. The team wants to avoid copying data outside Unity Catalog. Which approach best satisfies the auditor while preserving RAG functionality?

A.Move the clinical notes to an external location outside Unity Catalog and grant access via cloud IAM roles.
B.Create a column mask on the note text column that returns a redacted value for unauthorized users, and grant the RAG service principal access to a separate embeddings table derived from the notes.
C.Grant the RAG application's service principal SELECT on the table and rely on the application to redact note text in its responses.
D.Use a dynamic view that filters rows based on the user's group and grant the RAG service principal access to that view.
AnswerB

A column mask on the note text column ensures that only authorized users see raw PHI, satisfying the auditor. The RAG application can read embeddings from a derived table without needing raw text, so it continues to function. Unity Catalog enforces the mask at query time based on identity, providing auditable, centralized control.

Why this answer

A column mask on the sensitive note text column enforces redaction for unauthorized identities at query time, which is auditable. By having the RAG application consume a derived embeddings table instead of raw text, the application retains functionality while raw PHI remains protected. This keeps governance within Unity Catalog and satisfies the auditor's requirement.

Exam trap

The trap here is thinking that application-level redaction or IAM roles can substitute for Unity Catalog column masks when an auditor demands provable data-layer controls.

293
MCQmedium

An engineer is cleaning a Delta table of customer support transcripts stored in Unity Catalog. The table has a nested array column named messages, where each element has role and content fields. They need to keep only rows where at least one message has role equal to 'user' and content longer than 20 characters. Which transformation correctly expresses this filter?

A.Cast the messages array to a string and apply a LIKE pattern matching the literal text 'user' followed by any content.
B.Use the higher-order function exists on the messages array with a lambda testing role = 'user' AND length(content) > 20.
C.Use the transform higher-order function on messages and then check whether the resulting array is non-empty.
D.Explode the messages array, filter on role and content length, then group by the row identifier to rebuild the rows.
AnswerB

The exists higher-order function returns true when any array element satisfies the lambda predicate, which matches the requirement of at least one qualifying message. Applying it inside a filter keeps rows where the condition holds, and lambda access to role and content works directly on the nested struct elements without exploding the array.

Why this answer

Filtering rows based on whether any nested element meets a predicate is exactly what the exists higher-order function provides, and its lambda can test both role and content length in one pass. Exploding reshapes and shuffles data, string matching is ambiguous, and transform preserves array length so it cannot signal absence of a match.

Exam trap

The trap here is reaching for explode-and-regroup out of habit, when a higher-order array function expresses the same predicate without changing the table shape.

294
MCQmedium

When preparing a dataset for fine-tuning an LLM, you need to ensure the data is representative of the target domain. What is the most effective approach to detect and mitigate sampling bias in your training set using Databricks?

A.Increase the batch size during the fine-tuning process.
B.Perform statistical profiling on feature distributions and re-sample if necessary.
C.Change the model architecture to a larger parameter count.
D.Use a random seed to shuffle the data before splitting.
AnswerB

Statistical profiling reveals imbalances, allowing engineers to apply techniques like oversampling or undersampling to correct them. This quantitative approach is objective and standard for ensuring high-quality machine learning training sets. It prevents the model from favoring specific patterns simply because they were more frequent in the dataset.

Why this answer

Using exploratory data analysis (EDA) with Databricks SQL or Spark to analyze distribution statistics is the most effective way to identify bias. By comparing the distribution of the training set to a representative sample of real-world production data, engineers can identify under-represented categories. This ensures that the fine-tuned model performs reliably across all expected input scenarios, preventing the model from developing blind spots.

Exam trap

Candidates tend to focus exclusively on model hyperparameter tuning while ignoring raw dataset distributions, missing structural imbalances present in the training inputs.

295
MCQmedium

A developer is building a RAG application using Mosaic AI Model Serving. They need to ensure that the embedding model endpoint is strictly accessed only by specific service principals within the workspace. Which feature should the developer configure to enforce this security requirement?

A.Network Security Groups
B.Unity Catalog External Locations
C.Model Serving Permissions
D.Cluster Access Control Lists
AnswerC

Model Serving Permissions allow administrators to define precise access control lists for each inference endpoint. By assigning 'Can query' privileges only to the required service principals, the developer effectively secures the endpoint, ensuring that unauthorized entities cannot perform inference operations against the deployed embedding model in the production environment.

Why this answer

To restrict access to Mosaic AI Model Serving endpoints, developers should use Model Serving Permissions. By navigating to the Permissions tab of the specific endpoint, they can grant 'Can query' access exclusively to authorized service principals. This ensures that only authenticated and authorized applications can retrieve embeddings, protecting the model from unauthorized inference calls while maintaining a secure development lifecycle in Databricks.

Exam trap

Candidates often confuse workspace-level access controls or IAM roles with endpoint-specific permissions. They mistakenly select broad settings like 'Workspace Admin' or 'Can View' instead of the specific 'Can query' permission required for inference.

296
MCQmedium

A developer is building a retrieval-augmented generation (RAG) application on Databricks. They need to ensure that embeddings are updated automatically when the underlying Delta table changes. Which approach is the most efficient and scalable?

A.Write a manual PySpark job that runs every hour to scan the entire Delta table and recompute all embeddings.
B.Use a Databricks Job with a notebook task triggered by a cron schedule to process changes using a watermark.
C.Implement a Delta Live Tables pipeline using streaming tables and a Python UDF to generate embeddings on arrival.
D.Deploy a Unity Catalog volume to store the embeddings and trigger an external API call from an event-driven function.
AnswerC

Delta Live Tables streaming tables automatically handle incremental data ingestion and state management. By applying a transformation function during the stream, the pipeline computes embeddings only for new or updated records, significantly reducing compute overhead and ensuring the vector index is always current without manual maintenance or scheduling logic.

Why this answer

Delta Live Tables (DLT) with streaming tables allows for continuous data processing and incremental updates. By integrating embedding generation directly into the pipeline, the developer ensures that the vector database stays in sync with the source data without manual intervention or complex scheduling. This architecture minimizes latency and improves data consistency, which is a foundational requirement for production-grade RAG applications within the Databricks ecosystem.

Exam trap

Candidates often suggest manual batch jobs or triggers, which are less efficient and harder to scale than Delta Live Tables' native streaming capabilities for continuous data processing.

297
MCQhard

A team is using MLflow LLM Evaluation with the built-in answer_correctness metric to compare two prompt templates for a question-answering application. They notice that answer_correctness scores are nearly identical, but manual review shows one template produces answers that are factually correct yet omit key supporting details. Which additional built-in metric should they add to their evaluation to surface this difference?

A.groundedness
B.answer_completeness
C.toxicity
D.answer_similarity
AnswerB

answer_completeness evaluates whether the response covers all the key information present in the ground truth answer. Since the issue is that one template omits supporting details while remaining factually correct, this metric directly measures that gap and will differentiate the templates. It is the built-in metric designed to assess thoroughness against the reference.

Why this answer

When answers are factually correct but differ in how much supporting detail they include, a completeness metric is required. answer_completeness compares the response against the ground truth to check whether all key points are covered, making it the right built-in metric to expose the omission of supporting details that answer_correctness alone misses.

Exam trap

The trap here is treating factual correctness and completeness as the same thing, when a response can be correct yet incomplete and needs a dedicated completeness metric to detect it.

298
Multi-Selecthard

A team is preparing a large text corpus for embedding generation with a foundation model endpoint on Databricks. They must reduce token cost and improve retrieval quality before vectorization. Which two preprocessing steps should be applied to the raw text? (Choose two.)

Select 2 answers
A.Remove stop words and punctuation from every passage before generating embeddings.
B.Deduplicate near-identical passages and drop documents that fail a minimum content-length threshold.
C.Truncate every document to its first 512 characters to guarantee uniform chunk size.
D.Normalize whitespace and strip boilerplate such as navigation menus, headers, and repeated legal footers.
E.Convert all text to uppercase before tokenization to standardize casing.
AnswersB, D

Near-duplicate passages waste embedding calls and crowd the index with redundant vectors, so removing them lowers cost and improves result diversity. Dropping documents below a minimum length eliminates fragments and stubs that produce low-information embeddings, which raises overall retrieval precision without discarding meaningful content.

Why this answer

Removing boilerplate and normalizing whitespace cuts wasted tokens while sharpening the text signal, and deduplicating plus dropping undersized fragments avoids paying to embed redundant or low-information content. Uppercasing, stop-word stripping, and blind truncation either damage semantics or discard content, so they do not improve cost or retrieval quality.

Exam trap

The trap here is assuming aggressive token reduction like stop-word removal or truncation always lowers cost, when it can degrade embedding quality without meaningful savings.

299
MCQmedium

An AI engineer is developing a custom Databricks App using Mosaic AI Agent Framework and needs to deploy the application workspace securely. Which deployment artifact and configuration mechanism should the engineer use to define the app dependencies and entry point?

A.A setup.py file placed in the root directory configured to automatically trigger a cluster restart upon deployment.
B.An app.yaml configuration file defining the command, environment variables, and source path alongside the application source code.
C.A Databricks Asset Bundles deployment target pointing to an external Kubernetes cluster hosting the agent code.
D.A standard requirements.txt file executed through a Jupyter notebook scheduled via a workflow job.
AnswerB

The app.yaml file serves as the core manifest for Databricks Apps, enabling developers to specify entry-point commands, required Python versions, and runtime parameters. This file allows the Databricks platform to provision compute and host the application securely within the workspace.

Why this answer

Deploying Databricks Apps requires an app.yaml configuration file and source code bundled together. The app.yaml file specifies the runtime, environment variables, and command-line entry points required to execute the application properly inside the Databricks environment. Defining these parameters correctly ensures the Mosaic AI Agent dependencies are initialized and served without manual intervention.

Exam trap

Candidates often confuse Databricks App deployment with standard notebook deployment, forgetting that an 'app.yaml' file is strictly required to define the entry point and runtime environment.

300
MCQmedium

A developer is building a RAG application and notices that the retrieval step often returns irrelevant context. Which step in the pipeline should be improved to address this?

A.Increase the number of API calls to the LLM.
B.Refine the chunking strategy and embedding model quality.
C.Enable auto-scaling on the serving endpoint.
D.Switch to a larger, more expensive LLM.
AnswerB

The quality of retrieval is heavily dependent on how the source data is chunked and represented as vectors. Refining these parameters—such as using smaller, more meaningful segments or a higher-quality embedding model—is the standard approach to ensuring that the search engine finds the most relevant information for any given query.

Why this answer

Improving retrieval accuracy often involves enhancing the quality of the embeddings or the text chunking strategy. By adjusting how data is partitioned into chunks (e.g., overlapping chunks, semantic chunking) or refining the embedding model, the vector search index can better capture the semantic meaning of the documents. This is a common and critical improvement path in RAG development, as the quality of the retrieval is fundamentally limited by the index's representational capability.

Exam trap

Candidates often jump to increasing the number of retrieved chunks or changing the LLM model. These address symptoms rather than the root cause of poor semantic retrieval quality.

Page 3

Page 4 of 5

Page 5

All pages