Courseiva

Databricks Certified Generative AI Engineer Associate (Databricks-GenAI-Assoc) — Questions 76–150

330 questions total · 5pages · All types, answers revealed

Page 1

Page 2 of 5

Page 3
76
MCQhard

A generative AI engineer is designing a multi-stage RAG application on Databricks. The application first retrieves documents using Vector Search, then reranks them with a cross-encoder model, and finally calls a foundation model endpoint to generate an answer. The engineer wants to ensure that the entire pipeline is reproducible and that each stage can be independently versioned and deployed. Which design approach best meets these requirements?

A.Package each stage as an MLflow model with its own signature and dependencies, and compose them into a single MLflow pipeline that can be logged and served as one model.
B.Implement the pipeline as a single Python script that calls all stages sequentially, and version the script in Git.
C.Use MLflow Projects to define each stage as a separate project, and orchestrate them with a Databricks Job that passes artifacts between stages.
D.Deploy each stage as a separate Databricks Model Serving endpoint and have the client application call them in sequence.
AnswerA

MLflow models encapsulate code, environment, and signatures, enabling independent versioning. Composing them into a pipeline allows the entire multi-stage application to be logged as a single model, which can be served with Databricks Model Serving. This provides reproducibility and independent versioning of each stage while offering a unified deployment artifact. It is the recommended pattern for complex generative AI pipelines.

Why this answer

Packaging each stage as an MLflow model with its own signature and dependencies allows independent versioning. Composing them into a single MLflow pipeline enables logging and serving the entire multi-stage RAG application as one model, ensuring reproducibility. This approach is supported by Databricks Model Serving and MLflow, and it simplifies deployment and tracing compared to separate endpoints or scripts.

Exam trap

The trap here is assuming that separate endpoints or a single script provide sufficient versioning and reproducibility, when the key is to encapsulate each stage as an MLflow model and compose them into a unified pipeline.

77
MCQmedium

A GenAI engineer is designing a multi-turn chat application on Databricks. Users report that the assistant forgets details from earlier in long conversations and sometimes answers using only the most recent message. The team wants the model to reliably use facts stated several turns earlier without exceeding the model's context window. Which design should the engineer implement?

A.Implement a conversation memory strategy that summarizes or selectively retains earlier turns and injects the compressed history into each prompt within the context window.
B.Store the full conversation transcript in a Delta table and instruct the model to query the table during inference.
C.Increase the model's temperature so it explores more of the conversation history when generating a response.
D.Send only the latest user message to the model and rely on the model's pretrained knowledge to recall earlier facts.
AnswerA

A memory strategy that summarizes or selectively retains earlier turns keeps essential facts available while staying within the context window. By injecting the compressed history into each prompt, the model receives the earlier details it needs, which directly addresses the reported forgetting and avoids the cost and truncation problems of sending the full transcript.

Why this answer

Reliable multi-turn memory requires bringing earlier relevant turns into the prompt, either by summarizing them or by selectively retaining key facts. This keeps the payload within the context window while preserving details the user expects the assistant to remember. Temperature changes, sending only the latest message, or storing transcripts without retrieval all fail to put earlier facts in front of the model.

Exam trap

The trap here is assuming the model has inherent memory across API calls, when in fact each inference request is stateless and only sees the prompt it is given.

78
Multi-Selecthard

A GenAI engineer is designing a Databricks application that must ground answers in a large corpus of internal policy documents. The corpus is updated by a nightly Delta job, and the application must cite the source document for every answer. The engineer is deciding how to structure the retrieval and generation stages. Which TWO design choices best satisfy the grounding and citation requirements? (Choose two.)

Select 2 answers
A.Instruct the model to answer only from the retrieved context and to include the provided document identifiers in its response
B.Rely on the model's pretrained knowledge of internal policies to fill gaps when retrieval returns no relevant chunks
C.Store chunk text and source metadata such as document ID and section as columns alongside the embeddings in the Vector Search index
D.Cache the full corpus in the prompt on every request so the model always has complete policy context
E.Increase the model's temperature setting so the generation stage explores a wider range of possible answers
AnswersA, C

Constraining the prompt to the retrieved context reduces unsupported generation, and requiring the model to echo the supplied identifiers ties each statement back to a retrieved chunk. Because the identifiers originate from the index metadata rather than the model's memory, the citations remain verifiable against the source documents.

Why this answer

Grounding and citation depend on two things working together: retrieval that returns source identifiers with each chunk, and a generation prompt that restricts the model to that retrieved context while requiring it to surface the identifiers. Metadata columns supply verifiable provenance, and the constrained instruction keeps the answer tied to retrieved policy text instead of the model's own memory.

Exam trap

The trap here is believing a capable model can supply citations from memory, when only metadata carried through retrieval can produce citations that are verifiable against the source documents.

79
MCQhard

A developer is building an agentic workflow using Databricks and LangChain. The agent needs to decide whether to answer a user's query directly or call an external tool to retrieve additional information. The developer wants to ensure the agent's decisions are logged for debugging and auditing. Which approach should they take to achieve this?

A.Use MLflow's autologging for LangChain to automatically capture agent steps and tool calls.
B.Use Databricks Jobs to schedule the agent and rely on job run history for debugging.
C.Manually log each tool call using print statements to the driver logs.
D.Implement a custom logging function that writes agent decisions to a Delta table.
AnswerA

MLflow provides autologging for LangChain that automatically logs agent runs, including chain steps, tool invocations, and inputs/outputs. This creates a trace in MLflow, which can be viewed in the UI for debugging and auditing. It requires minimal code changes and is the recommended way to log agentic workflows on Databricks.

Why this answer

MLflow autologging for LangChain is the most efficient and integrated way to capture agent steps. It automatically logs tool calls, chain sequences, and inputs/outputs as MLflow traces, which can be inspected in the MLflow UI. This provides auditability and simplifies debugging without manual instrumentation.

Exam trap

The trap here is assuming that basic logging or job history is sufficient, but agentic workflows require detailed tracing of internal decisions, which MLflow autologging provides out of the box.

80
MCQmedium

A GenAI engineer is building a chatbot using Databricks Foundation Model APIs. They need the chatbot to maintain conversation context across multiple user turns and also allow the use of custom tools like a weather API. Which approach should they take?

A.Use the chat completion endpoint with a messages array that includes the full conversation history and define tools in the request.
B.Use the chat completion endpoint but only send the latest user message, relying on the model to remember previous turns.
C.Use the completions endpoint and concatenate all previous user inputs and model outputs into a single prompt string.
D.Use the embeddings endpoint to encode the conversation history and pass the embeddings to the model for context.
AnswerA

The chat completion endpoint accepts a messages array containing prior user and assistant messages, which provides conversational context. You can also specify tools (functions) in the request, and the model can return tool calls. This directly supports multi-turn context and custom tool integration, meeting both requirements.

Why this answer

The chat completion endpoint with a messages array is designed for multi-turn conversations and supports tool definitions. By including the full history, the model can reference prior context, and by defining tools, the model can request external API calls. This is the standard pattern for building conversational agents with Databricks Foundation Model APIs.

Exam trap

The trap here is assuming that the model retains memory between API calls, when in fact each request is stateless and must include the entire conversation history.

81
MCQhard

A generative AI engineer trains a model on a Delta table that contains customer support transcripts. Before registering the model, security requires that the training data be classified so that policies can be applied consistently across the lakehouse. The engineer wants to attach a governed label to the transcript column indicating it contains sensitive personal data. Which Unity Catalog feature should be used?

A.A governed tag applied to the column
B.A storage credential scoped to the transcript path
C.A comment added to the column definition
D.A row filter that excludes rows containing personal data
AnswerA

Governed tags in Unity Catalog attach classification metadata to securable objects, including individual columns, and can be used to drive policy decisions and discovery. Tagging the transcript column as sensitive personal data gives security a consistent label that downstream policies and audits can reference, which is exactly the classification requirement described.

Why this answer

Governed tags provide a structured, searchable classification layer in Unity Catalog that can be attached to columns and referenced by policy and discovery tooling. Applying a sensitivity tag to the transcript column gives security a consistent label for the training data, enabling uniform policy enforcement and auditability rather than relying on informal comments or access mechanisms.

Exam trap

The trap here is treating descriptive comments or access grants as classification, when only governed tags provide enforceable labels.

82
MCQeasy

A Generative AI engineer at a retail company has deployed a RAG chatbot on a Databricks Mosaic AI Model Serving endpoint. The team wants to capture per-request evaluation data, including the user's question, the retrieved context chunks, and the model's response, so they can analyze response quality over time. Which Databricks feature should they use to collect this data for downstream evaluation?

A.MLflow Model Registry
B.Delta Live Tables expectations
C.Inference tables
D.Unity Catalog lineage
AnswerC

Inference tables automatically log the request payload and the response payload for each call to a Mosaic AI Model Serving endpoint, which includes the user's question, retrieved context, and generated answer when the endpoint serves a RAG chain. This gives the team the raw per-request data needed to run evaluation and monitoring jobs later, without changing the client application.

Why this answer

Inference tables are the Databricks mechanism that logs request and response payloads for Mosaic AI Model Serving endpoints. Because the logged payloads include the user question, retrieved context, and generated answer, the team gains the raw material needed to run MLflow LLM Evaluation or custom monitoring jobs later against real production traffic.

Exam trap

The trap here is assuming that any Databricks governance or tracking feature automatically captures serving payloads, when only inference tables log request and response content for an endpoint.

83
Multi-Selectmedium

A team is preparing a Databricks Asset Bundle that deploys a GenAI application consisting of a job that builds a vector index and a Model Serving endpoint that hosts the agent. Before merging, they want the pipeline to validate and deploy the bundle to a staging workspace automatically. Which TWO bundle capabilities should the pipeline rely on? (Choose two.)

Select 2 answers
A.Run databricks bundle deploy with the staging target to create or update the job, endpoint, and related resources in the staging workspace.
B.Run databricks bundle schema to regenerate the JSON schema and commit it to the repository.
C.Run databricks bundle run to execute the vector index job as part of the merge validation.
D.Run databricks bundle generate to create resource definitions from existing workspace objects before each deployment.
E.Run databricks bundle validate to check the bundle configuration against the target workspace before deployment.
AnswersA, E

bundle deploy applies the resolved configuration to the selected target, creating or updating the declared resources and recording the deployment state. Pointing it at the staging target is how the pipeline promotes the bundle into that workspace.

Why this answer

A bundle pipeline validates the configuration against the target workspace and then deploys it to that target. Validation catches malformed or unresolvable configuration before anything changes, and deployment applies the resolved resources to the staging workspace, which together form the automated promotion path the team wants.

Exam trap

The trap here is confusing bundle run, which executes deployed resources, with bundle deploy, which publishes the bundle definition to a target workspace.

84
MCQmedium

A team runs a RAG chatbot whose MLflow evaluation with the built-in groundedness judge previously scored well. After they swap the retriever for a new embedding model, groundedness scores drop sharply even though the generator model and prompt are unchanged. They confirm the judge model itself is unchanged. Which action should they take FIRST to diagnose the regression?

A.Roll back the embedding model change immediately and re-run the evaluation to confirm the previous score returns.
B.Inspect the retrieved context chunks logged per evaluation row to verify whether the new retriever is returning passages that no longer support the generated claims.
C.Retrain the judge model on a fresh labeled dataset so its scoring distribution aligns with the new embedding model.
D.Increase the temperature of the generator model to encourage more diverse phrasing that the groundedness judge can score more reliably.
AnswerB

Groundedness judges whether each claim in the response is supported by the retrieved context. If the retriever changed, the logged retrieved chunks are the first artifact to inspect, since they directly feed the judge. Reviewing per-row traces confirms whether the regression originates in retrieval rather than generation.

Why this answer

Groundedness measures whether the answer is supported by the retrieved context, so a change to retrieval is the most likely culprit. Inspecting the logged retrieved chunks per evaluation row reveals whether the new embedding model returns less relevant or contradictory passages, which the judge then flags as unsupported claims. This isolates retrieval as the cause before any remediation.

Exam trap

The trap here is assuming a groundedness drop always indicates a hallucinating generator, when the metric is just as sensitive to degraded retrieved context.

85
MCQmedium

A team has deployed a model that is experiencing high latency. How should they identify if the bottleneck is the model inference or the preprocessing code?

A.Check the cluster cost in the Databricks billing console.
B.Use MLflow Tracing to inspect the execution pipeline.
C.Re-train the model on a larger dataset.
D.Increase the memory limit for the inference cluster.
AnswerB

MLflow Tracing provides detailed visibility into the duration of every step in the request flow, from input preprocessing to model inference and output post-processing. This allows engineers to pinpoint exactly where the latency is occurring, enabling them to optimize the specific components that are causing the performance delays.

Why this answer

Distinguishing between inference time and preprocessing time is vital for performance tuning. By using MLflow's tracing capabilities or custom timing logs, engineers can isolate specific segments of the request pipeline. This granular visibility allows for targeted optimization, such as optimizing the vector search or reducing prompt overhead, rather than blindly attempting to speed up the model itself, which is often significantly more resource-intensive.

Exam trap

Candidates often guess that general platform monitoring metrics are sufficient. They fail to realize that MLflow Tracing is specifically required to break down the request into discrete components like retrieval and generation.

86
MCQmedium

A GenAI engineer is building a RAG chatbot backed by a Databricks Vector Search index that is synced from a Delta table with Change Data Feed enabled. Source documents are frequently updated, and some are deleted. The engineer wants the index to reflect deletions and updates automatically with minimal operational overhead. Which configuration should be used?

A.Create a Direct Vector Access index and write a Structured Streaming job that upserts embeddings into it whenever the source table changes.
B.Create a Delta Sync index and enable Change Data Feed on the index itself so that deletions flow back into the source Delta table.
C.Create a Delta Sync index with a pipeline_type of TRIGGERED and schedule a Databricks Job to run a full refresh every 15 minutes.
D.Create a Delta Sync index with a pipeline_type of CONTINUOUS, pointing at the source Delta table, and let Databricks manage synchronization.
AnswerD

A CONTINUOUS Delta Sync index uses the source table's Change Data Feed to incrementally propagate inserts, updates, and deletes to the vector index without manual refreshes. This satisfies the requirement for automatic reflection of deletions and updates while minimizing operational overhead, because Databricks manages the sync pipeline and embedding computation end to end.

Why this answer

Automatic reflection of inserts, updates, and deletes from a Delta source into a Vector Search index is achieved with a Delta Sync index running in CONTINUOUS pipeline mode, which consumes the source table's Change Data Feed. Triggered refreshes, Direct Vector Access, and index-side Change Data Feed all either add manual work, introduce staleness, or rely on unsupported behavior.

Exam trap

The trap here is assuming that a scheduled full refresh or a Direct Vector Access index provides the same incremental delete handling as a continuous Delta Sync index.

87
MCQeasy

Which Databricks feature is best suited for maintaining the lineage of data used during the preparation of training sets for Generative AI?

A.Delta Lake Time Travel
B.Unity Catalog
C.Cluster Policies
D.MLflow Experiments
AnswerB

Unity Catalog captures fine-grained lineage information as data moves from raw ingestion to final training sets. It offers a unified view of dependencies, which is essential for auditing the training process. This visibility is crucial for ensuring that training data meets safety and compliance standards in enterprise environments.

Why this answer

Unity Catalog provides centralized governance, including automated data lineage tracking. This allows engineers to trace the source of data used for model training, which is vital for reproducibility and regulatory compliance. Knowing the exact provenance of training data helps in troubleshooting model bias and ensures that sensitive data sources are correctly identified and managed throughout the lifecycle of the model training process.

Exam trap

Candidates often suggest external logging tools or manual documentation, failing to recognize that Unity Catalog provides native, automated lineage tracking that is integrated directly into the data platform.

88
MCQmedium

A GenAI engineer is deploying a chain that calls an external LLM API. The chain must not block the serving thread while waiting on the remote API, and the endpoint must handle many concurrent requests. Which implementation approach should the engineer choose when logging the model?

A.Batch all incoming requests by writing them to a Delta table and have a separate job poll the table and call the external API.
B.Implement predict as an async function that awaits an async HTTP client, so the event loop can interleave other requests while waiting on the external API.
C.Use a synchronous HTTP client inside predict and increase the endpoint's scale-to-zero timeout so queued requests eventually complete.
D.Set the endpoint's concurrency to one and rely on the external API's own load balancing to absorb parallel requests.
AnswerB

Model Serving supports async predict, and an async HTTP client releases the event loop during network waits, allowing other requests to be handled concurrently on the same replica. This directly addresses the non-blocking requirement and improves throughput for I/O-bound chains. It is the recommended pattern for calling external LLM APIs from a served model.

Why this answer

Calling an external LLM API is I/O-bound, so the serving process should not block a thread while waiting. Writing predict as an async function with an async HTTP client lets the event loop serve other requests during those waits, improving concurrency on each replica. Synchronous clients, table-based queues, and single-concurrency settings all undermine throughput or break the synchronous API contract.

Exam trap

The trap here is treating a remote LLM call like local computation, where a synchronous client seems adequate, when in fact blocking I/O caps concurrency on the endpoint.

89
MCQmedium

A GenAI engineer is building a retrieval-augmented generation (RAG) application using Databricks Vector Search. During testing, they observe that for some queries, the retrieval step returns document chunks that are semantically similar but not actually relevant to the user's question, leading to poor answer quality. They want to improve retrieval precision without changing the embedding model. Which of the following approaches is most appropriate?

A.Reduce the number of retrieved chunks (top_k) to only the single highest-scoring chunk.
B.Switch the embedding model to a larger, more powerful model to improve semantic representation.
C.Enable hybrid search on the Vector Search index to combine keyword-based and vector similarity scores.
D.Increase the chunk size of the documents to include more context per retrieved chunk.
AnswerC

Hybrid search in Databricks Vector Search combines dense vector similarity with keyword-based (lexical) matching, which helps surface chunks that contain exact terms from the query. This improves precision when purely semantic matches drift to unrelated content. It does not require changing the embedding model and is a supported index configuration, making it the most direct and effective fix for the described issue.

Why this answer

Hybrid search enhances retrieval precision by blending vector similarity with keyword matching, ensuring that chunks containing exact query terms are ranked higher. This directly addresses the problem of semantically similar but irrelevant results without altering the embedding model. The other options either do not tackle the precision issue or violate the constraint of not changing the embedding model.

Exam trap

The trap here is assuming that any change to the retrieval pipeline, such as adjusting chunk size or top_k, will improve precision, when in fact only hybrid search directly combines lexical and semantic signals.

90
MCQhard

An engineer is troubleshooting a Vector Search index that fails to update. The index relies on a Delta table that is frequently updated. What is the most likely cause for the index failing to reflect new data?

A.The Delta table must be converted to a Parquet format.
B.The synchronization process was not triggered after the table update.
C.The embedding model version changed in the middle of a sync.
D.Unity Catalog permissions do not allow read access to the source table.
AnswerB

Vector Search indexes require an explicit trigger or a periodic sync configuration to ingest changes from the source Delta table. Without this action, the index does not automatically update, causing the RAG application to rely on outdated embeddings, which directly impacts the accuracy and relevance of the generative model's output.

Why this answer

Vector Search indexes are managed assets that require periodic synchronization with the source Delta table. If the synchronization process is not triggered automatically or via an API call, the index will remain stale. Understanding the lifecycle and sync mechanisms of Vector Search is essential for developers to ensure that RAG systems retrieve the most current information, as static indexes quickly lose relevance in dynamic data environments.

Exam trap

Test-takers often assume that updating the source Delta table automatically updates the Vector Search index in real time without needing a synchronization trigger.

91
MCQhard

You need to store embeddings generated by an LLM in a Delta table. Which data type is most efficient for storing these high-dimensional vector arrays in Databricks?

A.Store as a Base64 encoded string.
B.Store as an array of doubles.
C.Store as separate columns for each dimension (e.g., dim1, dim2, ...).
D.Store as a compressed binary blob (e.g., serialized byte array).
AnswerB

An array of doubles is the standard format for vector embeddings in Delta Lake. It is natively supported, memory-efficient, and easily accessible by Spark's vectorized query engine and popular libraries like FAISS or MosaicML Vector Search. This format provides the best balance between storage performance and computational efficiency for retrieval.

Why this answer

Storing embeddings as an 'Array' of 'Floats' (or 'Doubles') in a Delta table is the most efficient and native way to handle them. This allows Spark to utilize vectorized operations for similarity search calculations and enables compatibility with various indexing libraries. Using this approach ensures high performance during retrieval, minimizing the latency of finding the most relevant context for the LLM during generation.

Exam trap

Candidates often assume complex custom objects or string-serialized JSON are necessary. Delta natively supports arrays, which are significantly more efficient for vector math and similarity searches.

92
MCQmedium

When evaluating an LLM for deployment, what is a crucial 'non-functional' requirement that must be addressed?

A.The model's ability to learn from user feedback automatically.
B.The inference latency and throughput under expected load.
C.The number of parameters in the model architecture.
D.The language in which the model was initially trained.
AnswerB

Latency and throughput are critical non-functional requirements for production AI applications. If an application cannot respond within a timeframe acceptable to the user or handle the expected volume of concurrent requests, it will fail to meet business objectives, regardless of how accurate or intelligent the model is.

Why this answer

Non-functional requirements like latency, throughput, and cost are as important as the model's accuracy. In production, an accurate model that is too slow to provide a response is useless. Similarly, a high-performing model that is prohibitively expensive to run is not viable.

Balancing these factors is essential for ensuring that the generative AI application is both technically feasible and financially sustainable at scale.

Exam trap

Students often focus exclusively on accuracy metrics or benchmark scores, forgetting that operational factors like latency and throughput are essential non-functional requirements.

93
MCQmedium

A GenAI engineer is preparing a corpus of HTML product pages stored in a Unity Catalog volume for a RAG application. The pages contain navigation bars, script tags, and boilerplate footers that add noise to embeddings. Which Databricks-native approach best removes this noise while preserving the main article text before chunking?

A.Register the volume as a Delta table and run a SQL query using regexp_replace to strip all tags, then tokenize the result with the built-in Databricks tokenizer.
B.Increase the chunk size to 4096 tokens so that boilerplate is a smaller proportion of each chunk, and rely on the embedding model to ignore the noise.
C.Load the HTML files with spark.read.text and apply a VectorAssembler to convert the raw strings into dense vectors, then cluster and drop outliers.
D.Use an ai_parse_document or BeautifulSoup-style parser in a PySpark UDF to extract the main content, then write the cleaned text as a Delta table with a content column.
AnswerD

Parsing HTML with a document-aware parser lets you target the main article region and discard navigation, script, and footer nodes before text is written. Storing the cleaned content in Delta makes downstream chunking and embedding reproducible. This is the Databricks-native pattern for turning raw HTML in a volume into analysis-ready text.

Why this answer

Cleaning HTML before chunking requires a parser that understands document structure, not a regex or a numeric feature transformer. Extracting the main content with a document-aware parser and persisting the cleaned text to Delta produces a reliable, lineage-tracked input for chunking and embedding, which is the goal of data preparation for RAG on Databricks.

Exam trap

The trap here is assuming that a simple regex tag-stripping step or a larger chunk size is equivalent to semantic HTML cleaning.

94
Multi-Selectmedium

An organization is deploying an LLM application using Databricks Asset Bundles (DABs). Which TWO of the following are primary benefits of using DABs for the deployment process?

Select 2 answers
A.Automatic generation of unit tests for Python code.
B.Enables infrastructure-as-code (IaC) for Databricks resources.
C.Simplifies multi-environment deployment configuration.
D.Provides a drag-and-drop interface for building pipelines.
E.Automatically migrates all legacy data to Unity Catalog.
AnswersB, C

DABs allow engineers to define complex Databricks resources like Jobs, Pipelines, and Model Serving endpoints as code. By using YAML files, teams can manage their entire environment configuration in version control, enabling repeatable deployments and standardizing infrastructure across multiple workspaces through automation.

Why this answer

Databricks Asset Bundles streamline the development and deployment lifecycle by treating infrastructure and code as a single unit. They facilitate consistent environments across development, staging, and production by using YAML-based configuration files. This ensures that the deployment process is reproducible, automated via CI/CD pipelines, and manageable through version control, reducing human error and configuration drift in complex ML-driven applications.

Exam trap

Candidates sometimes focus on secondary features like UI ease-of-use, missing the primary architectural benefits of DABs: infrastructure-as-code capabilities and consistent multi-environment management for automated CI/CD pipelines.

95
MCQmedium

Refer to the exhibit. What is the impact of min_instances: 0 on this deployment?

A.The endpoint will be permanently disabled.
B.The endpoint will scale to zero when idle, reducing costs.
C.The endpoint will have higher latency than a fixed instance.
D.The endpoint will be restricted to CPU usage only.
AnswerB

This configuration allows the endpoint to release all compute resources when no requests are being processed. This is a best practice for cost management in AI deployments, as it prevents paying for expensive GPU instances when they are not actively serving traffic, enabling more sustainable and efficient cloud usage.

Why this answer

Setting min_instances to 0 in Databricks Model Serving enables 'scale-to-zero'. This feature is highly effective for cost optimization, as it automatically shuts down the GPU resources when no traffic is present. Upon receiving a new request, the endpoint will automatically scale up, though this may introduce a slight 'cold start' latency.

This trade-off is often acceptable for non-critical or low-traffic services to drastically reduce infrastructure costs.

Exam trap

Candidates often fear that 'min_instances: 0' will cause the model to be deleted or unavailable, failing to realize it is a standard cost-saving feature for serverless endpoints.

96
MCQhard

A team is building a RAG application using Databricks Vector Search with a Delta table as the source. They need the index to automatically reflect new and updated chunks as the source table changes, without rebuilding the entire index each time. Which configuration should they use?

A.Enable Delta Lake time travel on the source table and point the Vector Search index at a specific version using VERSION AS OF.
B.Use an external HNSW index built on a Databricks cluster and register it in Unity Catalog as a model, then query it with the model serving endpoint.
C.Create a Vector Search index with sync mode set to TRIGGERED or CONTINUOUS and specify a Delta table with Change Data Feed enabled as the source.
D.Create a standard Vector Search index and schedule a nightly job that drops the index and recreates it from the full Delta table.
AnswerC

Databricks Vector Search supports Delta Sync indexes that read Change Data Feed from the source Delta table to incrementally update embeddings. Enabling CDF and choosing a triggered or continuous sync keeps the index fresh without full rebuilds. This is the supported pattern for production RAG pipelines where source documents are frequently updated.

Why this answer

Delta Sync indexes in Databricks Vector Search rely on Change Data Feed to pick up inserts, updates, and deletes incrementally. Choosing a triggered or continuous sync mode keeps embeddings current without full recomputation. Enabling CDF on the source Delta table is a prerequisite, and this pattern is the documented way to keep a RAG index fresh as documents change.

Exam trap

The trap here is thinking that Vector Search automatically tracks changes to any Delta table, when incremental sync requires Delta Sync and Change Data Feed to be configured.

97
MCQhard

A team uses MLflow LLM Evaluation with an LLM judge to score a summarization agent. On a 400-row golden dataset the judge marks 96% of summaries as relevant. Spot-checking reveals the judge approves nearly every summary whenever the summary is fluent, even when key facts are missing. The team wants a defensible quality signal before approving a release. Which action best addresses this judge weakness?

A.Replace the relevance metric with a cheaper lexical overlap score such as ROUGE against the reference summaries.
B.Expand the golden dataset from 400 rows to 4,000 rows and rerun the same relevance metric.
C.Add a small human-labeled calibration set and compare the judge's verdicts against human labels before trusting the automated relevance scores.
D.Raise the judge's temperature so it explores a wider range of judgments and produces a more nuanced relevance distribution.
AnswerC

The judge exhibits a systematic bias toward fluency, so its scores cannot be trusted at face value. Labeling a representative subset by hand and measuring agreement between judge and human verdicts quantifies that bias and reveals which cases the judge mishandles. Once agreement is characterized, the team can report calibrated results, adjust thresholds, or supplement the judge with a fact-coverage metric, making the release decision defensible rather than resting on inflated automated scores.

Why this answer

An automated judge that approves nearly every fluent summary is a biased measurement instrument, and the fix is to calibrate it rather than to change its sampling or the dataset size. Human-labeling a representative subset and measuring judge-human agreement exposes the leniency and tells the team how much to trust the scores. Only then can thresholds be set or a complementary fact-coverage check be added, producing a release signal that can be defended.

Exam trap

The trap here is trying to fix a systematically lenient judge by enlarging the dataset or lowering rigor, when the real need is to calibrate the judge against human labels.

98
MCQmedium

You are deploying a RAG application. You need to ensure the model uses the most recent vector data without redeploying the model. What should you use?

A.Fine-tune the model with new data daily.
B.Implement a RAG pattern with a vector database index.
C.Re-register the model in MLflow with every update.
D.Embed all data into the model weights at deployment.
AnswerB

The RAG pattern retrieves relevant context from a vector database and feeds it into the model's prompt. Since the vector store can be updated independently of the model, this provides a scalable way to ensure the application always has access to the latest data without needing to redeploy models.

Why this answer

Decoupling the model logic from the data retrieval mechanism is a best practice. By using a retrieval-augmented generation pattern where the retriever fetches data dynamically from a vector store, the model remains static while the knowledge base evolves. This approach is essential for applications requiring real-time information, as it avoids the expensive and slow process of retraining or fine-tuning the base model whenever the underlying data changes.

Exam trap

Many candidates incorrectly think they need to fine-tune or retrain the foundational LLM whenever data changes, missing the efficiency of dynamic RAG architecture.

99
MCQmedium

Which of the following metrics is most effective for evaluating a RAG-based chatbot's ability to retrieve relevant context from a vector database during production monitoring?

A.Model Training Loss
B.Context Relevance
C.Inference Latency
D.Parameter Count
AnswerB

Context relevance measures how much of the retrieved information is actually necessary to answer the user query. Low scores indicate a failure in the embedding or vector search configuration, providing actionable insights for tuning the retriever component, which is a critical part of RAG evaluation.

Why this answer

Context relevance and faithfulness are key RAG metrics. Specifically, evaluating the retrieved context's alignment with the user query is essential for identifying retrieval failures. If the context is irrelevant, the model cannot generate a correct answer even with high-quality generative capabilities.

Monitoring this metric helps distinguish between errors caused by the retriever versus errors caused by the generator model during the production lifecycle.

Exam trap

Candidates frequently select generation metrics like faithfulness when asked specifically about evaluating how well the vector database retrieves information for the user query.

100
MCQmedium

Which security configuration is essential when deploying a model that accesses sensitive data stored in Unity Catalog?

A.Grant the endpoint access to the entire Unity Catalog.
B.Use a Service Principal with scoped permissions for the model.
C.Store credentials in plaintext in the model code.
D.Bypass Unity Catalog to use local file system access.
AnswerB

Using a dedicated Service Principal with finely-tuned permissions ensures that the serving endpoint can only access the specific data assets required for its tasks. This aligns with security best practices by minimizing risk and ensuring that the model does not have broad access to sensitive corporate information.

Why this answer

Accessing sensitive data requires strict adherence to the principle of least privilege. Assigning a specific Service Principal to the model serving endpoint ensures that it only has access to the data necessary for inference. This configuration is critical for security, as it limits the potential blast radius if the model endpoint is ever compromised, maintaining data privacy in line with organizational governance policies.

Exam trap

Candidates frequently select broad permissions like 'Workspace Admin' or 'User' roles, failing to recognize that Service Principals must be explicitly scoped to follow the principle of least privilege.

101
MCQhard

A GenAI engineer is building a retrieval-augmented chatbot whose answers must cite the exact source document and page. The team wants the chatbot's responses to include structured citations that downstream UIs can render. Which Databricks feature should the engineer use to return these structured citations from the model endpoint?

A.Databricks Vector Search's built-in reranking, which automatically appends source metadata to the model's generated answer.
B.MLflow Tracing with span attributes that capture the retrieved document IDs and page numbers for each LLM call.
C.Mosaic AI Agent Framework's document citation support, where the retriever returns chunks with metadata and the agent formats them into a structured citations field in the response.
D.Mosaic AI Model Serving's payload logging, which writes each request and response to an inference table that the UI can read for citations.
AnswerC

The Mosaic AI Agent Framework supports returning retrieved documents alongside the model's answer, exposing a structured citations payload that includes source identifiers and metadata such as page numbers. This is precisely the mechanism for delivering renderable citations from a deployed agent endpoint. It pairs the retriever's chunk metadata with the agent's response schema, matching the scenario's requirement.

Why this answer

Structured citations must be part of the model endpoint's response contract, not a side channel. The Mosaic AI Agent Framework's citation support lets the retriever supply chunk metadata and the agent emit a citations field the UI can render directly. Observability traces, retrieval reranking, and payload logging serve different purposes and cannot substitute for returning structured citations to the caller at inference time.

Exam trap

The trap here is equating observability artifacts such as MLflow traces or inference tables with the response payload, when only the agent's response schema can deliver citations to the UI.

102
MCQmedium

A team needs to provide fine-grained access control to an LLM application that uses a vector index stored in a Delta table. Which Unity Catalog feature should they use to restrict access to specific rows based on user identity?

A.Column-level masking
B.Row-level security (RLS)
C.Workspace-level permissions
D.Access Control Lists (ACLs)
AnswerB

Row-level security policies filter rows at query time based on user attributes. When an application queries the vector index, the Databricks engine automatically applies the filter, ensuring the RAG model only has access to authorized context, which is the standard mechanism for implementing secure data-aware applications.

Why this answer

Row-level security (RLS) in Unity Catalog allows administrators to define policies that filter data returned from tables based on the user's identity. By applying these policies to the Delta table used as a knowledge base, the team ensures that the LLM only retrieves information the user is authorized to see. This is essential for compliance and security in multi-tenant or sensitive information application environments.

Exam trap

Candidates often confuse Row-Level Security with Column-Level Security or Attribute-Based Access Control. RLS is specifically designed to filter rows based on the current user's identity.

103
MCQhard

A company wants to share a GenAI application's prompt templates and evaluation datasets with a partner organization. The partner uses its own Databricks account and must not access any other company data. The company also wants to revoke access easily. Which Unity Catalog feature should they use?

A.Export the prompt templates and evaluation datasets to cloud storage and share a bucket URL.
B.Create a new metastore for the partner and attach the company's workspace to it.
C.Grant the partner's users direct SELECT privileges on the company's Unity Catalog tables.
D.Delta Sharing with a share containing only the prompt templates and evaluation datasets.
AnswerD

Delta Sharing allows sharing specific tables or views with external organizations without copying data. By creating a share that includes only the prompt templates and evaluation datasets, the company limits access to those assets. Access can be revoked by removing the recipient from the share, meeting the easy revocation requirement.

Why this answer

Delta Sharing is designed for secure cross-organization data sharing without copying. By creating a share with only the prompt templates and evaluation datasets, the company limits exposure to those assets. The partner accesses the share from their own Databricks account, and the company can revoke access by removing the recipient, satisfying both isolation and revocation needs.

Exam trap

The trap here is assuming that direct permission grants or cloud storage exports can achieve cross-account sharing with easy revocation, when Delta Sharing is the purpose-built feature.

104
MCQmedium

A data engineer wants to share a specific subset of sensitive PII data with an external department using Unity Catalog. The engineering team must ensure the data is anonymized dynamically based on the user's role without creating physical copies. Which feature is most appropriate?

A.Create a temporary view that filters rows using a WHERE clause and grants SELECT access.
B.Apply a masking function to the column and use row filters in Unity Catalog.
C.Export the data to a new table with masked columns and grant access to the department.
D.Configure access control lists (ACLs) on the underlying S3 bucket or ADLS container.
AnswerB

Unity Catalog supports dynamic data masking and row filters, which apply policies at query time based on user identity. This enables granular control over data access without duplicating data or creating complex view hierarchies. It is the industry-standard method for enforcing privacy compliance across the Databricks Data Intelligence Platform.

Why this answer

Unity Catalog row-level security and column-level masking allow for dynamic data governance. By applying a masking function using 'current_user_role()' or 'is_account_group_member()', administrators can ensure that users see redacted data without duplicating the underlying storage. This approach centralizes governance policies, simplifies auditing, and ensures that sensitive data exposure is strictly controlled at the query execution time across all compute resources within the Unity Catalog metastore environment.

Exam trap

Candidates frequently choose to create physical duplicate tables with redacted columns, forgetting that Unity Catalog supports dynamic row filters and column masking natively.

105
Multi-Selecthard

A Generative AI engineer is setting up MLflow LLM Evaluation for a RAG pipeline on Databricks that answers questions about product manuals. The team wants to evaluate both the retrieval stage and the generation stage in a single evaluation run. Which TWO metric groups should they include in the evaluation configuration? (Choose two.)

Select 2 answers
A.System metrics such as CPU utilization and memory consumption of the serving endpoint
B.Retrieval metrics such as retrieval_relevance and retrieval_groundedness
C.Training metrics such as loss and learning rate from the fine-tuning job
D.Data freshness metrics such as table update timestamps in Unity Catalog
E.Generation metrics such as groundedness and relevance
AnswersB, E

Retrieval metrics score the chunks returned by the retriever against the question, so they directly evaluate the retrieval stage. Including them in the same run lets the team see whether the right passages were fetched before the generator produced its answer, which is necessary to attribute quality problems to retrieval rather than generation.

Why this answer

A complete RAG evaluation run needs metrics for both stages. Retrieval metrics score the chunks returned by the retriever, while generation metrics such as groundedness and relevance score the final answer against the question and context. Together they let the team pinpoint whether a quality issue originates in retrieval or in generation.

Exam trap

The trap here is mixing infrastructure or training telemetry into an evaluation run, when retrieval and generation quality metrics are the only ones that measure pipeline output.

106
MCQhard

A GenAI engineer prepares a Delta table of product descriptions that will feed a chunking and embedding pipeline. The descriptions are written in mixed languages, and the embedding model supports only English. The engineer needs to ensure that non-English rows are detected and routed for translation before embedding. Which approach is most appropriate?

A.Translate every description into English before chunking, regardless of its original language, to guarantee uniform embeddings.
B.Add a language detection UDF that runs ai_classify or a language-identification library on each description, write the detected language to a column, and split the table into English and non-English subsets.
C.Drop all rows where the description contains non-ASCII characters, since those rows are likely non-English.
D.Configure the embedding model endpoint to accept a language parameter and pass the detected locale from Unity Catalog tags on the table.
AnswerB

Detecting language per row and persisting the result creates an auditable routing column. Splitting the table lets English rows proceed directly to chunking and embedding while non-English rows go to a translation step. This is deterministic, testable data preparation and it keeps the pipeline extensible if more languages are added later.

Why this answer

When the embedding model is English-only, the pipeline must identify which rows need translation. Detecting language per row and storing the result as a column gives a deterministic routing signal, so English rows embed directly while non-English rows are sent to translation. This keeps the preparation step auditable and avoids both data loss and unnecessary translation cost.

Exam trap

The trap here is assuming that a table-level tag or a character-range filter can substitute for per-row language detection.

107
Multi-Selectmedium

A team is building a retrieval-augmented generation application on Databricks and wants to reduce hallucination by improving the quality of retrieved context before it reaches the LLM. Which TWO techniques should they apply? (Choose two.)

Select 2 answers
A.Increase the LLM temperature to encourage more diverse answers.
B.Apply a reranker model to the top-k retrieved chunks before passing them to the LLM.
C.Disable the Vector Search index and rely on keyword matching only.
D.Chunk documents with overlap and metadata so retrieval can filter by relevant attributes.
E.Store the entire document corpus in the prompt for every request.
AnswersB, D

A reranker scores retrieved candidates against the query with a cross-encoder, reordering them so the most relevant chunks appear first. This directly improves the precision of the context window and reduces the chance the LLM grounds its answer in loosely related passages. It is a standard post-retrieval step in Databricks RAG pipelines.

Why this answer

Reranking retrieved candidates with a cross-encoder and improving chunking with overlap and metadata both raise the relevance of the context supplied to the LLM, which directly reduces hallucination. Raising temperature, stuffing the full corpus, and dropping semantic search all degrade grounding rather than improve it.

Exam trap

The trap here is equating more context or more randomness with better answers, when grounding quality depends on retrieving fewer, more relevant chunks.

108
MCQmedium

A team runs an LLM-as-a-judge evaluation on Databricks using the built-in `mlflow.evaluate()` with `model_type="databricks-agent"`. They notice that the judge model, a serving endpoint, is producing scores that are consistently inflated compared to human review. Which configuration change should the team make first to improve the reliability of the evaluation?

A.Switch the judge to a different, stronger model endpoint and calibrate it using a small human-labeled dataset with known ground-truth scores.
B.Reduce the number of evaluation examples to only those where the judge and human agree, then report the average score.
C.Increase the temperature parameter of the judge model to 1.0 to encourage more diverse scoring.
D.Replace the LLM judge with a BLEU score comparison against a reference answer for every question.
AnswerA

Using a more capable judge model and calibrating it against human-labeled examples is the recommended way to reduce systematic bias. Calibration with ground-truth scores helps detect and correct consistent over- or under-scoring. This approach aligns with MLflow's guidance to validate LLM judges before relying on their outputs in production monitoring.

Why this answer

LLM judges can exhibit systematic bias, such as consistently inflated scores. The most effective first step is to use a stronger judge model and calibrate it against human-labeled data with known scores. This improves reliability and aligns with Databricks recommendations for validating judges before using them in production monitoring.

Simply changing temperature or filtering data does not address the root cause.

Exam trap

The trap here is assuming that increasing the judge model's temperature will make it more objective, when in fact it adds randomness and worsens reproducibility without fixing bias.

109
MCQhard

A GenAI engineer is building a chatbot using Databricks Model Serving and needs to ensure that the endpoint can only be invoked by users who have been granted the 'genai_users' group. The endpoint is registered in Unity Catalog. Which configuration is required to enforce this access control?

A.Grant USE CATALOG and USE SCHEMA on the catalog and schema containing the endpoint to the 'genai_users' group.
B.Grant EXECUTE on the model serving endpoint to the 'genai_users' group.
C.Attach the endpoint to a cluster that has been granted CAN ATTACH TO for the 'genai_users' group.
D.Set the endpoint's owner to the 'genai_users' group and rely on ownership for access.
AnswerB

In Unity Catalog, model serving endpoints are securable objects that support the EXECUTE privilege. Granting EXECUTE to the 'genai_users' group allows only those users to invoke the endpoint. Without this grant, other users cannot call the endpoint, thus enforcing the required access control.

Why this answer

Model serving endpoints registered in Unity Catalog are secured with the EXECUTE privilege. To allow only the 'genai_users' group to invoke the endpoint, an administrator must grant EXECUTE on the endpoint to that group. This ensures that only authorized users can call the endpoint, meeting the access-control requirement.

Exam trap

The trap here is assuming that USE CATALOG/USE SCHEMA or cluster permissions control endpoint invocation, but the EXECUTE privilege is the specific grant needed.

110
MCQhard

A team is deploying a RAG application using Databricks Model Serving with a foundation model endpoint and a Databricks Vector Search index. During load testing, they observe that p95 latency spikes when the retriever returns many chunks, and the LLM occasionally truncates context. They want to reduce latency while preserving answer quality. Which change is most appropriate?

A.Increase the number of retrieved chunks and rely on the LLM to ignore irrelevant content.
B.Switch the Vector Search index to a smaller embedding dimension and keep the same number of chunks.
C.Apply a reranker to the retrieved chunks and pass only the top few highest-scoring chunks to the LLM.
D.Lower the LLM temperature to zero to reduce response time.
AnswerC

A reranker scores retrieved chunks against the query with higher precision than the initial vector similarity, so the application can pass fewer but more relevant chunks to the LLM. This shortens the prompt, reducing latency and the chance of context truncation, while preserving answer quality because the most relevant evidence is retained. It directly addresses both symptoms observed during load testing.

Why this answer

Reranking retrieved chunks and passing only the top few to the LLM reduces prompt length, which lowers LLM latency and mitigates context truncation. Because the reranker prioritizes the most query-relevant chunks, answer quality is preserved even with fewer chunks. The other options either increase prompt size, change an unrelated parameter, or alter the index without addressing prompt length.

Exam trap

The trap here is treating temperature as a performance lever, when latency in RAG is dominated by prompt length and the volume of retrieved context, not by sampling randomness.

111
MCQeasy

A developer is using MLflow to track experiments for a generative AI application. They want to log a prompt template and its associated parameters so that they can reproduce the exact input to the model later. Which MLflow function should they use to log the prompt template as an artifact?

A.mlflow.log_artifact
B.mlflow.set_tag
C.mlflow.log_metric
D.mlflow.log_param
AnswerA

mlflow.log_artifact logs a file or directory as an artifact to the MLflow run. This is ideal for storing prompt templates as text files, allowing versioning and easy retrieval. It supports any file type and is the standard way to persist non-parameter assets like prompts, configuration files, or evaluation datasets.

Why this answer

Prompt templates are best stored as artifacts because they are files that may be reused, versioned, and referenced independently of the run. mlflow.log_artifact provides a robust way to persist such files, ensuring reproducibility and easy access later.

Exam trap

The trap here is confusing parameters with artifacts; parameters are for small configuration values, while artifacts are for files like prompt templates.

112
MCQmedium

When deploying a Generative AI application, why is it recommended to use a dedicated Serving Endpoint rather than a shared interactive cluster?

A.Shared clusters are cheaper to run for production.
B.Serving endpoints provide better isolation and predictable performance.
C.Shared clusters cannot be used for Python-based models.
D.Serving endpoints automatically train the model on new data.
AnswerB

Serving endpoints ensure that inference requests are not competing with interactive workloads or data science tasks for CPU and GPU resources. This isolation is critical for maintaining predictable, low-latency performance in production, which is a non-negotiable requirement for high-quality generative AI user experiences.

Why this answer

Dedicated serving endpoints provide resource isolation, performance predictability, and standardized API interfaces. Shared clusters are prone to resource contention, inconsistent environment states, and lack the operational features required for reliable production inference. Separating the serving infrastructure from development clusters is a standard architectural pattern that ensures the stability and scalability of AI-driven features in user-facing applications.

Exam trap

Candidates might assume interactive notebook clusters are acceptable for production due to lower initial setup complexity, ignoring scalability and cost.

113
MCQeasy

What is the primary benefit of using Unity Catalog when designing generative AI applications in Databricks?

A.It automatically generates Python code for model training.
B.It provides centralized governance, lineage, and access control for data assets.
C.It converts unstructured text into vectors automatically.
D.It eliminates the need for data preprocessing before RAG.
AnswerB

Unity Catalog enables secure, governed access to all data used in the AI lifecycle. By tracking data lineage from the source to the model, it ensures transparency and compliance. This centralized approach simplifies security management and audit readiness, which are essential when handling proprietary data in generative AI applications.

Why this answer

Unity Catalog provides a unified governance and security layer, ensuring that data used for training, fine-tuning, or RAG is access-controlled and lineage-tracked. This is crucial for compliance and reproducibility in AI projects. By centralizing permissions and audit logs, teams can securely manage access to sensitive data across their entire generative AI pipeline, which is a fundamental requirement for enterprise AI adoption.

Exam trap

Candidates often focus on the performance benefits of Unity Catalog, overlooking its primary enterprise role as a centralized governance and security layer for data assets.

114
MCQeasy

A logistics company monitors a Databricks-hosted RAG assistant that answers questions about shipping regulations. Over one week, the retrieval index was rebuilt after a document refresh, and the team observes that the rate of answers flagged as unsupported by retrieved context rose sharply while retrieval relevance scores stayed flat. Which metric should they inspect first to determine whether the regression originates in the retrieval stage or the generation stage?

A.The number of thumbs-up reactions collected from end users during the same week.
B.Endpoint request latency, comparing the p95 before and after the index rebuild.
C.Context recall, measured against a labeled set of questions with known ground-truth source documents.
D.Token usage per response, broken down by prompt and completion tokens on the serving endpoint.
AnswerC

Context recall checks whether the retriever surfaced the documents that actually contain the answer. Since relevance stayed flat but unsupported answers rose after an index rebuild, the retriever may now be returning topically similar but non-authoritative chunks, which recall against ground-truth sources detects immediately and separates retrieval failure from generation failure.

Why this answer

Context recall against a labeled ground-truth set directly tests whether the retriever brought back the documents that contain the answer. Because relevance measures topical similarity and can stay high even when authoritative chunks are missed, only recall can reveal that the rebuilt index now returns plausible but non-supporting passages, which would explain unsupported answers without any change in generation behavior.

Exam trap

The trap here is assuming flat relevance scores exonerate the retriever, when relevance measures topical similarity and can remain high while authoritative chunks are silently dropped after an index rebuild.

115
MCQeasy

Which Databricks feature is primary for managing the lifecycle, versioning, and deployment readiness of custom Generative AI models?

A.Unity Catalog Volumes
B.MLflow Model Registry
C.Databricks SQL Warehouse
D.Delta Live Tables
AnswerB

The MLflow Model Registry provides a centralized store for managing the full lifecycle of a model. It allows teams to version models, transition them between lifecycle stages like 'Staging' and 'Production', and maintain a comprehensive history of model development, which is essential for reliable and reproducible deployments in production environments.

Why this answer

MLflow Model Registry is the central feature for managing the lifecycle of models, including versioning and stage transitions. Understanding how to promote a model from 'Staging' to 'Production' is fundamental for ensuring that only tested and verified models reach the end-users. This workflow is critical for maintaining quality and stability in generative AI applications, as it provides a structured process for model evolution and deployment management.

Exam trap

Examinees often guess general MLflow tracking features instead of identifying the MLflow Model Registry as the dedicated tool for lifecycle and deployment readiness.

116
MCQmedium

A developer is building a RAG application. Which step is essential to prevent the model from hallucinating or providing outdated information?

A.Increase the temperature parameter of the model to 1.5.
B.Use a larger model to automatically detect and correct hallucinations.
C.Implement a real-time data ingestion pipeline for the vector database.
D.Force the model to always answer using a specific tone of voice.
AnswerC

Keeping the vector database up-to-date with the latest information is the most effective way to prevent hallucinations caused by outdated data. By ensuring that the RAG pipeline always retrieves the most current documents, the model is grounded in relevant, recent facts, which significantly improves the reliability and accuracy of the output.

Why this answer

Ensuring data freshness and providing high-quality, relevant context to the LLM during the retrieval step is critical. By implementing a robust data ingestion pipeline that updates vector indices in near real-time, the application ensures that the model is always informed by the most recent documents. This practice mitigates the risk of hallucinations by grounding the model's responses in verified, up-to-date business data, which is essential for building trustworthy GenAI solutions.

Exam trap

Candidates often choose manual static dataset uploads or model fine-tuning, overlooking that preventing hallucinations with fast-changing data specifically requires a real-time data ingestion pipeline into the vector database.

117
MCQmedium

An enterprise LLM application on Databricks is experiencing high latency and inconsistent responses. Which approach best enables observability to identify the root cause of these performance bottlenecks within the LLM pipeline?

A.Enable standard Spark UI logs to monitor cluster resource utilization during inference.
B.Review the Databricks SQL query history to identify slow retrieval operations.
C.Use MLflow Tracing to record inputs, outputs, and execution duration of each step in the chain.
D.Deploy an external monitoring agent to scan the model server network traffic.
AnswerC

MLflow Tracing captures end-to-end execution details for LLM workflows. It allows developers to visualize the entire dependency graph and latency of each step, enabling precise identification of bottlenecks in chains, prompt templates, or vector search lookups.

Why this answer

Integrating MLflow Tracing provides granular visibility into the execution flow, including model inputs, outputs, and latency per step. This is critical for diagnosing complex chains where intermediate calls may be stalling. By instrumenting the code, engineers can identify which specific retrieval step or prompt generation block is contributing to the high latency, allowing for targeted optimization of the RAG pipeline components.

Exam trap

Candidates often choose standard cluster log analysis or simple end-to-end latency metrics instead of MLflow Tracing to inspect individual steps within an LLM chain.

118
MCQeasy

A data scientist has developed a GenAI application locally and now needs to share it with stakeholders as an interactive web app inside the Databricks workspace. The app must use the workspace's existing authentication and compute. Which Databricks capability should be used?

A.Databricks Apps, which host the application within the workspace and integrate with workspace identity and compute.
B.A Databricks job with a notebook task scheduled to run every five minutes.
C.An MLflow experiment tracking run that logs the application as an artifact.
D.A Databricks notebook with the %run magic command to execute the application logic on demand.
AnswerA

Databricks Apps is the native capability for hosting interactive web applications inside the workspace. It integrates with workspace authentication, so stakeholders use their existing Databricks credentials, and it can access workspace compute and data resources without exposing separate credentials.

Why this answer

Databricks Apps is designed for hosting interactive applications directly in the workspace, using the workspace's identity and compute. This lets stakeholders access the GenAI app through a URL with existing authentication, avoiding the need to build custom hosting or manage separate credentials.

Exam trap

The trap here is confusing interactive development tools like notebooks with app-hosting services, when only Databricks Apps provides a shareable, authenticated web interface.

119
MCQeasy

When developing a Generative AI application in Databricks, which tool provides a collaborative environment for engineers to write code, visualize data, and document their experiments using Markdown?

A.Databricks File System (DBFS)
B.Databricks Notebooks
C.Unity Catalog Explorer
D.Databricks SQL Editor
AnswerB

Databricks Notebooks are the primary tool for collaborative data science and AI development. They support multi-language code execution, interactive visualizations, and Markdown documentation, providing the necessary features for engineers to build, document, and test GenAI models in a unified, shared environment that is accessible to the whole team.

Why this answer

Databricks Notebooks provide a versatile, collaborative environment designed for data science and AI development. They allow developers to combine code, SQL queries, visualizations, and rich text documentation (Markdown) in a single document. This makes them ideal for building GenAI applications, as engineers can document their model tuning processes, test retrieval logic, and share findings with stakeholders directly within the platform, facilitating a highly efficient and iterative development workflow.

Exam trap

Candidates often confuse Databricks Notebooks with external IDEs like VS Code or Databricks Repos, failing to recognize that Notebooks are the primary tool for integrated, collaborative experimentation with Markdown.

120
MCQeasy

An engineer is building a GenAI application that must return structured JSON output conforming to a specific schema so downstream systems can parse it reliably. The team wants to enforce the schema at generation time rather than post-processing free-form text. Which approach should they use?

A.Fine-tune the foundation model on examples of the desired JSON output.
B.Post-process the model output with a regular expression that extracts the first JSON object.
C.Set the LLM temperature to zero and instruct the model to return JSON in the system prompt.
D.Use a structured output mode or response format that constrains generation to the provided JSON schema.
AnswerD

Constrained decoding against a JSON schema, exposed by foundation model APIs as a structured output or response format setting, restricts token generation so the output must conform to the schema. This enforces field names, types, and required properties at generation time, eliminating the need for brittle post-processing. It is the appropriate mechanism when downstream systems require reliably parseable structured output.

Why this answer

Structured output or response format settings constrain the model's decoding process to the supplied JSON schema, guaranteeing that generated content parses and conforms. This is more reliable than prompt instructions, fine-tuning, or regex extraction because it operates at generation time and eliminates malformed or incomplete output. Downstream systems can then parse results without defensive fallback logic.

Exam trap

The trap here is assuming that a well-crafted prompt plus temperature zero guarantees valid JSON, when only constrained decoding against a schema enforces structural conformance.

121
MCQmedium

Which action must an administrator perform to allow a user to use Databricks SQL to query a table that is stored in an external storage location?

A.Grant the user the OWNER role for the storage credential.
B.Grant the user the READ FILES privilege on the external location.
C.Grant the user the ALL PRIVILEGES privilege on the metastore.
D.Add the user to the workspace-level admin group.
AnswerB

Querying an external table requires the user to have the READ FILES privilege on the external location object that manages the underlying storage path. This is in addition to the standard catalog, schema, and table privileges, ensuring that storage access is explicitly governed.

Why this answer

To query an external table, the user needs access to the storage location, the catalog, the schema, and the table. The administrator must grant the USAGE privilege on the catalog and schema, the SELECT privilege on the table, and the READ FILES privilege on the external location. This multi-layered approach ensures that data access is both logically and physically controlled.

Exam trap

Examinees frequently confuse table-level SELECT permissions with storage-level privileges, forgetting that querying external tables requires explicit READ FILES permissions on the external location.

122
MCQmedium

A team wants its Databricks App to read a secret that stores an external API key used by an agent tool. The secret is managed in a Databricks secret scope. Which approach correctly exposes the secret to the running app without putting the value in source control?

A.Pass the key as a bundle variable in databricks.yml so it is substituted during deployment.
B.Commit the API key into a .env file in the app repository and load it with a dotenv library at startup.
C.Store the key in a Delta table and have the app query the table with its service principal at startup.
D.Reference the secret scope and key in the app configuration's environment variables using the secrets-backed value syntax, then read it from the environment in code.
AnswerD

Databricks Apps can declare environment variables whose values are resolved from a secret scope using the secrets-backed reference syntax. At runtime the platform injects the resolved value into the process environment, so the code reads it like any environment variable and the secret never appears in source or logs. This is the supported way to pass secrets to a deployed app.

Why this answer

The app configuration can declare environment variables whose values are resolved from a Databricks secret scope using the secrets-backed reference syntax. The platform injects the resolved value at runtime, so the key is available to the process without ever being written into source, committed to a repository, or surfaced through deployment output.

Exam trap

The trap here is using bundle variables or a .env file for sensitive values, when secret scope references in the app configuration are the sanctioned mechanism.

123
MCQmedium

A financial services team runs a RAG assistant on Databricks that answers questions about internal policy documents. During a review, the team finds that for many questions the answer is factually correct but cites a document that does not actually contain the supporting statement. They want an MLflow LLM Evaluation metric that specifically detects when the response is not supported by the retrieved context. Which metric should they add to their evaluation run?

A.answer_similarity
B.groundedness
C.retrieval_precision
D.toxicity
AnswerB

The groundedness metric in MLflow LLM Evaluation judges whether the claims in the response are supported by the retrieved context. In this scenario the answers are correct but not backed by the cited documents, which is exactly the failure mode groundedness is designed to surface, making it the right metric to add to the evaluation run.

Why this answer

Groundedness evaluates whether the response's claims can be traced back to the retrieved context, which is precisely the gap identified in this review. Adding it to the MLflow LLM Evaluation run gives the team a metric that flags answers which sound correct but are not actually supported by the cited policy documents.

Exam trap

The trap here is confusing factual correctness with grounding, assuming that an accurate answer must also be supported by the retrieved context.

124
Multi-Selectmedium

A data team is developing a Mosaic AI Agent application. Which TWO of the following are mandatory for deploying this application using the Databricks Model Serving infrastructure?

Select 2 answers
A.A logged model signature defining input and output schemas
B.A manual provisioning of a dedicated GPU worker pool
C.A requirements.txt or conda environment file
D.A hardcoded API key for external vector store access
E.A pre-configured Kubernetes manifest file
AnswersA, C

A model signature is mandatory for Model Serving to validate incoming JSON payloads. It allows the serving endpoint to automatically enforce data types and structure, preventing malformed requests from reaching the model code. This ensures consistency and simplifies debugging during the inference request-response cycle.

Why this answer

Deploying Mosaic AI Agents requires defining a signature for the model input/output and ensuring the agent code is packaged with its dependencies. These steps are crucial because the Serving environment needs to understand how to route requests and instantiate the environment. Properly defining the signature ensures type safety during inference, while dependency management prevents runtime errors, ensuring the agent operates as expected within the serverless environment.

Exam trap

Candidates often select manual deployment scripts or external orchestration tools, missing the core mandatory package requirements like signatures and environment files.

125
MCQeasy

Which of the following describes the 'Gold' layer in a Medallion architecture, and why is it important for GenAI data preparation?

A.It contains raw, unprocessed data for initial exploration.
B.It stores transient data used only for debugging intermediate steps.
C.It is the final, curated state of data, optimized for consumption by ML models.
D.It is a temporary cache for speeding up cluster startup times.
AnswerC

The Gold layer provides clean, validated data. For GenAI, this means the text is pre-processed, chunks are optimized, and PII is scrubbed. By consuming data from the Gold layer, engineers ensure that their models are learning from the highest quality sources, which significantly improves overall application performance and reliability.

Why this answer

The Gold layer represents highly refined, business-level aggregates or prepared datasets ready for consumption. In GenAI, this is where the final, cleaned, and curated training sets (or vector-ready documents) reside. Having a Gold layer ensures that models are trained on validated, high-quality data, which is fundamental to building reliable, production-grade Generative AI applications that meet organizational standards for accuracy and data governance.

Exam trap

Candidates frequently mistake the Gold layer for the 'Silver' layer, which is cleaned but not necessarily aggregated or business-ready for final consumption by GenAI applications.

126
MCQmedium

A machine learning engineer needs to deploy a custom Mosaic AI Model Serving endpoint that requires access to a private internal database. Which mechanism should be used to securely manage the database credentials?

A.Hardcode the credentials as environment variables in the model serving endpoint deployment configuration.
B.Inject the credentials via a public GitHub repository linked to the Databricks Git folder.
C.Reference the credentials using the Databricks Secrets API within the model loading script.
D.Store the credentials in a plain-text file on the Unity Catalog volume attached to the serving endpoint.
AnswerC

The Databricks Secrets API allows code to retrieve sensitive values at runtime without exposing them in cleartext. This approach ensures that the credentials exist only in memory during the model execution phase. It provides a secure, auditable path for applications to authenticate with external services like private databases.

Why this answer

Databricks Secrets provide a centralized, secure way to store and reference sensitive information like database credentials. By using the secret scope, the engineer avoids hardcoding sensitive data into the model code or configuration files. This practice is essential for maintaining a secure MLOps lifecycle, as it prevents credential exposure and allows for granular access control via Databricks access control lists (ACLs) on the secret scope itself.

Exam trap

Test-takers frequently recommend hardcoding database credentials inside model artifacts or environment variables, ignoring secure workspace secret management practices.

127
Multi-Selectmedium

A platform team is preparing to deploy a GenAI chat application built on Mosaic AI Model Serving. They want the deployment to enforce least-privilege access for the application and to keep the serving environment reproducible. Which two practices should they implement? (Choose two.)

Select 2 answers
A.Allow the endpoint to install the latest version of every dependency at startup so the environment always has current patches.
B.Give the application principal workspace admin so it can create endpoints and modify cluster policies as needed during deployment.
C.Pin the model's dependency versions in the logged MLflow environment so the serving environment installs a known, reproducible set of packages.
D.Grant the application's service principal only the endpoint query permission and the minimum Unity Catalog privileges needed to read the model and its dependencies.
E.Disable Unity Catalog for the model and store it in the workspace model registry to simplify permission management.
AnswersC, D

Pinning dependencies in the model's logged environment ensures the serving container installs the same package versions used during development, which makes the deployment reproducible and avoids version drift that can break the agent. This directly supports the reproducibility objective and prevents surprises when the endpoint restarts or scales.

Why this answer

Least privilege is achieved by granting the application principal only endpoint query rights plus the minimum Unity Catalog reads for the model and its data. Reproducibility is achieved by pinning dependency versions in the logged MLflow environment so the serving container installs a known package set. Workspace admin, dropping Unity Catalog, and always-latest installs each trade away security or reproducibility, so they do not meet the two stated goals.

Exam trap

The trap here is equating 'convenient deployment' with broad admin rights or always-latest dependencies, when least privilege and pinned versions are what the scenario actually demands.

128
MCQeasy

A GenAI team is using MLflow to track experiments for a large language model. They want to ensure that only team members can view the experiment results and that the experiment artifacts are stored in a governed location. Which Unity Catalog integration should they use to manage MLflow experiments?

A.Use the MLflow tracking server with a personal access token for each team member
B.Store experiment artifacts in a Unity Catalog volume and use workspace ACLs on the notebook
C.Enable inference tables on the MLflow experiment to log access
D.Register the MLflow experiment in Unity Catalog and grant appropriate privileges on the experiment object
AnswerD

Unity Catalog supports registering MLflow experiments as securable objects. Once registered, you can grant privileges like USE SCHEMA and SELECT on the experiment to control access. Artifacts are stored in a Unity Catalog volume or external location, providing governance and auditability. This directly meets the requirements for access control and governed storage.

Why this answer

Registering an MLflow experiment in Unity Catalog makes it a securable object, allowing you to grant privileges to specific groups. Artifacts are stored in a governed location such as a Unity Catalog volume or external location. This provides both access control and governance, meeting the team's requirements.

Exam trap

The trap here is assuming that workspace ACLs or personal access tokens provide sufficient governance for MLflow experiments, when in fact Unity Catalog registration is required for fine-grained access control.

129
Multi-Selectmedium

A GenAI engineer is developing a conversational agent using Databricks. The agent must maintain context across multiple turns and retrieve relevant information from a knowledge base. They want to ensure the agent can handle follow-up questions that refer to previous exchanges. Which TWO techniques should they implement to manage conversation state and retrieval effectively? (Choose two.)

Select 2 answers
A.Store the conversation history in a Delta table and include it in the prompt for each turn.
B.Cache all responses from the LLM to avoid re-processing similar queries.
C.Use a separate Vector Search index for each conversation to retrieve relevant past messages.
D.Use a fixed window of the last three messages as context, ignoring older messages.
E.Rewrite follow-up questions into standalone queries using the conversation history before retrieving from the knowledge base.
AnswersA, E

Persisting conversation history in a Delta table allows the agent to retrieve past interactions and include them in the prompt, enabling context-aware responses. Delta tables provide ACID transactions and scalability, making them suitable for storing chat history. This approach ensures that follow-up questions can reference earlier turns, and the history can be trimmed or summarized to fit the context window.

Why this answer

Storing conversation history in a Delta table ensures persistence and allows inclusion of past turns in the prompt for context. Rewriting follow-up questions into standalone queries using that history improves retrieval accuracy by making the query self-contained. Together, these techniques enable the agent to handle multi-turn conversations effectively, where follow-up questions depend on earlier exchanges, and ensure the retrieval step has the necessary context.

Exam trap

The trap here is thinking that a separate Vector Search index per conversation or a fixed context window is sufficient, when scalable conversation state and query rewriting are the key techniques.

130
MCQhard

A Mosaic AI Agent application deployed as a Databricks App intermittently returns stale answers after the team updates the underlying vector index. The app caches a client to the serving endpoint and an index handle at module import time. Which change best resolves the staleness while keeping latency low?

A.Increase the app's compute size so the cached handle is refreshed more frequently by the runtime.
B.Disable caching entirely and rebuild the endpoint client and index handle on every request.
C.Redeploy the app after every index update so the module reinitializes with a fresh handle.
D.Initialize the endpoint client once but resolve the current index version per request, and invalidate the cached handle when the version changes.
AnswerD

Keeping the connection client but resolving the index version per request, with cache invalidation on version change, balances freshness and latency. The expensive client setup happens once, while the index reference is refreshed when the underlying data changes, so answers reflect the latest index without a redeploy. This directly addresses staleness while preserving low per-request overhead.

Why this answer

Staleness comes from caching an index handle at import time, so the fix is to keep the cheap-to-reuse client but resolve the index version per request and invalidate the handle when it changes. This preserves the latency benefit of a persistent client while ensuring answers reflect the latest index, avoiding both constant rebuilds and redeploy-on-every-update.

Exam trap

The trap here is assuming a redeploy or more compute fixes staleness, when the real issue is a long-lived cached handle that is never invalidated.

131
Multi-Selecthard

A Generative AI engineer is configuring a Mosaic AI Model Serving endpoint for a production-grade LLM. Which TWO of the following tasks are necessary to ensure effective monitoring and evaluation of the endpoint? (Choose two)

Select 2 answers
A.Enable Inference Tables on the serving endpoint.
B.Set up an automatic retraining job on the endpoint.
C.Implement an automated evaluation pipeline to score inference logs.
D.Increase the GPU count to maximize inference throughput.
E.Delete all request logs after 24 hours to save storage.
AnswersA, C

Enabling inference tables is the mandatory first step to store production request/response logs. Without this, you lack the raw material needed for post-hoc evaluation, analysis of failure modes, or detection of data drift. It creates a persistent data asset in Unity Catalog for ongoing oversight.

Why this answer

Effective monitoring requires both capturing data and evaluating that data against ground truth or automated metrics. Enabling inference tables provides the necessary raw data, while integrating with evaluation tools allows for systematic scoring of output quality. Together, these steps form the backbone of a robust monitoring strategy, ensuring that production drift is detected and quality is maintained according to business standards.

Exam trap

Candidates often select 'Model Training' or 'Feature Store' options, missing that inference monitoring requires capturing live request/response data via Inference Tables for evaluation.

132
MCQmedium

A data engineering team is deploying a RAG application using Mosaic AI Model Serving. They need to monitor the quality of the model's responses in production. Which Databricks feature should they use to capture and analyze inference data, such as requests, responses, and latency metrics?

A.Databricks SQL Alerts
B.Mosaic AI Inference Tables
C.Delta Live Tables Audit Logs
D.MLflow Experiment Tracking
AnswerB

Inference tables automatically log inference requests and responses to a Delta table in Unity Catalog. This enables seamless integration with monitoring tools for assessing model performance, latency, and throughput. It is the standardized method for capturing production data required to perform comprehensive model evaluation and drift detection.

Why this answer

Mosaic AI Model Serving provides built-in inference tables to automatically capture request and response logs. By enabling these tables, engineers can export data to a Unity Catalog table for analysis. This is critical for monitoring performance, data drift, and model quality over time.

Without this feature, teams lack the visibility required for production-grade LLM governance and continuous improvement cycles within the Databricks ecosystem.

Exam trap

Candidates often confuse inference tables with MLflow experiments or standard Delta tables created manually. They forget that Mosaic AI Inference Tables are a native, built-in feature specifically for capturing model serving request and response logs.

133
Multi-Selecthard

A team is deploying a GenAI application to a Databricks Model Serving endpoint and wants to ensure the deployment can be rolled back quickly and that traffic is shifted safely during updates. Which TWO practices should the team follow? (Choose two.)

Select 2 answers
A.Disable inference tables on the endpoint to reduce logging overhead during the update.
B.Deploy the new version by deleting the existing endpoint and creating a fresh one with the same name.
C.Set the endpoint's scale-to-zero behavior to always on to avoid cold starts during the update.
D.Configure the endpoint with multiple served entities and use traffic splitting to gradually shift traffic to the new model version.
E.Use Databricks Asset Bundles to define the endpoint configuration and deploy it through a CI/CD pipeline.
AnswersD, E

Model Serving supports multiple served entities on one endpoint with configurable traffic percentages. By routing a small percentage to the new version first, teams can validate behavior and shift traffic gradually, reducing risk and allowing a quick rollback by setting traffic back to the previous version.

Why this answer

Defining endpoints as code with Databricks Asset Bundles enables repeatable deployments and fast rollback by redeploying a prior version. Configuring multiple served entities with traffic splitting allows gradual rollout to the new version, so issues can be detected early and traffic reverted quickly, together providing safe updates and rollback.

Exam trap

The trap here is focusing on cost and latency knobs like scale-to-zero instead of the deployment controls, such as versioned configuration and traffic splitting, that actually enable safe rollback.

134
MCQmedium

A financial services company is designing a RAG assistant that must never return answers containing personally identifiable information from its knowledge base. The team plans to filter retrieved chunks before they are placed into the prompt. Which Databricks design element should they use to enforce this filtering consistently?

A.Add a system prompt instructing the model to ignore any personally identifiable information it sees in the context.
B.Store a sensitivity label or metadata field on each chunk in the Vector Search index and apply a metadata filter at query time to exclude restricted chunks.
C.Run a post-processing regex on the model output to remove patterns that resemble personally identifiable information.
D.Grant the serving endpoint a service principal that lacks access to the sensitive tables so retrieval returns no restricted rows.
AnswerB

Vector Search supports metadata filtering, so tagging chunks with a sensitivity attribute and filtering at query time prevents restricted content from ever entering the prompt. This enforces the policy at retrieval, which is earlier and more reliable than post-generation redaction, and it scales as new documents are labeled.

Why this answer

Metadata filtering in Vector Search allows the application to exclude chunks based on attributes such as sensitivity level before they reach the prompt. This enforces the policy at the retrieval boundary, reducing both leakage risk and token usage. Prompt instructions and output regexes act later and cannot guarantee that sensitive content never enters the model context.

Exam trap

The trap here is relying on prompt instructions or output filters for data protection, when the reliable control point is filtering retrieved chunks before prompt assembly.

135
MCQmedium

An organization requires that all GenAI models deployed in Databricks be tracked and managed with a unified registry for compliance. Which feature should the developer use?

A.Use Git branches to manage different model versions.
B.Unity Catalog Model Registry
C.Store model artifacts as Pickle files in a shared workspace folder.
D.Create a custom Python class to track model metadata in a Delta table.
AnswerB

Unity Catalog Model Registry allows for central management, governance, and deployment of models. It enforces consistent access policies, tracks the lineage of model artifacts, and manages the lifecycle stages of models, ensuring that only approved models reach production while maintaining a clear, auditable trail for compliance requirements.

Why this answer

Unity Catalog Model Registry serves as the centralized hub for governing, versioning, and deploying machine learning models across an organization. By using the Unity Catalog, developers ensure that every model has a documented lineage, access control, and deployment status. This is critical for regulatory compliance and enterprise security, as it prevents unvetted models from being deployed and provides a transparent audit trail of every model version currently in use.

Exam trap

Candidates frequently confuse legacy MLflow Model Registry workspaces with the organization-wide compliance and governance features provided by the Unity Catalog Model Registry.

136
MCQmedium

A GenAI engineer must build a fine-tuning dataset from 40 TB of raw JSONL conversation logs stored in cloud object storage. The logs are immutable and only ever read once, in full, during preprocessing. Which storage and access configuration should be used to minimize cost while keeping the data readable by Spark on Databricks?

A.Convert the JSONL to Parquet with an external tool, upload the Parquet back, then read it from the bucket.
B.Run COPY INTO to ingest the JSONL into a managed Delta table, then read the Delta table for preprocessing.
C.Register the cloud path as an external location in Unity Catalog and read the JSONL files directly with Spark.
D.Mount the object storage bucket to DBFS with a legacy mount and read files through the /mnt path.
AnswerC

Reading the immutable JSONL directly from its object-storage path through a Unity Catalog external location avoids copying 40 TB and avoids paying for a second persisted copy. Unity Catalog governs access while Spark reads the files in place, which matches the single-pass, read-only access pattern and minimizes both storage and egress cost.

Why this answer

Because the logs are immutable and read once in full, the lowest-cost approach is to leave them where they are and read them in place under Unity Catalog governance. Copying or converting the data creates a redundant full-size copy and extra compute for no reuse benefit, while legacy mounts sacrifice governance without reducing storage or access cost.

Exam trap

The trap here is assuming that ingesting raw logs into Delta is always the right first step, when single-pass immutable reads are cheaper and better governed when left in place.

137
Multi-Selecthard

A team is designing an LLM application that requires strict data privacy. Which TWO approaches ensure that sensitive data is not leaked during the model inference process or training?

Select 2 answers
A.Redact PII from prompts using local libraries before sending data to the model.
B.Enable public access on all input data buckets for faster retrieval.
C.Restrict training data access using Unity Catalog fine-grained permissions.
D.Store all API keys as plain text in notebook variables.
E.Use the model provider's default logging for all user queries.
AnswersA, C

Sanitizing prompts locally ensures that sensitive information is stripped away before it enters the model inference pipeline. This prevents the model from processing or accidentally memorizing PII, providing a critical layer of security that mitigates the risks associated with data leakage in third-party or internal LLM deployments.

Why this answer

Data privacy in GenAI necessitates both input sanitization and secure environment configuration. By using PII redaction libraries before sending tokens to the model and enforcing Unity Catalog access controls on training datasets, developers create a robust defense-in-depth posture. These techniques protect against prompt injection and unauthorized data exposure, which are critical security considerations when deploying large language models within enterprise-grade environments where compliance and data sovereignty are top priorities.

Exam trap

Candidates assume that server-side data encryption alone is sufficient for privacy, missing that client-side PII redaction and granular access controls are necessary to prevent leaks during inference and training.

138
MCQmedium

When integrating an external LLM via a Databricks Model Serving endpoint, how should the API credentials be managed to ensure they are not exposed in the application code?

A.Store credentials as environment variables in the notebook directly.
B.Use Databricks Secrets to reference credentials at runtime.
C.Hardcode the credentials in a hidden Python module.
D.Encrypt credentials and store them in a JSON file within the repo.
AnswerB

Databricks Secrets are designed to securely store and manage sensitive credentials. Using the secret utility API, developers can inject keys into their code at runtime without the keys ever being written to the source code or persisted in plain text, maintaining high security standards for application integrations.

Why this answer

Databricks Secrets provide a secure way to reference sensitive information like API keys without hardcoding them in notebooks or source files. By using the 'dbutils.secrets.get' function, the application pulls the key at runtime from a secure vault. This is a best practice for developers to ensure security and prevent credentials from being accidentally committed to version control systems or visible to unauthorized users within the workspace.

Exam trap

Candidates often suggest environment variables or hardcoded config files. Neither is secure in a collaborative Databricks workspace; Secrets are the only approved way to manage sensitive credentials.

139
MCQhard

Refer to the exhibit. An engineer is automating the deployment of a model to an endpoint using the Databricks CLI. Based on the error log provided, what is the most appropriate action to resolve this deployment failure?

A.Upgrade the model serving endpoint to a higher capacity cluster size.
B.Grant the service principal 'USE' and 'READ' permissions on the catalog, schema, and model.
C.Delete the model serving endpoint and recreate it using a personal access token.
D.Restart the Databricks workspace to force an update of the internal permission cache.
AnswerB

Unity Catalog requires explicit grants for service principals to access assets. The error indicates that the service principal lacks the necessary permissions to read the model metadata or download the model artifacts. Granting these specific privileges allows the serving infrastructure to access the model during deployment.

Why this answer

Deployment failures due to 'INSUFFICIENT_PERMISSIONS' often stem from the service principal used by the CI/CD pipeline lacking the necessary grants on the registered model in Unity Catalog. The engineer must ensure the principal has 'USE CATALOG', 'USE SCHEMA', and 'READ' permissions on the model version. Addressing this at the Unity Catalog level is the standard procedure for cross-service authorization in Databricks.

Exam trap

Candidates often try to resolve permission errors by checking workspace-level settings rather than focusing on the specific hierarchical grants (catalog, schema, model) required by Unity Catalog.

140
MCQmedium

A developer is creating a Databricks notebook to orchestrate a GenAI pipeline that includes data ingestion, vector index refresh, and model inference. They want to ensure that the pipeline can be easily tested and deployed across different environments. Which Databricks feature should they use to define the pipeline as code and manage deployments?

A.Databricks Jobs with notebook tasks and parameters.
B.Delta Live Tables with expectations.
C.MLflow Projects with a conda environment.
D.Databricks Asset Bundles (DABs).
AnswerD

Databricks Asset Bundles allow you to define jobs, notebooks, and other resources as code in YAML files. They support parameterization for different environments and can be deployed via CLI or CI/CD. This makes them ideal for managing GenAI pipelines across development, staging, and production, ensuring consistency and reproducibility.

Why this answer

Databricks Asset Bundles provide a declarative way to define and deploy Databricks resources, including jobs, notebooks, and configurations, as code. They support environment-specific parameters and CI/CD integration, making them the best choice for managing a GenAI pipeline across multiple environments.

Exam trap

The trap here is confusing orchestration tools like Jobs or Delta Live Tables with infrastructure-as-code solutions; Asset Bundles are specifically designed for packaging and deploying resources.

141
MCQmedium

A developer is using Databricks AI Functions to extract structured information from a large set of customer reviews stored in a Delta table. They want to apply a prompt to each review and store the results in a new column. Which function should they use?

A.ai_query()
B.ai_summarize()
C.ai_extract()
D.ai_classify()
AnswerA

ai_query() is a Databricks AI Function that allows you to run a prompt against a column of data in a SQL query. It sends the text to a specified model endpoint and returns the model's response, which can be stored in a new column. This is exactly suited for applying a prompt to each review in a Delta table.

Why this answer

Databricks AI Functions provide SQL-native access to generative AI models. ai_query() is the general-purpose function that accepts a prompt and a model endpoint, making it ideal for custom extraction tasks. It can be used in a SELECT statement to process each row and write the output to a new column. Other functions like ai_summarize and ai_classify are specialized and do not offer the same flexibility for structured extraction.

Exam trap

The trap here is assuming there is a dedicated ai_extract() function, when in fact ai_query() is the correct general-purpose function for custom prompts.

142
Multi-Selecthard

When designing an agentic workflow in Databricks, which TWO practices are essential to ensure the application remains observable and maintainable?

Select 2 answers
A.Enable MLflow tracing for all agent execution steps.
B.Avoid unit testing to speed up the development process.
C.Use automated evaluation suites to check for performance regressions.
D.Store all agent logic in a single, massive notebook file.
E.Disable logging to minimize storage consumption costs.
AnswersA, C

Tracing provides a detailed view of the agent's execution, including tool calls, reasoning steps, and model inputs/outputs. This visibility is essential for debugging complex agentic behaviors and ensures that developers can pinpoint where failures occur, ultimately making the agentic system easier to maintain and troubleshoot during production operations.

Why this answer

Observability and maintainability are critical for enterprise-grade AI agents. Using MLflow to log trace data provides visibility into the agent's decision-making process, which is necessary for debugging multi-step reasoning. Simultaneously, implementing robust evaluation suites ensures that agent updates do not introduce regressions in task performance.

These two practices form the foundation of a reliable AI operations strategy, enabling teams to confidently deploy and iterate on agentic workflows in production environments.

Exam trap

Test-takers often assume standard application logs are sufficient for complex multi-step agentic workflows, overlooking the necessity of deep execution tracing.

143
MCQmedium

A team's RAG evaluation shows high context recall but low answer correctness. Retrieved documents contain the needed facts, yet the generated answers frequently contradict them. Which single metric should they examine next to pinpoint whether the generator is ignoring or misusing the retrieved context?

A.Retrieval latency, to determine whether slow document fetches cause the generator to fall back on parametric memory.
B.Token count of the retrieved chunks, to verify whether the context window is overflowing and truncating documents.
C.Context precision, to check whether the most relevant chunks are ranked at the top of the retrieved list.
D.Groundedness, to check whether the claims in the generated answer are actually supported by the retrieved context.
AnswerD

Groundedness evaluates whether each statement in the answer is entailed by the retrieved context. High context recall with low correctness and low groundedness points to a generator that is not faithfully using the provided passages, whereas high groundedness would suggest a different failure mode such as a flawed ground-truth label.

Why this answer

With context recall high, retrieval supplied the necessary facts, so the failure lies downstream in generation. Groundedness directly measures whether the answer's claims are supported by the retrieved context, so a low groundedness score localizes the problem to the generator ignoring or misreading context. That distinguishes a generation-faithfulness issue from a retrieval-ranking or labeling issue.

Exam trap

The trap here is reaching for another retrieval metric when retrieval already proved sufficient via high context recall, so the investigation should move downstream.

144
MCQhard

A team deploys a Mosaic AI Agent application to a Databricks Model Serving endpoint. During load testing they observe that the first request after an idle period takes several seconds, while subsequent requests are fast. They want to eliminate this cold-start penalty for a latency-sensitive customer-facing application while keeping costs reasonable during off-peak hours. Which configuration should they apply?

A.Set the endpoint's minimum provisioned concurrency to a value greater than zero so at least one instance is always warm.
B.Reduce the agent's dependency footprint and re-log the model so the container image is smaller.
C.Enable inference tables on the endpoint to log requests and responses, which keeps the model warm.
D.Increase the endpoint's maximum concurrency per instance so a single warm instance can absorb all traffic.
AnswerA

Provisioned concurrency keeps a specified number of model instances loaded and ready, so requests never wait for a cold container to load the agent and its dependencies. Setting the minimum above zero removes the idle-time cold-start penalty while still allowing the endpoint to scale up under load. This is the intended control for latency-sensitive endpoints that must stay responsive.

Why this answer

Cold starts occur when the endpoint has scaled down and must load a fresh model container on the next request. Provisioned concurrency with a minimum above zero keeps at least one instance warm at all times, removing the first-request latency penalty while still permitting scale-up under load and scale-down to the configured minimum during off-peak hours. Concurrency, inference tables, and image size do not guarantee warm capacity.

Exam trap

The trap here is confusing concurrency limits or observability features with warm capacity, when only provisioned concurrency with a nonzero minimum prevents the endpoint from scaling fully to zero.

145
MCQmedium

A generative AI engineer is packaging a retrieval-augmented generation (RAG) application so it can be deployed as a Databricks App. The app reads the vector index name and the serving endpoint name from environment configuration so the same code can run in dev and prod. Which approach correctly supplies these values at deploy time using Databricks Asset Bundles?

A.Store the values in a notebook widget and have the app prompt the user for them at startup.
B.Declare bundle variables with defaults in databricks.yml, override them per target, and reference them in the app resource configuration.
C.Hard-code the dev and prod values in the app source and select a branch during deployment.
D.Write the values into a Delta table and have the app query that table on every request.
AnswerB

Bundle variables declared in databricks.yml with per-target overrides are the supported mechanism for environment-specific values. Referencing those variables inside the app resource passes the resolved endpoint and index names into the deployed app at deploy time, so the same source is promoted across targets without code edits and each target receives its own values.

Why this answer

Bundle variables with defaults in databricks.yml, overridden per target and referenced in the app resource, let one codebase deploy to dev and prod with different endpoint and index names. The bundle resolves values at deploy time and validates them, so promotion requires no source edits and each target receives the correct configuration without runtime prompts or extra lookups.

Exam trap

The trap here is assuming environment-specific settings must live in application code or be entered at runtime, when the bundle's variable and target mechanism is the intended injection point.

146
Multi-Selecthard

Which TWO of the following are primary components of the Unity Catalog identity model?

Select 2 answers
A.Workspace-local users
B.Account-level users
C.Groups
D.Cloud-provider-specific identities
E.Data asset aliases
AnswersB, C

Account-level users are the fundamental identity entity in Unity Catalog. By managing users at the account level, organizations ensure that a single identity is used across all workspaces, allowing for consistent permission assignment and easier lifecycle management of employees and service principals.

Why this answer

The Unity Catalog identity model is unified at the account level, meaning that identities are managed consistently across all workspaces. The core components are Users and Groups. By centralizing these identities in the account console, Databricks ensures that permissions are applied uniformly, regardless of which workspace a user is accessing, which is critical for large-scale enterprise governance and audit compliance.

Exam trap

Candidates frequently include 'Workspaces' or 'Metastores' as identity components, confusing the physical deployment infrastructure with the actual user/group identity objects managed in the account.

147
MCQmedium

Which workflow best describes the recommended CI/CD process for updating a Databricks Asset Bundle?

A.Make changes directly in the production workspace and export the updated YAML.
B.Update the local YAML, run 'bundle validate', and deploy to the target environment.
C.Use the Databricks UI to update the Job definition and update the YAML manually later.
D.Deploy the bundle to production without validation to speed up the delivery time.
AnswerB

This workflow follows standard DevOps practices: updating the configuration, validating it to ensure schema compliance, and deploying. Using the CLI ensures that the deployment process is repeatable and documented. Validation catches errors before they impact the environment, ensuring a smoother update process in production workspaces.

Why this answer

A robust CI/CD workflow for Databricks Asset Bundles involves validating the configuration, testing in a development workspace, and then deploying to production via an automated process. By validating the bundle before deployment, you catch syntax errors early. This pipeline-driven approach ensures that all changes are tracked in version control, reviewed through pull requests, and deployed consistently, minimizing the risks associated with manual workspace configuration changes.

Exam trap

Candidates often assume manual workspace changes are sufficient or forget the critical step of running 'bundle validate' before deployment, leading to syntax errors that only appear during the actual deployment process.

148
Multi-Selectmedium

A GenAI platform team is preparing a Unity Catalog schema to host a retrieval-augmented generation pipeline. They will store prompt templates and evaluation datasets as Delta tables, and they need to expose the pipeline to a group of application developers. The security team asks the platform team to describe how Unity Catalog privileges must be granted for the developers to query the tables. Which TWO statements correctly describe the required privilege model? (Choose two.)

Select 2 answers
A.Developers need CREATE TABLE on the catalog to read prompt templates stored as Delta tables
B.Developers need USE CATALOG on the catalog and USE SCHEMA on the schema before any table privileges take effect
C.Developers must be granted ALL PRIVILEGES on the catalog so that future objects are automatically accessible
D.Developers must own the schema so that they can create their own tables without further grants
E.Developers must be granted SELECT on each table, or SELECT on the schema to cover all current and future tables
AnswersB, E

Unity Catalog privilege evaluation is hierarchical. A principal must have USE CATALOG on the containing catalog and USE SCHEMA on the containing schema for table-level grants to be usable. Without those, even an explicit SELECT grant on a table is ineffective because the traversal of the namespace is blocked, so this is a genuine prerequisite for the developers.

Why this answer

Reading a table in Unity Catalog requires traversing the namespace and holding a data privilege. USE CATALOG on the catalog and USE SCHEMA on the schema make the table addressable, and SELECT on the table, or on the schema for broad coverage, authorizes the read. Together these two statements describe the minimum privilege set that lets the developers query the RAG pipeline's tables without over-granting administrative rights.

Exam trap

The trap here is granting broad catalog-level privileges such as ALL PRIVILEGES instead of the narrow USE CATALOG, USE SCHEMA, and SELECT combination required for read access.

149
MCQhard

Refer to the exhibit. An engineer wants to perform a canary deployment by routing 10% of traffic to a new version (version 6). How should the JSON traffic configuration be modified?

A.Add a new route for version 6 and set its traffic_percentage to 10, then set version 5 to 90.
B.Set the traffic_percentage of version 6 to 10 and keep version 5 at 100.
C.Remove version 5 and set version 6 to 10.
D.Create a new endpoint entirely for version 6 and switch the DNS record.
AnswerA

This configuration correctly implements a 90/10 traffic split. By splitting the traffic, the engineer can observe the performance of version 6 with real-world data without risking the entire workload. Once validation is complete, the percentages can be adjusted until the new version receives 100% of the traffic.

Why this answer

Canary deployments involve splitting traffic between an established model version and a new candidate. To achieve this, the 'routes' list must include both versions, with the 'traffic_percentage' values summing to 100%. This controlled rollout allows teams to monitor the performance and accuracy of the new model on a small segment of production traffic before committing the full load to the new version.

Exam trap

Candidates frequently forget that traffic percentages must sum to 100%. They often modify the new version's percentage without adjusting the existing version, leading to invalid configuration errors in the deployment.

150
MCQmedium

A GenAI engineering team has deployed a customer-support RAG chain on Databricks and registered it in Unity Catalog. They now want MLflow 3 to automatically score every production request for groundedness and relevance without writing custom scoring code, and to persist those assessments against the logged traces. Which approach should they use?

A.Configure the endpoint's autoscaling policy to collect evaluation metrics alongside throughput metrics.
B.Create a Databricks SQL dashboard that re-runs mlflow.evaluate() on the inference table every five minutes.
C.Enable the system table system.serving.endpoint_usage and query it for groundedness scores.
D.Attach MLflow scorers to the deployed agent so that assessments are computed on live traces and stored in the MLflow experiment.
AnswerD

MLflow 3 monitoring lets you attach built-in scorers such as groundedness and relevance directly to a deployed agent. The scorers run asynchronously on live traces and write assessment records back to the MLflow experiment, so no custom scoring code is required and results stay linked to each trace.

Why this answer

MLflow 3 production monitoring is designed for exactly this need: built-in scorers are attached to a deployed agent rather than invoked manually, and assessments are recorded on live traces inside the MLflow experiment. This removes custom scoring code, keeps quality signals tied to each request, and enables continuous monitoring without batch jobs or dashboard workarounds.

Exam trap

The trap here is assuming that any Databricks observability surface, such as system tables or dashboards, can produce semantic quality scores, when only MLflow scorers attached to the agent compute groundedness and relevance.

Page 1

Page 2 of 5

Page 3

All pages