Courseiva

Databricks Certified Generative AI Engineer Associate (Databricks-GenAI-Assoc) — Questions 301–330

330 questions total · 5pages · All types, answers revealed

Page 4

Page 5 of 5

301
MCQmedium

A developer is configuring a RAG application and needs to ensure that the LLM response is based on specific, trusted document snippets. Which technique, when implemented correctly, helps mitigate hallucination by grounding the response in provided context?

A.Increasing the model's temperature parameter
B.Retrieval Augmented Generation (RAG)
C.Fine-tuning the base model on all internal documents
D.Using a larger foundation model without context injection
AnswerB

RAG grounds the model's output by providing relevant, factual information from trusted sources within the prompt. This context-based approach limits the model's tendency to hallucinate by forcing it to answer based on the provided document snippets, which are retrieved via similarity search before the generation step occurs.

Why this answer

Retrieval Augmented Generation (RAG) is the primary technique for grounding LLM responses. By retrieving relevant, trusted documents from a vector store based on a user's query and injecting those documents into the LLM's prompt, the developer forces the model to synthesize an answer based on specific retrieved context rather than relying solely on its internal training data. This significantly reduces hallucinations and increases the accuracy and relevance of the generated responses.

Exam trap

Candidates often confuse RAG with Fine-tuning or Prompt Engineering. They assume the model's internal weights are being updated, when RAG is strictly about providing external context at inference time.

302
Multi-Selectmedium

A GenAI engineer is designing a RAG application on Databricks that must support multi-turn conversations where users refer to earlier messages, and the application must keep responses grounded in retrieved documents. Which TWO design elements are required to meet these requirements? (Choose two.)

Select 2 answers
A.Increase the LLM temperature so the model can explore multiple interpretations of ambiguous follow-up questions.
B.Store all conversation turns in a Delta table and query that table at inference time for each user message.
C.Disable retrieval for follow-up questions and rely on the model's parametric knowledge.
D.Rewrite the user query using conversation history before performing vector search retrieval.
E.Maintain conversation history and include relevant prior turns in the prompt sent to the model.
AnswersD, E

In multi-turn conversations, the latest user message often lacks the keywords needed for retrieval because it references prior context. Rewriting the query with conversation history produces a self-contained query that retrieves the correct documents. Without this step, vector search may return irrelevant chunks and the model's answer will not be grounded in the right evidence, breaking the grounding requirement.

Why this answer

Multi-turn conversation support requires passing relevant prior turns to the model so it can resolve references, and grounding requires that retrieval use a query rewritten with conversation context so the right documents are fetched. Together these elements ensure the model understands follow-up intent and answers from authoritative retrieved content. The other options either reduce determinism, add storage without context, or remove retrieval.

Exam trap

The trap here is treating conversation history storage as sufficient for multi-turn RAG, when the history must be used both in the prompt and in query rewriting before retrieval.

303
MCQhard

A GenAI engineer registers a fine-tuned model in Unity Catalog and wants a downstream application to call it through the Databricks Model Serving endpoint without embedding a long-lived personal access token in the application. The application runs on Azure Databricks and must authenticate as its own identity, and the security team requires that credentials be short-lived and automatically rotated. Which authentication approach should the engineer implement?

A.Create a Databricks service principal, generate a personal access token for it, and store that token in the application's environment variables
B.Use OAuth machine-to-machine authentication with the service principal's client ID and client secret to obtain a short-lived token
C.Configure the application to use the workspace user's username and password for basic authentication
D.Embed a Microsoft Entra ID managed identity token directly in the application source code
AnswerB

Databricks supports OAuth 2.0 client credentials flow for service principals, where the application exchanges its client ID and secret for an access token that expires in about an hour. The application authenticates as its own identity and never stores a durable bearer token. This matches the requirement for short-lived, automatically refreshed credentials and is the recommended pattern for unattended workloads.

Why this answer

OAuth machine-to-machine authentication lets an application present a service principal's client ID and secret to the Databricks OAuth token endpoint and receive an access token with a limited lifetime. The application has its own Unity Catalog identity, permissions can be granted to that principal on the registered model and serving endpoint, and token refresh happens automatically. This is the supported pattern for non-interactive workloads that must avoid long-lived secrets.

Exam trap

The trap here is treating a service principal's personal access token as equivalent to OAuth client credentials, when only the latter yields automatically rotated short-lived tokens.

304
MCQeasy

A data engineer needs to prepare a Delta table of customer reviews for embedding generation. The reviews contain HTML tags, inconsistent whitespace, and mixed casing that hurt embedding quality. Which preparation step should be applied before generating embeddings?

A.Apply a hash of the raw review text as the embedding input to guarantee deterministic vectors.
B.Increase the embedding vector dimension to capture the additional characters in the raw text.
C.Store the reviews in a Parquet file instead of Delta to improve text compression before embedding.
D.Normalize the text by stripping HTML tags, collapsing whitespace, and applying consistent casing.
AnswerD

Removing HTML tags, collapsing whitespace, and normalizing casing reduces noise that would otherwise dilute the embedding signal. Embedding models encode the literal tokens they receive, so markup and erratic spacing consume vector capacity without adding semantic value. Cleaning the text first yields more consistent and comparable vectors across the corpus.

Why this answer

Embedding models encode the exact tokens they receive, so HTML markup, irregular whitespace, and inconsistent casing add noise that dilutes semantic signal. Normalizing the text before embedding removes that noise and produces vectors that better reflect meaning. Storage format and vector dimension do not change the input text, so they cannot solve the quality problem.

Exam trap

The trap here is assuming that a storage or model configuration change can compensate for dirty input text, when the embedding model only sees the characters it is given.

305
MCQmedium

An application requires low-latency retrieval of RAG metadata stored in Databricks. Which storage approach balances performance and cost while ensuring seamless integration with Unity Catalog?

A.Store embeddings in a flat Parquet file on DBFS accessed via direct Spark reads.
B.Utilize a third-party managed vector database outside of the Databricks environment.
C.Deploy a Databricks Vector Search index on a Delta table with automatic synchronization.
D.Implement a custom KNN algorithm using standard Python libraries on a single node cluster.
AnswerC

Databricks Vector Search provides a managed, scalable service specifically optimized for low-latency similarity search. It automatically handles the synchronization between your Delta table and the vector index, ensuring that as your source data updates, your search capability remains current without manual intervention or pipeline management.

Why this answer

Vector Search indexes are designed for low-latency similarity search, which is critical for RAG applications. By leveraging Databricks Vector Search, you avoid the overhead of custom search implementations. This approach ensures that embeddings are automatically synchronized with Delta tables, providing a scalable solution that integrates directly with Unity Catalog's security and governance framework, which is essential for maintaining consistent data access policies across enterprise environments.

Exam trap

Candidates mistakenly choose manual embedding generation jobs or standalone external databases, ignoring the native integration of Databricks Vector Search with Delta tables.

306
Multi-Selecthard

A Generative AI engineer is evaluating a RAG pipeline using MLflow LLM Evaluation on Databricks. They want to assess both the retrieval quality and the generation quality. Which TWO built-in evaluation metrics should they use? (Choose two.)

Select 2 answers
A.`perplexity`
B.`toxicity`
C.`ROUGE`
D.`relevance`
E.`groundedness`
AnswersD, E

Relevance assesses how well the retrieved documents match the user's query, directly evaluating retrieval quality. In MLflow LLM Evaluation, the relevance metric can be computed for the retrieval step. It helps identify whether the retriever is fetching appropriate context. This is essential for diagnosing retrieval failures in a RAG pipeline.

Why this answer

For a RAG pipeline, retrieval quality is assessed by relevance, which measures how well retrieved documents match the query. Generation quality is assessed by groundedness, which checks if the answer is supported by the retrieved context. Both are built-in metrics in MLflow LLM Evaluation and are specifically designed for RAG evaluation on Databricks.

Exam trap

The trap here is selecting common NLP metrics like ROUGE or perplexity, which are not part of MLflow's built-in RAG evaluation suite and do not measure retrieval or groundedness.

307
MCQmedium

A data engineer is assembling a fine-tuning dataset from a Delta table of conversation transcripts. Each transcript contains a list of message objects with a role and a content field, and the training job requires one row per conversation with the messages serialized into the expected format. Which transformation should the engineer apply?

A.Use the explode function to expand each message object into a separate row before writing the dataset.
B.Use the collect_list aggregate to group the message objects back into an array per conversation and serialize that array into the required format.
C.Use the pivot function to turn message roles into columns with content values in the cells.
D.Use the flatten function to collapse the nested message array into a single string per row.
AnswerB

collect_list aggregates the individual message rows back into an ordered array grouped by conversation identifier, which is exactly the nested structure the training format expects. Serializing that array into the required representation produces one training example per conversation, matching the job's input contract.

Why this answer

The training job requires one row per conversation with the message sequence preserved, so the individual message rows must be grouped back together. collect_list aggregates messages into an ordered array per conversation, and serializing that array yields the nested format the job expects. Exploding, pivoting, or flattening would alter the structure and lose the required per-conversation layout.

Exam trap

The trap here is reaching for explode out of habit when flattening data, even though this scenario requires the opposite operation of aggregating messages back into one row per conversation.

308
MCQmedium

An engineer is building a multi-turn GenAI assistant on Databricks. The assistant must answer follow-up questions that reference earlier turns, such as 'what about its warranty?', while keeping each request within the model's context limit. Which design should the engineer implement?

A.Send only the current user message to the model on every turn and rely on the model's built-in memory of prior requests.
B.Maintain the conversation in a Delta table, then on each turn build a prompt from a bounded recent window of turns plus a summary of older turns, along with retrieved context.
C.Persist the conversation in a Delta table and, on each turn, resend the entire raw history along with retrieved context, trimming nothing.
D.Encode the entire conversation history into a single embedding and prepend that vector to the prompt so the model can reconstruct prior turns.
AnswerB

Combining a recent-turn window with a rolling summary of older turns preserves the antecedents needed for follow-up questions while bounding prompt size. Retrieving relevant context per turn keeps answers grounded. This design directly meets both requirements: coherent handling of references like 'its warranty' and staying within the model's context limit as the conversation lengthens.

Why this answer

Bounding prompt size while preserving conversational antecedents is achieved by keeping a recent window of turns and summarizing older ones, with per-turn retrieval for grounding. Stateless endpoints, untrimmed history replay, and vector-encoded history all fail to give the model usable prior context within the context limit.

Exam trap

The trap here is assuming the serving endpoint retains state between requests, which leads to designs that either omit prior turns entirely or replay unbounded history.

309
MCQhard

Refer to the exhibit. An engineer is testing a model endpoint. The outputs are too brief and often stop mid-sentence. What is the most likely cause?

A.The temperature is too low for the model to generate full sentences.
B.The top_p parameter is too high, leading to early termination.
C.The max_new_tokens limit is too low for the expected response length.
D.The model version is incompatible with the specified parameters.
AnswerC

The 'max_new_tokens' parameter restricts the number of tokens the model is permitted to generate. If the model's intended response exceeds this count, it stops generation. This configuration is the direct cause of the truncation observed, and it must be increased to accommodate longer, more complete responses.

Why this answer

The 'max_new_tokens' parameter is set to 512, which imposes a hard limit on the length of the generated output. If the model needs more tokens to complete its thought, it will be abruptly cut off. Increasing this limit is the logical troubleshooting step when encountering truncation issues in generative model responses, ensuring the model has sufficient space to finalize its output properly.

Exam trap

Candidates often blame the model's training or temperature settings for truncated output, failing to recognize that max_new_tokens is a hard configuration limit that forces an early stop.

310
MCQmedium

Refer to the exhibit. What is the cause of this error when logging a RAG chain to MLflow?

A.The vector index is offline.
B.The model was logged without a signature defining input and output types.
C.The user has insufficient Unity Catalog permissions.
D.The LLM endpoint is overloaded.
AnswerB

MLflow Models require a signature to define the expected input and output structure. This signature is critical for serving, as it allows the platform to validate incoming requests. Without it, the model cannot be registered or deployed to a serving endpoint, resulting in the error shown during the logging process.

Why this answer

This error occurs because the model signature is missing or incorrectly defined, which is required for MLflow to perform input validation and inference. When serving a RAG chain, MLflow needs to know exactly what the input schema looks like so it can properly route requests and manage data types. Providing a valid signature ensures the model is deployable and functional.

Exam trap

Candidates often assume the error is related to model size or environment libraries. They overlook the critical requirement that MLflow must have a defined input/output schema for inference.

311
Multi-Selecthard

A GenAI engineer is designing a retrieval-augmented generation application whose source documents are long PDFs. Early testing shows that answers are vague because retrieved chunks contain several unrelated topics, and the language model frequently cites content that does not support its claims. The engineer wants to improve chunk quality before indexing. Which TWO changes should the engineer make to the ingestion pipeline? (Choose two.)

Select 2 answers
A.Split documents along semantic boundaries such as sections and headings instead of a fixed character count
B.Store the full PDF text in a single chunk per document
C.Increase the embedding dimension of the model used for the index
D.Lower the temperature setting on the language model
E.Add overlapping context between adjacent chunks
AnswersA, E

Semantic splitting keeps each chunk focused on a single topic, so retrieved context is coherent and the language model is less likely to blend unrelated statements. Fixed-size splitting is what produced multi-topic chunks in this scenario, and respecting document structure directly addresses that root cause.

Why this answer

Chunk quality is determined at ingestion: splitting on semantic boundaries keeps each chunk topically focused, and overlap preserves meaning across boundaries. Together they raise the precision of retrieved context so the model has grounded material to cite. Embedding dimension, temperature, and oversized single chunks do not correct incoherent chunking.

Exam trap

The trap here is treating vague answers as a model-tuning problem, when the evidence points to ingestion-stage chunking that mixes topics and drops boundary sentences.

312
MCQmedium

A generative AI team runs a RAG chatbot whose Mosaic AI Model Serving endpoint is monitored in Unity Catalog inference tables. Over two weeks, the percentage of user questions that receive a refusal answer ('I don't have enough information') climbs from 4% to 31%, while retrieval latency and token counts stay flat. The team wants the earliest actionable signal that the retrieval corpus has gone stale rather than the prompt or model. Which monitoring signal should they inspect first?

A.The average generation time per completion recorded by the serving endpoint's request logs.
B.The endpoint's provisioned concurrency utilization and queue depth metrics from the serving endpoint logs.
C.The distribution of retrieval scores for the top-k chunks returned by the Vector Search index, tracked over time in the inference table.
D.The count of prompt tokens consumed per request, compared week over week in the billing usage table.
AnswerC

A rising refusal rate with unchanged latency and token counts points at the retrieval stage, not generation. If top-k similarity scores drift downward or their spread collapses, the index is returning progressively weaker matches for the same query mix, which is the earliest measurable evidence that the corpus has gone stale. Inspecting score distributions in the logged inference table isolates retrieval quality before blaming prompt or model changes.

Why this answer

Refusals rising while latency and token counts hold steady isolates the problem to retrieval quality rather than the model or the serving tier. Retrieval score distributions logged to the inference table show whether top-k chunks are becoming weaker matches, which is the earliest signal that the indexed corpus no longer covers current user questions. Checking capacity, token cost, or generation time cannot distinguish stale content from a healthy pipeline.

Exam trap

The trap here is assuming a rising refusal rate must be a prompt-engineering or model problem, when stable latency and token counts actually point to degraded retrieval relevance.

313
MCQmedium

A generative AI team is building a customer-support assistant on Databricks. The assistant must answer questions using the company's private knowledge base, and the team wants to minimize latency while ensuring that the LLM only uses retrieved documents. They plan to use Databricks Vector Search with a Delta table as the source. Which design choice best balances low latency and grounded responses?

A.Use a keyword-based search over the Delta table and pass the top 10 matching rows to the LLM, bypassing Vector Search entirely.
B.Use a Vector Search index with a relevance score threshold and retrieve the top 5 chunks, then include those chunks in the prompt with instructions to answer only from the provided context.
C.Store the entire knowledge base in a single prompt and use a long-context LLM to answer questions without retrieval.
D.Configure a Vector Search index with a small embedding dimension and retrieve the top 20 chunks, then pass all chunks directly to the LLM without filtering.
AnswerB

This approach grounds the LLM by supplying only the most relevant retrieved chunks and explicitly instructing it to rely on that context. Limiting to top 5 with a relevance threshold reduces prompt size and latency while filtering out weak matches. Databricks Vector Search supports similarity search with scores, enabling this pattern, and the instruction reduces hallucination risk.

Why this answer

Grounding requires retrieving only relevant context and instructing the model to use it. A Vector Search index with a relevance threshold and a small top-k (e.g., 5) reduces prompt size and latency while filtering weak matches. Including those chunks with an explicit instruction to answer from context limits hallucination.

This combination balances performance and accuracy better than unfiltered or keyword-only retrieval.

Exam trap

The trap here is assuming that retrieving more chunks always improves grounding, when actually larger top-k without a relevance threshold increases latency and introduces noise that can cause the model to ignore the provided context.

314
MCQmedium

An engineer is preparing to deploy a Mosaic AI Agent to a Model Serving endpoint. The agent depends on a custom Python library that is not available on PyPI. Which approach ensures the library is available at serving time?

A.Copy the library source into the notebook that defines the agent and rely on the notebook's sys.path.
B.Install the library manually on the driver node of the cluster used to log the model.
C.Add the library to the workspace's global init script so all clusters install it at startup.
D.Include the library as a wheel file in the MLflow model's requirements or artifacts so it is installed in the serving environment.
AnswerD

MLflow models capture their Python dependencies, and Model Serving builds the container using those dependencies. Including the custom wheel as an artifact and referencing it in the model's requirements ensures the library is installed in the serving environment, making it available to the agent at inference time.

Why this answer

Model Serving builds its container from the MLflow model's declared environment, so custom libraries must be packaged as artifacts and referenced in the model's requirements. This ensures the dependency is installed in the serving environment and the agent can import it reliably.

Exam trap

The trap here is assuming that installing a library on a development cluster will carry over to the serving endpoint, when serving environments are built solely from the logged model's dependency specification.

315
MCQhard

Refer to the exhibit. A developer encounters this error when trying to call a Model Serving endpoint from a job. Which action should the developer take to resolve this authorization failure?

A.Increase the 'workload_size' in the model configuration.
B.Grant 'Can query' permissions to the service principal in the endpoint settings.
C.Update the workspace token expiration policy.
D.Restart the serving endpoint to refresh security policies.
AnswerB

This is the correct approach to fix an authorization error. By updating the access control list (ACL) for the specific Model Serving endpoint to include the service principal with 'Can query' rights, the security policy is satisfied, allowing the job to successfully execute inference calls against the LLM.

Why this answer

The error indicates a lack of 'Can query' permission for the service principal or user attempting to access the endpoint. To resolve this, the developer must update the permissions for the specific endpoint within the Databricks workspace. By granting the required 'Can query' privilege to the service principal running the job, the application will regain authorization to perform inference against the model, satisfying the security policy requirements.

Exam trap

Test-takers often look for cluster-level spark configurations or IAM role modifications when encountering endpoint authorization errors, overlooking workspace access control lists.

316
MCQhard

A team runs an offline evaluation of a RAG chatbot using MLflow LLM Evaluation with an LLM judge. The same evaluation dataset produces noticeably different scores when re-run on different days, even though the application code and retrieved documents are unchanged. Which action best addresses this score instability?

A.Increase the number of retrieved chunks per query so the judge has more context
B.Switch all metrics from LLM-judged to exact-match string comparison
C.Reduce the evaluation dataset to a single representative example
D.Pin the judge model version and set temperature to zero for evaluation runs
AnswerD

Varying scores on identical inputs are a classic symptom of a nondeterministic judge: the judge model version may silently change, and sampling temperature introduces randomness in its verdicts. Pinning the judge model version and setting temperature to zero makes the scoring function repeatable, so differences in scores reflect real changes in the application rather than judge noise. This directly targets the instability described.

Why this answer

Identical inputs producing different scores across runs indicate the judge itself is nondeterministic, typically through model version drift or sampling temperature. Pinning the judge model version and forcing temperature to zero makes scoring reproducible, so observed differences can be attributed to the application. The other options either change the system under test, destroy evaluation fidelity, or hide variance rather than eliminate it.

Exam trap

The trap here is blaming the application or the retrieval pipeline for score drift, when unchanged inputs and code point squarely at the LLM judge's own nondeterminism.

317
MCQmedium

An engineer is designing a RAG application where the LLM must answer questions using only the retrieved context and must refuse to answer when the context is insufficient. Which prompt design approach best enforces this behavior?

A.Increase the number of retrieved chunks to ten so the model always has enough context to answer.
B.Use a system prompt that instructs the model to answer only from the provided context and to respond with a fixed refusal phrase when the context does not contain the answer.
C.Set the model's temperature to zero and rely on deterministic decoding to prevent unsupported answers.
D.Append the user's question twice in the prompt to reinforce the instruction to stay grounded.
AnswerB

A system prompt sets the model's operating constraints and is the most direct way to instruct it to rely solely on the retrieved context. Including an explicit refusal phrase gives the model a defined fallback, reducing hallucinations. This approach is deterministic in structure and works across providers that support system-level instructions.

Why this answer

The most reliable way to constrain a model to retrieved context is a system prompt that explicitly limits answers to that context and defines a refusal response when the information is missing. This gives the model a clear rule and a fallback, reducing hallucinations. Retrieval volume, temperature, and question repetition do not establish grounding boundaries.

Exam trap

The trap here is believing that deterministic decoding settings such as temperature zero eliminate hallucinations, when grounding must be enforced through prompt instructions.

318
MCQhard

An engineer maintains a GenAI application that uses a Databricks Asset Bundle to deploy a Mosaic AI Agent serving endpoint. A new model version has been logged and validated, and the team wants to roll it out to production with the ability to revert quickly if quality regressions appear. Which deployment approach best satisfies this?

A.Keep the existing endpoint unchanged and run the new model version only in notebooks until the team is confident, then deploy it later.
B.Create a new serving endpoint version by updating the endpoint's served model entity to the new model version, then verify quality and roll back by pointing the entity back to the previous version.
C.Deploy the new model version to a second workspace and route production traffic there through a global load balancer.
D.Delete the production endpoint and create a brand-new endpoint with the new model version, then update all clients to the new URL.
AnswerB

Updating the served entity on the existing endpoint switches traffic to the new model version while preserving the endpoint URL and configuration, enabling a fast rollback by reverting the served entity to the prior version. This gives the team a quick, low-risk path to production with a clear revert option if regressions appear.

Why this answer

Updating the served model entity on an existing endpoint moves production traffic to the validated model version while keeping the endpoint URL and configuration stable, so clients are unaffected. If quality regresses, reverting the served entity to the previous version restores prior behavior quickly. Recreating endpoints, notebook-only testing, and cross-workspace routing each introduce downtime, delay, or unnecessary complexity instead of a clean, reversible cutover.

Exam trap

The trap here is assuming a new model version requires a new endpoint, when the served entity on an existing endpoint can be updated and reverted in place.

319
MCQeasy

An engineer has finished building a RAG chatbot and wants to expose it as a Databricks App so business users can reach it through a browser. The app needs a Python web server and a command that starts it. Which artifact in the app's project layout defines the runtime command and dependencies used when the app is deployed?

A.The app.yaml file, which declares the command to run and the environment or dependencies for the app.
B.The MLflow model signature stored with the logged model, which encodes the serving command.
C.The requirements.txt file alone, which both lists dependencies and specifies the server start command.
D.The databricks.yml bundle file, which contains the Python entrypoint and pip requirements for the app.
AnswerA

Databricks Apps use an app.yaml file to define how the app starts and what it needs. It carries the run command, the source path, and environment or dependency references, so the platform knows how to launch the web server on deploy. This is the correct artifact for specifying startup behavior and dependencies in the app project layout.

Why this answer

A Databricks App is launched according to its app.yaml configuration, which specifies the command that starts the web server along with source and dependency information. The bundle file orchestrates deployment, requirements files list packages, and model signatures describe schemas, but only the app configuration tells the platform how to run the application.

Exam trap

The trap here is conflating deployment orchestration files with the app's own runtime configuration, when only the app configuration declares the start command.

320
MCQmedium

A GenAI engineer has built a retrieval-augmented generation (RAG) application using Databricks Vector Search and a Databricks-hosted LLM served via Mosaic AI Model Serving. Users report that responses are sometimes irrelevant or cite incorrect document passages. The engineer wants to systematically improve answer quality by identifying which retrieved chunks are actually being used by the LLM. Which approach should the engineer take to capture the relationship between retrieved context and the generated response for later evaluation?

A.Configure the Model Serving endpoint to log all requests and responses to a Delta table for offline analysis.
B.Increase the Vector Search index's embedding dimension to improve semantic matching.
C.Add a reranker model after Vector Search to reorder retrieved chunks by relevance score.
D.Enable MLflow Tracing on the RAG chain to log each retrieval and generation step with inputs and outputs.
AnswerD

MLflow Tracing captures spans for retrieval and LLM calls, including retrieved chunks and the final response, enabling correlation analysis. This directly addresses the need to see which context influenced the answer and supports systematic evaluation and debugging of the RAG pipeline.

Why this answer

MLflow Tracing instruments each stage of a RAG pipeline, recording retrieved chunks and the LLM's output as linked spans. This gives the engineer the data needed to see which context was passed and potentially used, enabling targeted improvements and evaluation. Other options either change retrieval behavior or log only endpoint-level data, lacking the granular linkage required.

Exam trap

The trap here is assuming that endpoint request logging alone provides enough detail to know which retrieved chunks influenced the answer, when it only captures the final prompt and response.

321
MCQhard

A data engineer is preparing a large corpus of customer support transcripts stored as Parquet in a Unity Catalog volume. Before generating embeddings, each transcript must be tokenized and truncated to a maximum token length. The engineer wants to use a Databricks-native approach that runs distributed across the cluster and avoids pulling the full corpus into a single node. Which approach best satisfies these requirements?

A.Use spark.read.parquet to load the corpus and apply a pandas_udf that tokenizes and truncates each transcript.
B.Use dbutils.fs.cp to copy the Parquet files to a local path, then run a Python tokenizer loop over each file.
C.Create a Delta table from the Parquet files and query it with a SQL UDF that calls a tokenizer from a Python library.
D.Read the Parquet files with pandas.read_parquet on the driver, tokenize, then write the truncated text back to the volume.
AnswerA

A pandas UDF runs vectorized tokenization inside Spark executors, so the corpus stays distributed and each partition processes its own rows. This avoids collecting the data to the driver and scales with cluster size. Tokenizing and truncating in the same distributed step is exactly the preparation needed before embedding generation.

Why this answer

Distributed tokenization of a large corpus requires the work to run inside Spark executors, which a pandas UDF provides by operating vectorized on each partition. Reading with pandas on the driver or looping over files locally centralizes the workload and breaks at scale. A pandas UDF keeps the corpus distributed while applying the tokenizer and truncation per row.

Exam trap

The trap here is equating native Python tokenization with distributed execution, when only executor-side constructs like pandas UDFs keep the corpus distributed.

322
Multi-Selectmedium

A GenAI engineer is deploying a RAG application that uses Databricks Vector Search and a Foundation Model API. The solution must comply with governance policies that require all data access and model invocations to be auditable and access-controlled at a fine-grained level. Which two Unity Catalog features should the engineer leverage to meet these requirements? (Choose two.)

Select 2 answers
A.Unity Catalog volumes to store the model weights and configuration files.
B.Cluster policies to restrict the types of clusters that can access the data.
C.Unity Catalog privileges on the Vector Search index and the source Delta table.
D.Inference tables for the Foundation Model API endpoint to log requests and responses.
E.Databricks SQL dashboards to visualize access patterns and model usage.
AnswersC, D

Unity Catalog privileges on the Vector Search index and the source Delta table allow fine-grained access control. You can grant SELECT on the index to specific groups and restrict access to the underlying table. This ensures that only authorized users can retrieve data and perform searches, meeting the access-control requirement for the RAG application.

Why this answer

Unity Catalog privileges on the Vector Search index and source table provide fine-grained access control, ensuring only authorized users can retrieve data. Inference tables for the Foundation Model API automatically log requests and responses, including user identity, creating an audit trail. Together, these features meet the access-control and auditability requirements for the RAG application.

Exam trap

The trap here is assuming that storing model weights in Volumes or using cluster policies provides governance, but these do not enforce access control or audit model invocations.

323
MCQhard

A team is deploying a LLM-based application using Databricks Model Serving. They want to implement robust observability and monitoring for their endpoint. Which TWO features should they utilize to track performance and quality metrics? (Select TWO)

A.Model Inference Tables
B.Model Serving Monitoring Tab
C.Unity Catalog Lineage
D.Databricks Delta Sharing
E.Workspace-level Audit Logs
AnswerA, B

Inference tables provide a structured way to capture all requests and responses sent to a serving endpoint. This data is written directly into Unity Catalog, allowing developers to perform SQL analysis on input prompts and model outputs to evaluate quality, detect drift, and perform offline debugging of model performance.

Why this answer

To effectively monitor a LLM application, teams must integrate both system-level performance metrics and application-level quality telemetry. Inference tables automatically log input and output data for analysis, while the native 'Monitoring' tab in the Model Serving UI provides latency and throughput metrics. Combined, these features provide a comprehensive view of how the model is performing, identifying bottlenecks in latency or degradation in response quality over time.

Exam trap

Candidates tend to confuse generic cluster monitoring tools with LLM-specific telemetry features, missing that Model Inference Tables and the Serving Monitoring Tab are specifically built for tracking endpoint inputs, outputs, and quality metrics.

324
MCQeasy

A support team operates a Databricks-hosted RAG assistant and wants end-user feedback to feed their monitoring dashboards. They plan to add a thumbs-up and thumbs-down control to the chat UI and log each vote alongside the request ID. What is the primary value of collecting this feedback for the evaluation and monitoring workflow?

A.It guarantees that groundedness scores will improve over time
B.It provides real-world signal that can be joined to traces for triage and dataset curation
C.It replaces the need for any offline evaluation dataset
D.It automatically fine-tunes the underlying foundation model
AnswerB

Thumbs votes are weak labels, but joined to request IDs and traces they point directly at production interactions that went wrong or right. Teams use negatively rated requests to build triage queues and to curate new evaluation examples, closing the loop from production back into offline testing. This is the core value: grounding monitoring and dataset growth in actual user experience.

Why this answer

End-user thumbs votes are weak but real labels tied to specific production requests. Joined with traces and request IDs, they drive triage of failing interactions and curation of new offline evaluation examples, connecting production monitoring to the evaluation dataset. They do not replace datasets, train models automatically, or improve metrics on their own.

Exam trap

The trap here is overvaluing raw user feedback as a complete evaluation or training signal, when it is sparse, biased, and useful mainly for triage and dataset curation.

325
MCQeasy

An engineer is designing a Databricks GenAI application that summarizes customer meeting notes. The notes contain personally identifiable information, and company policy requires that the summarization model never sends note text to an external provider. The team already has a fine-tuned open model registered in Unity Catalog. Which deployment choice satisfies the policy?

A.Serve the registered open model on a Databricks Model Serving endpoint and call it from the application
B.Use Databricks Foundation Model APIs with the pay-per-token option for a supported open model
C.Export the notes to a local workstation, run an open model there, and upload the summaries back
D.Call a third-party foundation model API directly from the notebook and pass the notes in the request body
AnswerA

Hosting the registered open model on a Model Serving endpoint keeps all inference inside the Databricks workspace, so note text never leaves the company's boundary. The endpoint exposes a standard API the application can call, and Unity Catalog governance continues to apply to the model artifact and its access controls.

Why this answer

The policy requires that note text never reach an external provider, so inference must run inside the company's Databricks environment. Serving the registered open model on a Databricks Model Serving endpoint keeps prompts and responses within the workspace while still offering a scalable API, and Unity Catalog continues to govern access to the model.

Exam trap

The trap here is equating a Databricks-hosted model API with in-workspace inference, when only serving the model inside the workspace guarantees the data boundary the policy demands.

326
MCQeasy

Which Databricks feature is specifically designed to allow developers to programmatically manage and version their entire data and AI infrastructure as code?

A.Unity Catalog Volumes
B.Databricks Asset Bundles
C.SQL Warehouses
D.Delta Sharing
AnswerB

Databricks Asset Bundles allow users to define their Databricks resources in code and deploy them using the Databricks CLI. This approach is the standard for infrastructure-as-code in the Databricks ecosystem, providing a unified way to manage complex deployments across different environments with automated CI/CD integration.

Why this answer

Databricks Asset Bundles (DABs) enable developers to package and deploy their projects—including jobs, pipelines, and models—using a declarative, version-controlled approach. This is fundamental for modern MLOps and DataOps workflows, as it replaces manual UI configuration with repeatable, automated deployments. By treating infrastructure as code, teams improve reliability, auditability, and speed when promoting applications from development to production environments.

Exam trap

Candidates often confuse 'Databricks Asset Bundles' with 'Delta Live Tables' or 'Workflows', assuming the latter are the primary tools for infrastructure versioning rather than just data pipeline execution.

327
MCQmedium

Which of the following describes the purpose of using a 'Feature Store' when preparing data for Generative AI applications?

A.To store raw text files for long-term archival.
B.To provide a unified interface for consistent feature computation during training and inference.
C.To serve as a high-performance database for LLM model weights.
D.To replace the need for data cleaning and preprocessing entirely.
AnswerB

The primary goal of a feature store is to eliminate inconsistencies in feature logic. By centralizing this, organizations ensure that the features the model was trained on are computed in exactly the same way when deployed. This is foundational for building stable, repeatable machine learning pipelines in Databricks.

Why this answer

The Feature Store acts as a centralized repository for standardized, reusable features. It ensures that the same logic is used for data preparation during both training and inference. This consistency is critical for preventing training-serving skew, where the model's performance in production differs from its training performance because the input data was transformed differently, ensuring the model remains accurate and reliable in real-world scenarios.

Exam trap

Candidates often mistakenly believe the Feature Store is primarily for model storage, ignoring its primary role in ensuring identical feature computation logic between training and real-time inference pipelines.

328
MCQhard

Refer to the exhibit. The developer is attempting to log a custom model to the Unity Catalog. Which configuration is missing to ensure the model is registered correctly under the specified Unity Catalog location?

A.mlflow.set_registry_uri("databricks")
B.mlflow.set_registry_uri("databricks-uc")
C.mlflow.set_tracking_uri("databricks-uc")
D.mlflow.set_experiment("uc_experiment")
AnswerB

The 'databricks-uc' registry URI tells MLflow to route the model registration process to Unity Catalog. This is the mandatory configuration for using Unity Catalog's model registry features, which are required for cross-workspace model sharing and centralized governance, ensuring the model artifact is properly stored and discoverable.

Why this answer

To log a model to Unity Catalog, the registry must be specified as 'databricks-uc'. The default behavior for 'databricks' tracking URI is to log to the workspace model registry. By explicitly setting the registry to the Unity Catalog destination, the developer ensures the model is governed by the correct security and lineage frameworks, allowing for seamless downstream deployment and access control within the enterprise environment.

Exam trap

Candidates often assume the default MLflow registry URI automatically points to Unity Catalog, failing to realize that an explicit set_registry_uri('databricks-uc') call is required to shift from the workspace registry.

329
MCQeasy

What is the primary purpose of the 'bundle validate' command in the Databricks Asset Bundles CLI?

A.To execute the deployment and verify it in the workspace.
B.To check for schema compliance and configuration errors.
C.To format the YAML files according to Databricks standards.
D.To automatically update the bundle version to the latest release.
AnswerB

This command validates the YAML files against the expected schema for DABs. It checks for missing fields, incorrect types, and invalid resource references. By catching these issues early, developers can fix errors locally before the configuration is pushed to the CI/CD pipeline or the target environment.

Why this answer

The 'bundle validate' command performs a schema check and structural analysis on your configuration files. It ensures that the YAML is syntactically correct and that all referenced resources exist or will be created correctly. This step is crucial for preventing common errors before you attempt a deployment, saving time and avoiding partially successful deployments that could leave the environment in an inconsistent state.

Exam trap

Candidates often confuse bundle validate with bundle deploy, thinking that validation actually pushes changes to the remote workspace instead of only checking local syntax and schema.

330
MCQmedium

A GenAI engineer is building a retrieval-augmented generation (RAG) application using Databricks Vector Search. They notice that the retriever sometimes returns irrelevant chunks that hurt answer quality. They want to add a reranking step to improve the relevance of retrieved documents before passing them to the LLM. Which component should they add to their RAG pipeline?

A.A larger embedding model
B.A cross-encoder reranker
C.A fine-tuned LLM
D.A BM25 retriever
AnswerB

A cross-encoder reranker scores each query-document pair jointly, capturing fine-grained interactions and reordering the initial retrieval results by relevance. In Databricks, you can integrate a reranker model (e.g., from Hugging Face) as a separate component in the RAG chain. This directly addresses the issue of irrelevant chunks by refining the top-k results before generation.

Why this answer

A cross-encoder reranker is specifically designed to reorder retrieved documents by jointly encoding the query and each document, producing a relevance score. This step filters out irrelevant chunks and improves the quality of context passed to the LLM. In Databricks, you can incorporate a reranker as a separate component in a LangChain or custom RAG pipeline, often using a model served via Mosaic AI Model Serving or an external endpoint.

Exam trap

The trap here is assuming that a larger embedding model or a fine-tuned LLM can fix retrieval relevance without a dedicated reranking step.

Page 4

Page 5 of 5

All pages