Courseiva

CCNA Design Applications Questions

57 questions · Design Applications topic · All types, answers revealed

1
MCQeasy

A team is designing a GenAI application that must call an external LLM provider from a Databricks notebook and from a Model Serving endpoint. The security team requires that the provider API key never appear in notebook source code, Git history, or the model artifact, and that the same governed credential be reused by both the notebook and the serving endpoint. Which design satisfies these requirements?

A.Store the API key in a Unity Catalog connection that supports external services, and reference that connection from both the notebook and the Model Serving endpoint.
B.Hard-code the API key as a widget default in the notebook and mirror it into the serving endpoint's environment variables so both components share one value.
C.Store the API key in a Databricks secret scope and reference it with dbutils.secrets.get in the notebook, then pass the resolved value into the serving endpoint's environment variables.
D.Place the API key in a cluster-scoped Spark configuration on the interactive cluster, and have the serving endpoint read the same configuration at request time.
AnswerA

A Unity Catalog connection centralizes credential storage and access control, so the API key never appears in code or artifacts. Both notebooks and Model Serving endpoints can reference the same governed connection, giving consistent secret rotation and auditability. This directly meets the requirement that one governed credential be reused without exposing the plaintext anywhere in source or configuration.

Why this answer

Centralizing the external provider credential in a Unity Catalog connection keeps plaintext out of source code, Git history, and artifacts while letting both notebooks and Model Serving endpoints reference the same governed object. Secret scopes with manual injection, widget defaults, and cluster Spark configurations either leak the value or fail to reach the serving runtime.

Exam trap

The trap here is treating any secret scope reference as sufficient, when the requirement also demands that the same governed credential be consumable by a Model Serving endpoint, which a manually resolved secret does not satisfy.

2
MCQmedium

When designing a production RAG application, which technique is most effective for preventing the LLM from hallucinating based on outdated information?

A.Increasing the temperature parameter of the model to maximum.
B.Enforcing a streaming data pipeline to keep the vector index updated.
C.Restricting the LLM to a specific list of keywords for its output.
D.Adding a long system prompt instructing the model not to lie.
AnswerB

Keeping the vector index synchronized with the source data via a streaming pipeline ensures that the context retrieved during RAG is always current. This minimizes the risk of the model using outdated information, which is a major source of hallucinations in production systems that rely on rapidly changing business data.

Why this answer

Implementing a robust data pipeline that refreshes the vector index in real-time or near-real-time ensures that the retrieved context is current. When coupled with source attribution, this allows the system to verify findings against the latest data. This approach is essential for maintaining accuracy, as stale information in the vector database directly leads to hallucinations that can damage user trust in the AI application.

Exam trap

Candidates often assume that simply increasing the LLM's context window or using a more advanced model will solve hallucinations, ignoring that the root cause is the outdated source data itself.

3
MCQhard

An engineer has a Databricks Vector Search index that stores chunk embeddings for a technical manual. Users report that queries containing exact part numbers return irrelevant chunks because the embedding model blurs numeric tokens. The engineer wants retrieval to consider both dense vector similarity and exact keyword matching on the same Delta table without building a separate search system. Which Databricks Vector Search feature should the engineer enable?

A.A larger embedding model dimension
B.Hybrid search on the index
C.A higher number of returned results per query
D.Delta table change data feed on the source table
AnswerB

Hybrid search in Databricks Vector Search combines dense vector similarity with keyword-based matching over the indexed text, so exact tokens such as part numbers contribute to ranking alongside semantic similarity. Enabling it on the existing index lets the engineer improve recall for exact identifiers without standing up a separate keyword search system or duplicating the source data.

Why this answer

Hybrid search augments dense vector similarity with keyword matching, which restores recall for exact tokens like part numbers while preserving semantic matching for natural-language queries. The other options affect index freshness, capacity, or result volume, but none of them introduce the exact-match signal that the scenario requires.

Exam trap

The trap here is treating retrieval quality as a function of embedding size or result count, when the missing capability is a keyword matching signal that only hybrid search provides.

4
MCQmedium

When designing a RAG application for a regulated industry, what is the most important Unity Catalog feature for ensuring data compliance during retrieval?

A.Catalog-level tagging for organization.
B.Row-level security and column masking.
C.Personal Access Tokens (PATs) for service principals.
D.Automatic data lineage tracking.
AnswerB

Row-level security and column masking are essential for ensuring that users only retrieve authorized information from the source Delta tables. When the index is queried, these Unity Catalog policies apply, ensuring that the RAG application respects organizational access policies, which is mandatory for compliance in regulated industries.

Why this answer

Unity Catalog's attribute-based access control (ABAC) or row-level security allows for granular control over which users can access specific pieces of data. In regulated industries, it is critical that the RAG model only retrieves information that the querying user is explicitly authorized to view, preventing sensitive data leakage and ensuring auditability at the catalog level.

Exam trap

Candidates often choose general encryption or workspace-level access. In regulated industries, security must be applied at the data level (rows/columns) to ensure users only see what they are authorized for.

5
MCQmedium

A generative AI engineer is building a RAG chain on Databricks using LangChain. The chain must call an external LLM provider via a Databricks Model Serving endpoint that proxies the provider, and the engineer wants to avoid hardcoding credentials in notebooks. Which approach should the engineer use to authenticate calls from the chain to the serving endpoint?

A.Pass the provider API key as a plaintext argument in the LangChain ChatOpenAI constructor and rely on notebook access controls to protect it.
B.Use a personal access token generated for a service principal and embed it in the chain's prompt template so the LLM can forward it to the provider.
C.Configure the serving endpoint with the provider credential stored as a secret in the endpoint's environment, and call the endpoint using the Databricks SDK or LangChain's Databricks LLM class with workspace authentication.
D.Store the provider API key in a Databricks secret scope and reference it with dbutils.secrets.get inside the notebook before instantiating the LLM client.
AnswerC

Model Serving endpoints can hold provider credentials as secrets in their environment, so the client only needs a Databricks token. LangChain's Databricks LLM integration and the Databricks SDK both use workspace authentication, keeping the provider key out of notebooks and enabling centralized governance and rotation.

Why this answer

The serving endpoint is the correct place to hold the external provider credential because it centralizes secret storage, rotation, and auditing. Clients then authenticate to the endpoint with workspace identity rather than the provider key. LangChain's Databricks LLM integration and the Databricks SDK both support this pattern, keeping credentials out of notebooks and enabling Unity Catalog governance over the endpoint.

Exam trap

The trap here is assuming the provider API key must be retrieved in the notebook, when the serving endpoint is designed to store and use that credential on the client's behalf.

6
MCQmedium

A GenAI engineer is building a customer-support assistant on Databricks. The assistant must invoke an external LLM endpoint that requires an API token, and the token must not appear in notebook source or logs. The engineer wants to store the token once and let notebooks and Databricks Jobs reference it by name. Which Databricks capability should the engineer use to satisfy this requirement?

A.A Unity Catalog volume mounted at /Volumes/secure/token.txt
B.A workspace environment variable set in the cluster Spark configuration
C.A personal access token stored in the Databricks CLI configuration file
D.Databricks secrets backed by a secret scope
AnswerD

Databricks secrets store sensitive values in a secret scope and expose them only through dbutils.secrets.get or secret references, so the token never appears in notebook source, logs, or job definitions. This satisfies the requirement of storing the token once and referencing it by name from notebooks and jobs while keeping it out of version control.

Why this answer

Databricks secrets provide a governed, auditable store for sensitive values and expose them through dbutils.secrets.get or secret references, so the external LLM token is never embedded in code or logs. Other storage locations either expose plaintext to readers or serve a different purpose, such as authenticating to the workspace rather than to an external provider.

Exam trap

The trap here is assuming any file or variable accessible from a notebook is safe for credentials, when only a purpose-built secret store redacts values and enforces access control and auditing.

7
MCQhard

In a RAG application, which architectural component ensures that the system handles changes in the source data effectively?

A.The LLM's internal knowledge base.
B.The Delta Lake Change Data Feed (CDF).
C.The prompt engineering template.
D.The user session cache.
AnswerB

CDF provides a structured way to track changes in Delta tables. By leveraging CDF, the vector search index can incrementally ingest only the delta between the old and new states of the data. This is significantly more efficient than full re-indexing and ensures high data consistency with minimal compute cost.

Why this answer

The Delta Lake Change Data Feed (CDF) or the automatic synchronization of Vector Search indexes is designed specifically to track and apply changes (inserts, updates, deletes) to downstream systems. This ensures that the vector representation of the data is always in sync with the source of truth, avoiding the need for expensive full re-indexes and ensuring that retrieved information is always up-to-date.

Exam trap

Candidates often select manual scheduling tools or general caching mechanisms instead of recognizing Delta Lake Change Data Feed as the native engine tracking incremental table changes.

8
Multi-Selectmedium

Which TWO factors should be considered when choosing an embedding model for a RAG application?

Select 2 answers
A.The model's native dimensionality and its effect on vector index storage size.
B.The model's ability to generate creative, hallucinatory content.
C.The compatibility of the model with the target languages of the documents.
D.The maximum number of characters allowed in a single prompt.
E.The color profile of the documents being ingested.
AnswersA, C

Higher dimensionality embeddings provide better semantic resolution but significantly increase memory usage and storage costs. Choosing the right dimensionality is a trade-off between the precision of the similarity search and the infrastructure overhead required to maintain the index, which directly impacts the scalability and cost of the RAG application.

Why this answer

The choice of embedding model determines the quality of semantic retrieval and the computational requirements of the system. By aligning the model's language support, dimensionality, and performance with the application's needs, you ensure a balanced design. It is also vital to consider the model's licensing and whether it needs to be fine-tuned to capture domain-specific terminology that general models might misinterpret.

Exam trap

Candidates often focus solely on model accuracy. They neglect the practical constraints of vector index storage size and language compatibility, which are critical for production scalability and performance.

9
MCQmedium

An engineer is designing a GenAI application where a Databricks workflow must call an external LLM provider that enforces a strict rate limit and occasionally returns transient errors. The engineer wants the application to degrade gracefully instead of failing the whole job. Which design choice best addresses this requirement?

A.Increase the driver node type on the job cluster
B.Cache all prompts and responses in a Delta table before calling the provider
C.Set the job to retry the entire task on failure
D.Wrap the external call in retry logic with exponential backoff and a fallback response path
AnswerD

Retries with exponential backoff absorb transient rate-limit and server errors without immediately failing the job, and a fallback path lets the application return a degraded but useful response when retries are exhausted. Together they prevent a single external hiccup from aborting the entire workflow.

Why this answer

Resilience to external rate limits and transient failures comes from handling the call itself: bounded retries with exponential backoff smooth over temporary rejections, and a fallback response keeps the application functional when the provider stays unavailable. Infrastructure sizing, task-level retries, and response caching do not address the failure at its source.

Exam trap

The trap here is reaching for cluster or job-level scaling to solve an external dependency problem, when the failure originates at the API boundary and must be handled in the calling code.

10
MCQmedium

A team is deploying a Databricks RAG chatbot that must serve interactive traffic with low latency while also allowing the data science team to test prompt variations safely. The engineer wants the production endpoint to keep serving stable traffic even while a new prompt template is being evaluated. Which approach best meets these requirements?

A.Lower the production endpoint's concurrency limit and run the experiment on the same endpoint during off-peak hours
B.Edit the prompt template in the production endpoint's registered model version and restart the endpoint
C.Add the new prompt as a branch inside the production chain and select it at runtime using a random number generator
D.Create a separate Model Serving endpoint for the experimental prompt and route a small percentage of traffic to it for comparison
AnswerD

Isolating the experimental prompt on its own serving endpoint lets the team evaluate it with real or synthetic traffic without risking the production endpoint's stability. Traffic can be split at the application layer for A/B comparison, and if the experiment degrades, production remains unaffected because the two endpoints scale and version independently.

Why this answer

Separating the experimental prompt onto its own Model Serving endpoint preserves production stability while enabling controlled comparison. The team can direct a small share of traffic to the experimental endpoint, measure quality and latency, and promote the prompt only after it proves better, all without redeploying or interrupting the production endpoint.

Exam trap

The trap here is assuming a single endpoint can host both stable and experimental prompts safely, when shared capacity and shared code paths let an experiment degrade production traffic.

11
Multi-Selectmedium

Which THREE strategies improve the quality of retrieval in a Databricks Vector Search-based RAG application?

Select 3 answers
A.Implementing hybrid search using both semantic vectors and keyword-based filtering.
B.Increasing the chunk size to include the entire dataset in a single vector.
C.Adding relevant metadata tags to documents to enable targeted filtering.
D.Optimizing the chunking strategy to maintain context boundaries.
E.Removing all stop words from the documents during the ingestion phase.
AnswersA, C, D

Hybrid search combines the strengths of semantic understanding with the precision of keyword matching. This ensures that specific technical terms or IDs are found accurately while still capturing the intent behind user queries, significantly improving retrieval quality for complex domain-specific datasets where semantic similarity alone might lead to irrelevant results.

Why this answer

Improving retrieval quality requires a combination of high-quality data chunking, metadata filtering, and optimized search parameters. These strategies directly influence the relevance and accuracy of retrieved context, which is the foundation of a successful RAG application. Mastering these techniques allows engineers to fine-tune system behavior to meet specific user needs and reduce hallucinations during the generation phase.

Exam trap

Candidates often assume that only vector embeddings matter, ignoring how hybrid search and metadata filtering are essential for combining semantic relevance with precise keyword matches to boost retrieval quality.

12
MCQeasy

Which Databricks component is recommended for orchestrating the end-to-end RAG pipeline, including data ingestion, transformation, and vector indexing?

A.Databricks SQL.
B.Databricks Workflows.
C.The Databricks File System (DBFS).
D.Unity Catalog.
AnswerB

Databricks Workflows provides the orchestration layer needed to schedule and manage tasks like data ingestion, transformation, and vector index updates. Its ability to manage dependencies between tasks and provide monitoring is critical for keeping the RAG pipeline functional, up-to-date, and reliable in a production environment.

Why this answer

Databricks Workflows allows for the scheduled and triggered execution of tasks, such as notebook runs or Delta Live Tables pipelines. This is the recommended way to orchestrate complex RAG pipelines where data must be cleaned, chunked, and indexed systematically. By using Workflows, engineers ensure that the entire pipeline is reproducible, monitorable, and resilient to failures, which is essential for maintaining a production-grade AI application.

Exam trap

Candidates tend to confuse cluster management tools with orchestration tools, selecting cluster settings instead of Databricks Workflows for running multi-step ETL and RAG pipelines.

13
Multi-Selecthard

A team is designing a Databricks GenAI application that must return grounded answers with citations to source documents. The application uses Databricks Vector Search for retrieval and a Foundation Model API for generation. Which TWO design choices are required to return accurate citations alongside each answer? (Choose two.)

Select 2 answers
A.Instruct the model in the prompt to reference the provided chunks by their identifiers and include those identifiers in the response.
B.Set the model's temperature to zero to guarantee that citations are factually correct.
C.Increase the embedding model's dimensionality to the maximum supported value so retrieval is more precise.
D.Enable Foundation Model API pay-per-token billing so the model has access to citation formatting.
E.Store source metadata such as document ID, title, and chunk position alongside each embedding in the Vector Search index.
AnswersA, E

Even with metadata available, the model must be told to attribute its statements to specific chunks. Prompting the model to cite chunk identifiers, and structuring the prompt so each chunk carries a visible ID, gives the model the information it needs to ground its answer and lets the application resolve those IDs into human-readable citations.

Why this answer

Citations require two things working together: source metadata attached to each retrieved chunk so the application knows where content came from, and prompt instructions that tell the model to attribute statements to specific chunk identifiers. Precision tuning, billing mode, and temperature settings do not supply the provenance information or the attribution behavior needed to render citations.

Exam trap

The trap here is assuming that better retrieval quality automatically yields citations, when provenance metadata and explicit attribution instructions are what actually enable them.

14
MCQmedium

An organization needs to build a RAG application on Databricks that minimizes data egress and maximizes security by keeping all data within the workspace perimeter. Which architectural pattern best satisfies this requirement?

A.Call external LLM APIs from a standard Python notebook without VPC constraints.
B.Export data to an external vector database and use a cloud-hosted LLM.
C.Deploy an embedding model on Mosaic AI Model Serving and utilize Databricks Vector Search.
D.Use a public LLM endpoint with a public bucket to store the vector index.
AnswerC

This approach keeps all data processing, storage, and inference within the Databricks workspace perimeter. By using internal serving for embeddings and native vector search capabilities, the architecture minimizes egress traffic and simplifies security policy enforcement, ensuring that data never leaves the protected environment during the RAG retrieval process.

Why this answer

Utilizing Mosaic AI Model Serving with private endpoints and leveraging Vector Search indexes ensures that both the embedding model and the retrieval process occur within the Databricks control plane. By avoiding external API calls to third-party providers, the organization maintains strict governance, data residency compliance, and lower latency for inference, which is critical for enterprise-grade generative AI applications handling sensitive corporate documents.

Exam trap

Candidates often suggest external API-based embedding models, ignoring the requirement to keep data within the workspace perimeter to minimize egress and satisfy strict residency requirements.

15
Multi-Selectmedium

Which TWO factors should be prioritized when selecting an embedding model for a domain-specific RAG application on Databricks?

Select 2 answers
A.The total number of parameters in the model regardless of the domain.
B.The semantic relevance of the model to the target domain's terminology.
C.The availability of the model on the public Hugging Face repository.
D.The computational resource requirements for inference latency.
E.The color scheme of the model's documentation page.
AnswersB, D

Embedding models must map domain-specific terms to accurate vector representations to ensure relevant document retrieval. A model that has not been trained or fine-tuned on the specific domain’s jargon will produce poor vector alignments, leading to inaccurate RAG responses, regardless of the model's performance on general-purpose benchmarks.

Why this answer

Selecting the right embedding model requires balancing semantic accuracy within the specific domain and the operational cost of maintaining the model. By prioritizing domain-specific performance and resource efficiency, teams ensure that the RAG pipeline provides relevant search results without incurring excessive latency or compute costs, which are foundational for sustaining long-term generative AI production workloads.

Exam trap

Candidates often prioritize model popularity or parameter count, ignoring the critical balance between domain-specific semantic relevance and the operational latency costs of the model.

16
MCQeasy

A team is deploying a RAG chatbot that answers questions from internal policy documents. The documents change frequently, and the team wants retrieval to reflect updates within minutes without re-running a full embedding job. The source Delta table already has change data feed enabled. Which approach should the team use to keep the Databricks Vector Search index current?

A.Increase the cluster size of the vector search endpoint
B.Configure the index to sync from the Delta table so changes are propagated incrementally
C.Drop and recreate the vector search index on a schedule
D.Add a metadata filter for the document modification timestamp
AnswerB

A vector search index that syncs from a Delta table with change data feed enabled picks up inserts, updates, and deletes incrementally, so newly modified policy chunks become searchable shortly after they land. This meets the requirement of reflecting updates within minutes without recomputing embeddings for the entire corpus.

Why this answer

A Delta-synced vector search index consumes change data feed events to apply inserts, updates, and deletes incrementally, keeping embeddings aligned with the source table. Rebuilding the index or tuning the endpoint addresses cost and throughput rather than freshness, and metadata filters only narrow an already-indexed set.

Exam trap

The trap here is confusing index freshness with endpoint capacity, when freshness depends on how changes flow from the Delta table into the index rather than on how much compute serves queries.

17
MCQhard

An engineer is designing a Databricks RAG application where the same retrieved context must be reused across several prompt variants during evaluation. They want to reduce token cost and improve consistency between variants. Which design choice best supports this?

A.Use a higher temperature during evaluation so the model explores more responses and the best one can be selected.
B.Compute the retrieval step once, store the retrieved documents in a Delta table keyed by query, and have each prompt variant read from that table.
C.Configure the model serving endpoint to cache identical requests so repeated prompts with the same context are served from cache.
D.Increase the chunk size in the Vector Search index so that each retrieval returns more context, reducing the need to call the retriever multiple times.
AnswerB

Caching retrieval results in a Delta table decouples retrieval from generation, so multiple prompt variants consume the same context without repeating embedding lookups or Vector Search calls. This reduces token and compute cost and ensures every variant is evaluated against identical context, which makes comparisons fair and reproducible.

Why this answer

Separating retrieval from generation and persisting the retrieved context lets several prompt variants share one retrieval result. This lowers cost because embeddings and Vector Search are invoked once, and it improves consistency because all variants see the same evidence. The alternatives change generation or chunking behavior but do not eliminate redundant retrieval.

Exam trap

The trap here is confusing generation-side caching with retrieval reuse, when the cost and consistency problem originates in repeated retriever calls.

18
MCQhard

A generative AI engineer is designing a multi-stage RAG application on Databricks. The application first retrieves documents using Vector Search, then reranks them with a cross-encoder model, and finally calls a foundation model endpoint to generate an answer. The engineer wants to ensure that the entire pipeline is reproducible and that each stage can be independently versioned and deployed. Which design approach best meets these requirements?

A.Package each stage as an MLflow model with its own signature and dependencies, and compose them into a single MLflow pipeline that can be logged and served as one model.
B.Implement the pipeline as a single Python script that calls all stages sequentially, and version the script in Git.
C.Use MLflow Projects to define each stage as a separate project, and orchestrate them with a Databricks Job that passes artifacts between stages.
D.Deploy each stage as a separate Databricks Model Serving endpoint and have the client application call them in sequence.
AnswerA

MLflow models encapsulate code, environment, and signatures, enabling independent versioning. Composing them into a pipeline allows the entire multi-stage application to be logged as a single model, which can be served with Databricks Model Serving. This provides reproducibility and independent versioning of each stage while offering a unified deployment artifact. It is the recommended pattern for complex generative AI pipelines.

Why this answer

Packaging each stage as an MLflow model with its own signature and dependencies allows independent versioning. Composing them into a single MLflow pipeline enables logging and serving the entire multi-stage RAG application as one model, ensuring reproducibility. This approach is supported by Databricks Model Serving and MLflow, and it simplifies deployment and tracing compared to separate endpoints or scripts.

Exam trap

The trap here is assuming that separate endpoints or a single script provide sufficient versioning and reproducibility, when the key is to encapsulate each stage as an MLflow model and compose them into a unified pipeline.

19
MCQmedium

A GenAI engineer is designing a multi-turn chat application on Databricks. Users report that the assistant forgets details from earlier in long conversations and sometimes answers using only the most recent message. The team wants the model to reliably use facts stated several turns earlier without exceeding the model's context window. Which design should the engineer implement?

A.Implement a conversation memory strategy that summarizes or selectively retains earlier turns and injects the compressed history into each prompt within the context window.
B.Store the full conversation transcript in a Delta table and instruct the model to query the table during inference.
C.Increase the model's temperature so it explores more of the conversation history when generating a response.
D.Send only the latest user message to the model and rely on the model's pretrained knowledge to recall earlier facts.
AnswerA

A memory strategy that summarizes or selectively retains earlier turns keeps essential facts available while staying within the context window. By injecting the compressed history into each prompt, the model receives the earlier details it needs, which directly addresses the reported forgetting and avoids the cost and truncation problems of sending the full transcript.

Why this answer

Reliable multi-turn memory requires bringing earlier relevant turns into the prompt, either by summarizing them or by selectively retaining key facts. This keeps the payload within the context window while preserving details the user expects the assistant to remember. Temperature changes, sending only the latest message, or storing transcripts without retrieval all fail to put earlier facts in front of the model.

Exam trap

The trap here is assuming the model has inherent memory across API calls, when in fact each inference request is stateless and only sees the prompt it is given.

20
Multi-Selecthard

A GenAI engineer is designing a Databricks application that must ground answers in a large corpus of internal policy documents. The corpus is updated by a nightly Delta job, and the application must cite the source document for every answer. The engineer is deciding how to structure the retrieval and generation stages. Which TWO design choices best satisfy the grounding and citation requirements? (Choose two.)

Select 2 answers
A.Instruct the model to answer only from the retrieved context and to include the provided document identifiers in its response
B.Rely on the model's pretrained knowledge of internal policies to fill gaps when retrieval returns no relevant chunks
C.Store chunk text and source metadata such as document ID and section as columns alongside the embeddings in the Vector Search index
D.Cache the full corpus in the prompt on every request so the model always has complete policy context
E.Increase the model's temperature setting so the generation stage explores a wider range of possible answers
AnswersA, C

Constraining the prompt to the retrieved context reduces unsupported generation, and requiring the model to echo the supplied identifiers ties each statement back to a retrieved chunk. Because the identifiers originate from the index metadata rather than the model's memory, the citations remain verifiable against the source documents.

Why this answer

Grounding and citation depend on two things working together: retrieval that returns source identifiers with each chunk, and a generation prompt that restricts the model to that retrieved context while requiring it to surface the identifiers. Metadata columns supply verifiable provenance, and the constrained instruction keeps the answer tied to retrieved policy text instead of the model's own memory.

Exam trap

The trap here is believing a capable model can supply citations from memory, when only metadata carried through retrieval can produce citations that are verifiable against the source documents.

21
MCQmedium

A GenAI engineer is building a RAG chatbot backed by a Databricks Vector Search index that is synced from a Delta table with Change Data Feed enabled. Source documents are frequently updated, and some are deleted. The engineer wants the index to reflect deletions and updates automatically with minimal operational overhead. Which configuration should be used?

A.Create a Direct Vector Access index and write a Structured Streaming job that upserts embeddings into it whenever the source table changes.
B.Create a Delta Sync index and enable Change Data Feed on the index itself so that deletions flow back into the source Delta table.
C.Create a Delta Sync index with a pipeline_type of TRIGGERED and schedule a Databricks Job to run a full refresh every 15 minutes.
D.Create a Delta Sync index with a pipeline_type of CONTINUOUS, pointing at the source Delta table, and let Databricks manage synchronization.
AnswerD

A CONTINUOUS Delta Sync index uses the source table's Change Data Feed to incrementally propagate inserts, updates, and deletes to the vector index without manual refreshes. This satisfies the requirement for automatic reflection of deletions and updates while minimizing operational overhead, because Databricks manages the sync pipeline and embedding computation end to end.

Why this answer

Automatic reflection of inserts, updates, and deletes from a Delta source into a Vector Search index is achieved with a Delta Sync index running in CONTINUOUS pipeline mode, which consumes the source table's Change Data Feed. Triggered refreshes, Direct Vector Access, and index-side Change Data Feed all either add manual work, introduce staleness, or rely on unsupported behavior.

Exam trap

The trap here is assuming that a scheduled full refresh or a Direct Vector Access index provides the same incremental delete handling as a continuous Delta Sync index.

22
MCQhard

A team is deploying a RAG application using Databricks Model Serving with a foundation model endpoint and a Databricks Vector Search index. During load testing, they observe that p95 latency spikes when the retriever returns many chunks, and the LLM occasionally truncates context. They want to reduce latency while preserving answer quality. Which change is most appropriate?

A.Increase the number of retrieved chunks and rely on the LLM to ignore irrelevant content.
B.Switch the Vector Search index to a smaller embedding dimension and keep the same number of chunks.
C.Apply a reranker to the retrieved chunks and pass only the top few highest-scoring chunks to the LLM.
D.Lower the LLM temperature to zero to reduce response time.
AnswerC

A reranker scores retrieved chunks against the query with higher precision than the initial vector similarity, so the application can pass fewer but more relevant chunks to the LLM. This shortens the prompt, reducing latency and the chance of context truncation, while preserving answer quality because the most relevant evidence is retained. It directly addresses both symptoms observed during load testing.

Why this answer

Reranking retrieved chunks and passing only the top few to the LLM reduces prompt length, which lowers LLM latency and mitigates context truncation. Because the reranker prioritizes the most query-relevant chunks, answer quality is preserved even with fewer chunks. The other options either increase prompt size, change an unrelated parameter, or alter the index without addressing prompt length.

Exam trap

The trap here is treating temperature as a performance lever, when latency in RAG is dominated by prompt length and the volume of retrieved context, not by sampling randomness.

23
MCQeasy

What is the primary benefit of using Unity Catalog when designing generative AI applications in Databricks?

A.It automatically generates Python code for model training.
B.It provides centralized governance, lineage, and access control for data assets.
C.It converts unstructured text into vectors automatically.
D.It eliminates the need for data preprocessing before RAG.
AnswerB

Unity Catalog enables secure, governed access to all data used in the AI lifecycle. By tracking data lineage from the source to the model, it ensures transparency and compliance. This centralized approach simplifies security management and audit readiness, which are essential when handling proprietary data in generative AI applications.

Why this answer

Unity Catalog provides a unified governance and security layer, ensuring that data used for training, fine-tuning, or RAG is access-controlled and lineage-tracked. This is crucial for compliance and reproducibility in AI projects. By centralizing permissions and audit logs, teams can securely manage access to sensitive data across their entire generative AI pipeline, which is a fundamental requirement for enterprise AI adoption.

Exam trap

Candidates often focus on the performance benefits of Unity Catalog, overlooking its primary enterprise role as a centralized governance and security layer for data assets.

24
MCQeasy

An engineer is building a GenAI application that must return structured JSON output conforming to a specific schema so downstream systems can parse it reliably. The team wants to enforce the schema at generation time rather than post-processing free-form text. Which approach should they use?

A.Fine-tune the foundation model on examples of the desired JSON output.
B.Post-process the model output with a regular expression that extracts the first JSON object.
C.Set the LLM temperature to zero and instruct the model to return JSON in the system prompt.
D.Use a structured output mode or response format that constrains generation to the provided JSON schema.
AnswerD

Constrained decoding against a JSON schema, exposed by foundation model APIs as a structured output or response format setting, restricts token generation so the output must conform to the schema. This enforces field names, types, and required properties at generation time, eliminating the need for brittle post-processing. It is the appropriate mechanism when downstream systems require reliably parseable structured output.

Why this answer

Structured output or response format settings constrain the model's decoding process to the supplied JSON schema, guaranteeing that generated content parses and conforms. This is more reliable than prompt instructions, fine-tuning, or regex extraction because it operates at generation time and eliminates malformed or incomplete output. Downstream systems can then parse results without defensive fallback logic.

Exam trap

The trap here is assuming that a well-crafted prompt plus temperature zero guarantees valid JSON, when only constrained decoding against a schema enforces structural conformance.

25
MCQmedium

A financial services company is designing a RAG assistant that must never return answers containing personally identifiable information from its knowledge base. The team plans to filter retrieved chunks before they are placed into the prompt. Which Databricks design element should they use to enforce this filtering consistently?

A.Add a system prompt instructing the model to ignore any personally identifiable information it sees in the context.
B.Store a sensitivity label or metadata field on each chunk in the Vector Search index and apply a metadata filter at query time to exclude restricted chunks.
C.Run a post-processing regex on the model output to remove patterns that resemble personally identifiable information.
D.Grant the serving endpoint a service principal that lacks access to the sensitive tables so retrieval returns no restricted rows.
AnswerB

Vector Search supports metadata filtering, so tagging chunks with a sensitivity attribute and filtering at query time prevents restricted content from ever entering the prompt. This enforces the policy at retrieval, which is earlier and more reliable than post-generation redaction, and it scales as new documents are labeled.

Why this answer

Metadata filtering in Vector Search allows the application to exclude chunks based on attributes such as sensitivity level before they reach the prompt. This enforces the policy at the retrieval boundary, reducing both leakage risk and token usage. Prompt instructions and output regexes act later and cannot guarantee that sensitive content never enters the model context.

Exam trap

The trap here is relying on prompt instructions or output filters for data protection, when the reliable control point is filtering retrieved chunks before prompt assembly.

26
MCQmedium

Which design pattern is best for protecting the LLM from prompt injection attacks when building a customer-facing chatbot on Databricks?

A.Embedding all user inputs into the vector database.
B.Using a guardrail service to validate and sanitize user inputs.
C.Relying on the LLM’s internal safety training to reject injections.
D.Sending all inputs through a standard SQL query first.
AnswerB

Guardrail services or classification models act as a security gateway, analyzing inputs for patterns indicative of injection attacks. By blocking or sanitizing these inputs before they are passed to the primary LLM, the system prevents unauthorized instructions from executing, which is a foundational requirement for securing public-facing chatbots.

Why this answer

Implementing an input-filtering layer using a separate, lightweight classification model or guardrail service effectively detects and blocks malicious prompts before they reach the LLM. This proactive defense is vital for securing AI applications, as prompt injection can lead to unauthorized data access or malicious behaviors. This layering approach creates a defense-in-depth architecture that keeps the generative model safe from adversarial inputs.

Exam trap

Candidates often assume that prompt engineering or system instructions are sufficient to prevent injection, forgetting that malicious users can easily bypass these through prompt manipulation.

27
Multi-Selecthard

A team is designing a Databricks RAG application that must return answers with citations to source documents. They plan to use Databricks Vector Search and want the LLM to reference specific chunks. Which two design choices are required to produce reliable citations? (Choose two.)

Select 2 answers
A.Increase the LLM temperature to encourage the model to generate more detailed citations.
B.Store a stable document identifier and chunk metadata alongside each embedding in the Vector Search index.
C.Enable inference tables on the Vector Search endpoint to capture citation metadata.
D.Set the Vector Search index to use a cosine similarity metric instead of dot product.
E.Include the retrieved chunk identifiers and source references in the prompt context so the LLM can cite them in its response.
AnswersB, E

Citations require mapping a retrieved chunk back to its source document. Including a stable document ID and chunk metadata in the index payload allows the retriever to return that identifier with each result. Without this, the LLM cannot reference a specific source, and citations become guesses or are omitted entirely.

Why this answer

Reliable citations require two things: the retriever must return source identifiers, and those identifiers must reach the LLM in the prompt. Storing document IDs and chunk metadata in the Vector Search index enables the first, and including them in the prompt context enables the second. Similarity metrics, temperature, and inference tables do not provide the linkage needed for citations.

Exam trap

The trap here is focusing on retrieval tuning knobs such as similarity metrics while overlooking that citations depend on carrying source metadata from the index into the prompt.

28
MCQmedium

Refer to the exhibit. The Databricks Vector Search index is currently in an 'ONLINE' status with incremental sync. If a developer performs a DELETE operation on the source Delta table, what happens to the vector index?

A.The record remains in the index until a full re-index is manually triggered.
B.The system throws an error because vector indexes do not support deletes.
C.The record is automatically removed from the index during the next sync cycle.
D.The index enters a 'PENDING_DELETE' state and stops accepting queries.
AnswerC

Incremental sync mode ensures that changes to the Delta table, including deletions, are propagated to the vector index. By monitoring the Delta log, the service identifies the deleted record and updates the index accordingly, keeping the vector store accurate with minimal latency without requiring manual maintenance.

Why this answer

When using incremental synchronization, the Databricks Vector Search service monitors the Delta table's transaction log for changes. A delete operation recorded in the Delta log will trigger a corresponding removal in the index. This ensures the index remains a true representation of the source data, which is critical for compliance and data quality, ensuring the RAG application does not retrieve outdated or removed sensitive information.

Exam trap

Candidates assume deletions require a full manual re-index of the vector store, forgetting that incremental sync automatically processes table transaction logs.

29
MCQeasy

An engineer is designing a GenAI application that must call an external LLM provider's API. The provider key must not appear in notebook code, job logs, or Git. The team already uses Databricks and wants the key to be injected into the serving endpoint at runtime. Which Databricks capability should the engineer use?

A.Hard-coding the provider key in the model signature so the endpoint receives it as a request parameter.
B.Databricks secrets stored in a secret scope, referenced from the serving endpoint configuration using secret references.
C.A Unity Catalog volume containing a text file with the provider key, mounted into the serving endpoint.
D.Environment variables defined in the cluster's Spark configuration and read by the notebook at runtime.
AnswerB

Databricks secret scopes store credentials in an encrypted backend, and secret references let the serving endpoint configuration pull the key at runtime without exposing it in code, logs, or Git. This directly satisfies the requirement that the provider key never appear in notebook code or job logs while still being available to the endpoint.

Why this answer

Databricks secret scopes with secret references keep credentials encrypted and out of source control, logs, and notebooks, while still allowing Model Serving to inject them at runtime. Spark config values, Unity Catalog volume files, and model signature parameters all expose the key through metadata, logs, or request payloads and do not provide runtime injection for serving endpoints.

Exam trap

The trap here is confusing governed storage such as Unity Catalog volumes with secret management, even though volumes are not designed to protect credentials from readers.

30
MCQmedium

Which Databricks feature is specifically designed to monitor model quality and drift in production?

A.Unity Catalog audit logs.
B.MLflow Model Monitoring.
C.The Databricks SQL query history.
D.Cluster event logs.
AnswerB

MLflow Model Monitoring is the dedicated Databricks component for tracking the health of deployed models. It allows engineers to monitor metrics, detect performance degradation, and identify drift in model outputs. This is essential for maintaining high-quality generative AI applications and responding effectively to changes in data or user behavior.

Why this answer

MLflow Model Monitoring (now part of Mosaic AI) provides the necessary tooling to track model performance metrics and identify drift over time. For generative AI, this is critical because language models can exhibit performance degradation or shifts in behavior when input distributions change. Proactive monitoring allows engineers to detect these shifts early and trigger retraining or fine-tuning, ensuring consistent performance in the production environment.

Exam trap

Candidates often choose generic logging tools or Unity Catalog governance features when asked specifically about tracking production model performance drift.

31
MCQeasy

A data science team wants to expose a RAG chain as a REST API so that an external web application can send questions and receive answers. The chain is developed with Databricks LangChain integrations and must be deployed with autoscaling and built-in monitoring. Which Databricks capability should they use?

A.A Databricks job that runs the chain on a schedule and writes answers to a Delta table.
B.Model Serving with a custom MLflow pyfunc model that wraps the chain.
C.An all-purpose cluster with a Flask app started manually on the driver node.
D.A Databricks SQL warehouse with a query that calls the chain through a UDF.
AnswerB

Databricks Model Serving deploys MLflow models as scalable REST endpoints and supports custom pyfunc models, which can wrap a LangChain chain. It provides autoscaling, request logging, and integration with inference tables for monitoring, matching the requirement to serve the chain as an API with operational visibility.

Why this answer

Model Serving is the Databricks capability for hosting models and custom Python logic as REST endpoints with autoscaling and monitoring. Wrapping the LangChain chain in an MLflow pyfunc model makes it deployable through that service, so the web application can call a stable API. Scheduled jobs, SQL warehouses, and manual Flask apps do not provide the same managed serving characteristics.

Exam trap

The trap here is equating any Databricks compute that can run Python with a production API endpoint, overlooking that Model Serving is the managed serving layer.

32
Multi-Selecthard

A generative AI engineer is designing a RAG application on Databricks that uses a foundation model served via Databricks Model Serving. The application must handle peak loads gracefully and provide consistent response times. The engineer is evaluating design patterns for scaling and reliability. Which TWO design choices should the engineer implement? (Choose two.)

Select 2 answers
A.Configure the Model Serving endpoint with autoscaling enabled and set a minimum and maximum number of concurrent requests per replica.
B.Deploy the model to a single large GPU replica with maximum memory to handle all traffic without autoscaling.
C.Implement client-side retries with exponential backoff and a timeout for each request to the serving endpoint.
D.Precompute all possible user questions and their answers and store them in a lookup table for instant retrieval.
E.Cache the LLM responses for identical user queries in a Delta table and serve them without invoking the model.
AnswersA, C

Autoscaling allows the endpoint to add or remove replicas based on load, which helps maintain consistent latency during peaks. Setting concurrency bounds per replica controls how many requests each replica handles, preventing overload. Together they balance throughput and response time. This is a supported configuration for Databricks Model Serving and directly addresses the reliability and scaling requirements.

Why this answer

To handle peak loads and maintain consistent response times, the endpoint must scale elastically and the client must handle transient failures. Autoscaling with concurrency limits per replica ensures the serving infrastructure adapts to load, while client-side retries with backoff and timeouts improve resilience. These two choices together provide both server-side elasticity and client-side fault tolerance, which are essential for production reliability.

Exam trap

The trap here is focusing on caching or precomputation as primary scaling strategies, when the core requirements are elastic endpoint scaling and resilient client behavior under load.

33
MCQmedium

Refer to the exhibit. What is the correct way to log a custom RAG chain so it can be loaded using the provided code?

A.mlflow.log_artifact(local_path="chain.py")
B.mlflow.pyfunc.log_model(artifact_path="my_model", python_model=my_chain_object)
C.mlflow.register_model("path/to/chain")
D.mlflow.sklearn.log_model(my_chain_object)
AnswerB

This method correctly logs a custom Python object (the RAG chain) as an MLflow model. By defining the chain as a Python model, it becomes compatible with the MLflow Model Serving platform, allowing for seamless deployment to production endpoints where it can be invoked via standard API calls.

Why this answer

To load a custom chain using the 'models:/' URI, the chain must be logged as an MLflow model. Using 'mlflow.pyfunc.log_model' with the appropriate artifact parameters allows the entire RAG pipeline—retrieval logic, prompt handling, and LLM inference—to be captured as a single deployable unit. This ensures consistency between local development and production serving environments.

Exam trap

Candidates often confuse 'log_model' with 'save_model' or try to log raw files manually. Using 'mlflow.pyfunc.log_model' is the specific standard for capturing custom RAG chains as deployable artifacts.

34
MCQmedium

A generative AI engineer is designing a Databricks RAG application that ingests a Delta table containing 400 million support articles into a Databricks Vector Search index. The team wants the lowest-latency online serving with the smallest possible index while preserving retrieval quality for the most common queries. Which design decision best meets these requirements?

A.Store all 400 million embeddings in a Unity Catalog volume and query them directly with a Python UDF at request time.
B.Replace Databricks Vector Search with a keyword-only BM25 index over the raw article text to avoid embedding storage costs.
C.Use Databricks Vector Search with a Delta Sync index, and configure a smaller index by filtering out low-traffic or stale articles before syncing so only high-value documents are embedded.
D.Enable the vector index with the vector search endpoint sized for the full 400 million rows and keep all embeddings in a single index.
AnswerC

Filtering the source Delta table to high-value, frequently accessed articles before the Delta Sync index is built reduces the number of embeddings stored and searched, which lowers latency and index size. Vector Search Delta Sync indexes stay current with the source table, so the team can prune low-traffic content while keeping retrieval quality for the common queries the business cares about.

Why this answer

Pruning low-value content before building a Delta Sync index directly reduces the number of embeddings stored and searched, which shrinks the index and lowers serving latency. Because Vector Search Delta Sync indexes continue to track the source Delta table, the team retains freshness and can still serve high-quality results for the queries that matter most, matching both the performance and quality goals.

Exam trap

The trap here is assuming that a RAG application must always embed every row of the source table, when pre-filtering the Delta table before the Delta Sync index is created is a valid way to shrink the index and improve latency.

35
MCQhard

A GenAI engineer is building a Databricks RAG application where the retrieval step returns the top-5 chunks for each user question. The engineer wants to add a second LLM call that evaluates whether each retrieved chunk contains enough information to answer the question, and then filters out chunks that fail this evaluation before passing the remaining chunks to the final answer-generation prompt. Which design pattern is the engineer implementing?

A.Chain-of-thought prompting
B.Retrieval-augmented generation with relevance grading
C.Fine-tuning the embedding model on domain data
D.Hybrid search combining vector similarity with BM25 keyword matching
AnswerB

The engineer is adding an LLM-based relevance grader between retrieval and generation. Each retrieved chunk is scored for whether it can actually answer the question, and only chunks that pass the grade are forwarded. This pattern improves answer precision by removing semantically similar but non-answering passages before the final prompt is assembled, which is exactly what the scenario describes.

Why this answer

The defining behavior in the scenario is a second LLM call that grades each retrieved chunk for answerability and drops the ones that fail. That is relevance grading layered on top of retrieval-augmented generation, which raises precision by keeping only context that can actually support the answer. The other techniques alter retrieval ranking, prompt reasoning, or embeddings but do not insert an evaluative filtering step between retrieval and generation.

Exam trap

The trap here is assuming that any technique which improves retrieval quality, such as hybrid search or embedding fine-tuning, also performs per-chunk relevance filtering with an LLM.

36
MCQhard

A GenAI application uses a Databricks Model Serving endpoint hosting a foundation model, and the team wants structured JSON output that conforms to a fixed schema for downstream parsing. Responses sometimes include prose or markdown fences that break the parser. Which design change is most likely to produce reliably parseable output?

A.Add a regex-based post-processor that strips markdown fences and extracts the first balanced JSON object from the raw response.
B.Lower the temperature to zero and add the phrase 'return only JSON' to the end of the user message.
C.Define the expected structure with a response format or structured-output parameter on the serving request so the endpoint constrains generation to the schema.
D.Increase max_tokens so the model has room to finish the JSON object without being truncated mid-structure.
AnswerC

Constrained decoding against a declared response schema forces the model's output to match the expected structure, eliminating prose and markdown fences at generation time rather than relying on post-processing. This is the most reliable design change because it enforces conformance at the source and keeps downstream parsing deterministic for the fixed schema.

Why this answer

Declaring a response format or structured-output schema makes the endpoint constrain decoding so the generated text matches the required structure, removing prose and markdown fences by construction. Lowering temperature, raising the token limit, or stripping fences afterward only reduce or mask the symptom without guaranteeing conformance.

Exam trap

The trap here is treating prompt wording such as 'return only JSON' as equivalent to schema enforcement, when only a declared response format constrains decoding.

37
MCQmedium

An engineer needs to ensure that only authorized users can access the RAG chatbot. Which Databricks security feature should be used to enforce this access control?

A.Notebook tags.
B.Cluster environment variables.
C.Unity Catalog access controls.
D.The public internet firewall.
AnswerC

Unity Catalog provides centralized access control that is essential for securing all data assets within Databricks. By defining clear permissions, engineers can ensure that only authenticated and authorized users have access to the data powering the chatbot, which is a fundamental requirement for any secure, enterprise-ready application.

Why this answer

Unity Catalog's access control lists (ACLs) provide a robust framework for managing permissions at various levels, including catalog, schema, table, and function. By integrating identity management with these controls, engineers can define granular access to the data and the endpoints behind the chatbot. This ensures that sensitive information is only accessible by verified users, maintaining compliance and enterprise-level data security throughout the system.

Exam trap

Candidates often choose workspace-level permissions or generic IAM roles. However, Unity Catalog provides the specific, granular object-level access controls required to secure data assets directly within the RAG pipeline architecture.

38
Multi-Selecthard

When designing an application that requires fine-tuning a small model (like Llama-3-8B) on Databricks, which THREE factors must be considered to ensure a successful training job?

Select 3 answers
A.Ensuring the training data is sufficiently representative of the target task.
B.Selecting a GPU-accelerated instance type sufficient for the model's memory.
C.Disabling checkpointing to save storage costs.
D.Configuring appropriate learning rates to prevent catastrophic forgetting.
E.Increasing the batch size to the maximum possible for all datasets.
AnswersA, B, D

Data quality and representativeness are the most critical factors for fine-tuning success. If the training data does not cover the breadth of scenarios the model will encounter, the fine-tuned model will fail to generalize. Ensuring the dataset is high-quality and well-curated is the first step toward effective model adaptation.

Why this answer

Training small models on Databricks requires careful consideration of compute, data quality, and hyperparameter tuning. By ensuring sufficient GPU memory, clean data, and properly tuned learning rates, engineers can successfully adapt the model to new domains. These factors are critical to avoid common pitfalls like catastrophic forgetting or overfitting, which can render the fine-tuned model less useful than its base version.

Exam trap

Candidates sometimes forget hyperparameter tuning specifics, ignoring learning rate configuration which directly causes catastrophic forgetting during fine-tuning.

39
MCQmedium

A GenAI engineer is building a Databricks RAG application that answers questions over a Delta table containing 40 million support tickets. Users report that simple keyword lookups return irrelevant results because the tickets use inconsistent terminology. The engineer needs semantic retrieval that stays synchronized as new tickets stream in every few minutes. Which Databricks component should be used to serve this retrieval layer?

A.A Feature Store table that materializes ticket embeddings refreshed by a nightly job
B.A Databricks Vector Search index with a Delta Sync index over the tickets table
C.A Delta Live Tables pipeline that computes embeddings and writes them to a Parquet directory
D.A Databricks SQL warehouse running a LIKE query over the ticket text column
AnswerB

A Delta Sync index continuously tracks the source Delta table and automatically updates embeddings as new tickets arrive, so semantic similarity search stays current without manual reindexing. This directly addresses the inconsistent terminology because retrieval is based on embedding proximity rather than literal keyword matching, and it scales to millions of rows on Databricks-managed infrastructure.

Why this answer

Semantic retrieval over frequently changing data requires both embedding-based similarity and automatic synchronization with the source table. A Databricks Vector Search Delta Sync index provides exactly this: it computes and maintains embeddings as the Delta table changes and exposes a low-latency endpoint for similarity queries, so terminology variations are handled by vector proximity rather than literal string matching.

Exam trap

The trap here is assuming that a scheduled embedding job plus a plain table is equivalent to a managed vector index, when only a Delta Sync index keeps retrieval current within minutes without manual reindexing.

40
MCQhard

A generative AI engineer is designing an agent on Databricks that uses a LangChain agent with tool-calling capabilities. The agent must call a Databricks SQL warehouse to run queries and a Vector Search index for retrieval. Which design consideration is most important for controlling agent behavior in production?

A.Define clear tool descriptions and limit the agent's available tools to only those required for the task.
B.Use the largest available foundation model for the agent to maximize reasoning quality.
C.Enable verbose logging of every intermediate step and store the logs in a Delta table for later review.
D.Increase the agent's maximum iterations to a very high number so it can always complete complex tasks.
AnswerA

The agent selects tools based on their descriptions, so precise descriptions and a minimal toolset reduce incorrect tool selection and unintended actions. Limiting tools also narrows the blast radius if the agent misbehaves. This is a core design principle for production agents that must be predictable and governable.

Why this answer

Agent behavior is driven by tool descriptions and the set of tools available. Clear descriptions help the model choose the right tool, and a minimal toolset reduces the chance of unintended calls. Iteration limits, logging, and model size affect performance or observability but do not directly control which tools the agent invokes.

Exam trap

The trap here is assuming that a more capable model or more iterations will make an agent reliable, when reliability comes from precise tool definitions and a constrained toolset.

41
MCQmedium

When evaluating the performance of a RAG application, which metric is most useful for measuring the quality of the retrieved context?

A.Total number of documents retrieved.
B.The total latency of the RAG pipeline.
C.Context Precision.
D.The number of parameters in the LLM.
AnswerC

Context Precision evaluates the ratio of relevant documents within the retrieved set. High precision ensures that the LLM is provided with high-quality, actionable context, which is the cornerstone of effective RAG. Monitoring this metric helps identify when the retrieval strategy needs adjustment to improve the application's overall accuracy.

Why this answer

Context Precision measures how much of the retrieved information is actually relevant to the user query. This is a critical metric for RAG systems because high retrieval precision directly reduces the noise in the LLM input, leading to more accurate and focused answers. Measuring this allows engineers to iteratively improve the retrieval pipeline, which is essential for maximizing the utility of the application.

Exam trap

Candidates frequently confuse evaluation metrics, mixing up context precision (what was retrieved vs. relevant) with generation metrics like answer correctness or faithfulness.

42
Multi-Selecthard

An engineer is designing a Databricks RAG application that must support multi-turn conversations where follow-up questions refer to earlier turns. They want the retrieval step to remain accurate as the conversation progresses. Which TWO design elements should they include? (Choose two.)

Select 2 answers
A.Maintain a conversation history buffer and condense prior turns into a standalone query before invoking Vector Search.
B.Increase the number of retrieved chunks for every turn to compensate for conversational ambiguity.
C.Include the previous assistant response as additional context in the prompt sent to the generation model.
D.Store the conversation history in the prompt and pass it unchanged to the retriever as the search query.
E.Disable metadata filtering during multi-turn conversations so more historical documents are eligible.
AnswersA, C

Follow-up questions often contain pronouns or omitted subjects that are ambiguous in isolation. Condensing the history into a self-contained query gives the retriever enough context to find relevant chunks, improving recall for multi-turn interactions. This is a standard pattern in conversational RAG design on Databricks.

Why this answer

Conversational RAG needs two coordinated elements: a rewritten, standalone query for retrieval, and conversation history for generation. Rewriting resolves references so Vector Search returns on-topic chunks, while including prior turns helps the model produce coherent follow-ups. Simply enlarging the retrieval window or removing filters does not address the ambiguity that multi-turn questions introduce.

Exam trap

The trap here is assuming that adding more conversation text to the retriever query improves multi-turn accuracy, when it actually dilutes the embedding and harms relevance.

43
MCQhard

A GenAI engineer is designing an agentic application on Databricks that uses a foundation model to decide which external tools to call. The team wants the agent to be able to invoke a Databricks SQL warehouse query and a Python function registered as a Unity Catalog function, and they need the model to select tools based on natural language requests. Which design element is required for the model to select and invoke these tools correctly?

A.Enable inference tables on the model serving endpoint to capture tool calls.
B.Store the tool list in a Unity Catalog volume and have the agent read it before each request.
C.Provide tool definitions with names, descriptions, and parameter schemas to the model in the request.
D.Fine-tune the foundation model on the team's historical tool-call logs.
AnswerC

Function-calling capable models select tools based on the tool definitions supplied in the request, which include the tool name, a natural language description, and a JSON schema for parameters. Without these definitions, the model has no knowledge of available tools or how to structure arguments. Supplying accurate definitions for the SQL warehouse query and the Unity Catalog function enables correct tool selection and argument generation.

Why this answer

Function-calling models rely on tool definitions passed in the request, including name, description, and parameter schema, to decide which tool to invoke and how to format arguments. Providing these definitions for the SQL warehouse query and Unity Catalog function gives the model the runtime contract it needs. Fine-tuning, volume storage, and inference tables do not supply that contract and therefore cannot enable correct tool selection.

Exam trap

The trap here is assuming that fine-tuning or logging replaces the runtime tool definitions, when the model needs explicit function schemas in every request to invoke tools correctly.

44
MCQmedium

Which approach is recommended for managing secrets, such as API keys for external LLM providers, in Databricks?

A.Hardcoding keys in a configuration file within a Repo.
B.Storing keys in Databricks Secrets.
C.Passing keys as plain text arguments in a job parameter.
D.Storing keys in a public S3 bucket with restricted access.
AnswerB

Databricks Secrets provides an encrypted, centralized way to manage sensitive keys. By using 'dbutils.secrets.get', you can inject these values into your code at runtime without them ever appearing in cleartext within your scripts, notebooks, or version control, ensuring a robust security posture for your production application.

Why this answer

Databricks Secrets provides a secure, encrypted storage mechanism for sensitive information, preventing credentials from being hardcoded in notebooks or scripts. This is a foundational security best practice. By using scope-based secret management, you ensure that only authorized users or service principals can access the credentials, keeping your RAG application secure and compliant with enterprise security standards.

Exam trap

Candidates often suggest environment variables or hardcoding secrets in notebooks, mistakenly believing these are acceptable for internal development, which violates basic Databricks security compliance standards.

45
MCQmedium

A GenAI team at a retail company has built a RAG chatbot on Databricks that answers customer questions from a product catalog stored in Delta Lake. The catalog is updated nightly, and the team wants the chatbot to reflect those updates without manual intervention. They are deciding how to keep the Vector Search index synchronized. Which approach should they use?

A.Use a Databricks SQL scheduled query to copy the catalog table into a Parquet file and point the retriever at that file each night.
B.Configure a Vector Search Delta Sync index on the catalog table so the index is refreshed automatically when the source table changes.
C.Create a standard Vector Search index and schedule a daily notebook that deletes the index and recreates it from the catalog table.
D.Enable Change Data Feed on the catalog table and rely on it to update the Vector Search index automatically.
AnswerB

A Delta Sync index in Databricks Vector Search tracks a Delta table as its source and can automatically sync new or changed rows, so nightly catalog updates flow into the index without a separate pipeline. This matches the requirement for unattended freshness and keeps retrieval aligned with the Delta Lake table.

Why this answer

A Delta Sync index is the Databricks Vector Search construct that keeps an index aligned with a Delta table, including incremental changes. Because the catalog is a Delta table that changes nightly, configuring Delta Sync lets the chatbot retrieve current product information without a custom rebuild job. The other choices either fail to update the index or introduce unnecessary full recomputation.

Exam trap

The trap here is assuming that enabling Change Data Feed on the source table is sufficient to refresh a Vector Search index, when the index itself must be configured for Delta Sync.

46
MCQeasy

What is the primary function of the 'Prompt' object in a Databricks RAG architecture?

A.To store the user's chat history for session persistence.
B.To format instructions and retrieved context for LLM consumption.
C.To index documents for faster similarity search.
D.To calculate the latency of the RAG pipeline.
AnswerB

The prompt is responsible for assembling retrieved documents and user queries into a format the model understands. This ensures that the context is properly delimited, the role of the model is defined, and the task is clearly stated, which is critical for generating high-quality, relevant, and accurate answers.

Why this answer

The prompt acts as the interface between the data retrieved from the vector index and the LLM's generative capability. It organizes the context, instructions, and user request into a coherent message structure. Proper prompt design is essential to focus the LLM's attention, reduce extraneous output, and ensure that the generative output remains grounded in the provided evidence.

Exam trap

Candidates think the prompt executes generation or runs vector searches, missing its actual role as a structured template organizing context and instructions for the LLM.

47
MCQmedium

An engineer is designing a RAG application that uses Databricks Vector Search to retrieve documents and a foundation model endpoint to generate answers. The team wants to log all user queries, retrieved documents, and generated responses for auditing and continuous improvement. They also need to monitor for drift in retrieval quality over time. Which Databricks capability should they integrate into the application design?

A.Delta Live Tables with expectations to validate the retrieved documents.
B.Databricks SQL dashboards querying the source Delta table for document updates.
C.Unity Catalog audit logs and lineage for the Vector Search index.
D.MLflow Tracing with inference tables on the model serving endpoint.
AnswerD

MLflow Tracing captures detailed spans for each step of a generative AI application, including retrieval and LLM calls, enabling end-to-end observability. Inference tables on the serving endpoint automatically log request and response payloads for auditing. Together they provide the query, retrieved documents, and generated response logging, plus monitoring for drift. This is the recommended Databricks approach for tracing and logging generative AI applications.

Why this answer

MLflow Tracing provides detailed spans across the RAG pipeline, capturing queries, retrieved documents, and LLM responses. Inference tables on the serving endpoint automatically log request and response payloads for auditing. This combination enables both observability and drift monitoring.

Other options focus on data pipelines, dashboards, or governance, which do not capture the runtime application behavior needed for continuous improvement.

Exam trap

The trap here is confusing governance auditing (Unity Catalog) or data pipeline quality (Delta Live Tables) with application-level tracing and payload logging, which are needed to monitor retrieval quality and improve the RAG application.

48
MCQmedium

You are building an application that uses Model Serving to host a fine-tuned LLM. Which configuration is required to optimize for high-concurrency request throughput?

A.Set the serving endpoint to use CPU-only instances for all LLM inference tasks.
B.Enable 'Provisioned Throughput' on the serving endpoint configuration.
C.Disable auto-scaling and fix the number of replicas to one.
D.Configure the endpoint to use an external API key for every request.
AnswerB

Provisioned Throughput provides dedicated capacity for foundation models, which is essential for high-concurrency scenarios. It ensures that the model has the necessary resources to handle concurrent requests without performance degradation, making it the correct choice for scaling LLM applications beyond simple development or low-traffic testing environments.

Why this answer

Enabling Provisioned Throughput for Model Serving allows for dedicated resources that can handle high-concurrency workloads efficiently. This is crucial for applications serving many users simultaneously. By optimizing resource allocation, you prevent bottlenecks during peak usage, ensuring consistent performance and minimizing request latency, which directly impacts the user experience and the scalability of the overall application architecture.

Exam trap

Candidates often select standard auto-scaling for high-concurrency needs, failing to realize that Provisioned Throughput is specifically required to guarantee performance for high-load production LLM workloads.

49
Multi-Selectmedium

A GenAI engineer is designing a RAG application on Databricks that must support multi-turn conversations where users refer to earlier messages, and the application must keep responses grounded in retrieved documents. Which TWO design elements are required to meet these requirements? (Choose two.)

Select 2 answers
A.Increase the LLM temperature so the model can explore multiple interpretations of ambiguous follow-up questions.
B.Store all conversation turns in a Delta table and query that table at inference time for each user message.
C.Disable retrieval for follow-up questions and rely on the model's parametric knowledge.
D.Rewrite the user query using conversation history before performing vector search retrieval.
E.Maintain conversation history and include relevant prior turns in the prompt sent to the model.
AnswersD, E

In multi-turn conversations, the latest user message often lacks the keywords needed for retrieval because it references prior context. Rewriting the query with conversation history produces a self-contained query that retrieves the correct documents. Without this step, vector search may return irrelevant chunks and the model's answer will not be grounded in the right evidence, breaking the grounding requirement.

Why this answer

Multi-turn conversation support requires passing relevant prior turns to the model so it can resolve references, and grounding requires that retrieval use a query rewritten with conversation context so the right documents are fetched. Together these elements ensure the model understands follow-up intent and answers from authoritative retrieved content. The other options either reduce determinism, add storage without context, or remove retrieval.

Exam trap

The trap here is treating conversation history storage as sufficient for multi-turn RAG, when the history must be used both in the prompt and in query rewriting before retrieval.

50
MCQmedium

An application requires low-latency retrieval of RAG metadata stored in Databricks. Which storage approach balances performance and cost while ensuring seamless integration with Unity Catalog?

A.Store embeddings in a flat Parquet file on DBFS accessed via direct Spark reads.
B.Utilize a third-party managed vector database outside of the Databricks environment.
C.Deploy a Databricks Vector Search index on a Delta table with automatic synchronization.
D.Implement a custom KNN algorithm using standard Python libraries on a single node cluster.
AnswerC

Databricks Vector Search provides a managed, scalable service specifically optimized for low-latency similarity search. It automatically handles the synchronization between your Delta table and the vector index, ensuring that as your source data updates, your search capability remains current without manual intervention or pipeline management.

Why this answer

Vector Search indexes are designed for low-latency similarity search, which is critical for RAG applications. By leveraging Databricks Vector Search, you avoid the overhead of custom search implementations. This approach ensures that embeddings are automatically synchronized with Delta tables, providing a scalable solution that integrates directly with Unity Catalog's security and governance framework, which is essential for maintaining consistent data access policies across enterprise environments.

Exam trap

Candidates mistakenly choose manual embedding generation jobs or standalone external databases, ignoring the native integration of Databricks Vector Search with Delta tables.

51
MCQmedium

An engineer is building a multi-turn GenAI assistant on Databricks. The assistant must answer follow-up questions that reference earlier turns, such as 'what about its warranty?', while keeping each request within the model's context limit. Which design should the engineer implement?

A.Send only the current user message to the model on every turn and rely on the model's built-in memory of prior requests.
B.Maintain the conversation in a Delta table, then on each turn build a prompt from a bounded recent window of turns plus a summary of older turns, along with retrieved context.
C.Persist the conversation in a Delta table and, on each turn, resend the entire raw history along with retrieved context, trimming nothing.
D.Encode the entire conversation history into a single embedding and prepend that vector to the prompt so the model can reconstruct prior turns.
AnswerB

Combining a recent-turn window with a rolling summary of older turns preserves the antecedents needed for follow-up questions while bounding prompt size. Retrieving relevant context per turn keeps answers grounded. This design directly meets both requirements: coherent handling of references like 'its warranty' and staying within the model's context limit as the conversation lengthens.

Why this answer

Bounding prompt size while preserving conversational antecedents is achieved by keeping a recent window of turns and summarizing older ones, with per-turn retrieval for grounding. Stateless endpoints, untrimmed history replay, and vector-encoded history all fail to give the model usable prior context within the context limit.

Exam trap

The trap here is assuming the serving endpoint retains state between requests, which leads to designs that either omit prior turns entirely or replay unbounded history.

52
MCQhard

Refer to the exhibit. An engineer is testing a model endpoint. The outputs are too brief and often stop mid-sentence. What is the most likely cause?

A.The temperature is too low for the model to generate full sentences.
B.The top_p parameter is too high, leading to early termination.
C.The max_new_tokens limit is too low for the expected response length.
D.The model version is incompatible with the specified parameters.
AnswerC

The 'max_new_tokens' parameter restricts the number of tokens the model is permitted to generate. If the model's intended response exceeds this count, it stops generation. This configuration is the direct cause of the truncation observed, and it must be increased to accommodate longer, more complete responses.

Why this answer

The 'max_new_tokens' parameter is set to 512, which imposes a hard limit on the length of the generated output. If the model needs more tokens to complete its thought, it will be abruptly cut off. Increasing this limit is the logical troubleshooting step when encountering truncation issues in generative model responses, ensuring the model has sufficient space to finalize its output properly.

Exam trap

Candidates often blame the model's training or temperature settings for truncated output, failing to recognize that max_new_tokens is a hard configuration limit that forces an early stop.

53
MCQmedium

Refer to the exhibit. What is the cause of this error when logging a RAG chain to MLflow?

A.The vector index is offline.
B.The model was logged without a signature defining input and output types.
C.The user has insufficient Unity Catalog permissions.
D.The LLM endpoint is overloaded.
AnswerB

MLflow Models require a signature to define the expected input and output structure. This signature is critical for serving, as it allows the platform to validate incoming requests. Without it, the model cannot be registered or deployed to a serving endpoint, resulting in the error shown during the logging process.

Why this answer

This error occurs because the model signature is missing or incorrectly defined, which is required for MLflow to perform input validation and inference. When serving a RAG chain, MLflow needs to know exactly what the input schema looks like so it can properly route requests and manage data types. Providing a valid signature ensures the model is deployable and functional.

Exam trap

Candidates often assume the error is related to model size or environment libraries. They overlook the critical requirement that MLflow must have a defined input/output schema for inference.

54
Multi-Selecthard

A GenAI engineer is designing a retrieval-augmented generation application whose source documents are long PDFs. Early testing shows that answers are vague because retrieved chunks contain several unrelated topics, and the language model frequently cites content that does not support its claims. The engineer wants to improve chunk quality before indexing. Which TWO changes should the engineer make to the ingestion pipeline? (Choose two.)

Select 2 answers
A.Split documents along semantic boundaries such as sections and headings instead of a fixed character count
B.Store the full PDF text in a single chunk per document
C.Increase the embedding dimension of the model used for the index
D.Lower the temperature setting on the language model
E.Add overlapping context between adjacent chunks
AnswersA, E

Semantic splitting keeps each chunk focused on a single topic, so retrieved context is coherent and the language model is less likely to blend unrelated statements. Fixed-size splitting is what produced multi-topic chunks in this scenario, and respecting document structure directly addresses that root cause.

Why this answer

Chunk quality is determined at ingestion: splitting on semantic boundaries keeps each chunk topically focused, and overlap preserves meaning across boundaries. Together they raise the precision of retrieved context so the model has grounded material to cite. Embedding dimension, temperature, and oversized single chunks do not correct incoherent chunking.

Exam trap

The trap here is treating vague answers as a model-tuning problem, when the evidence points to ingestion-stage chunking that mixes topics and drops boundary sentences.

55
MCQmedium

A generative AI team is building a customer-support assistant on Databricks. The assistant must answer questions using the company's private knowledge base, and the team wants to minimize latency while ensuring that the LLM only uses retrieved documents. They plan to use Databricks Vector Search with a Delta table as the source. Which design choice best balances low latency and grounded responses?

A.Use a keyword-based search over the Delta table and pass the top 10 matching rows to the LLM, bypassing Vector Search entirely.
B.Use a Vector Search index with a relevance score threshold and retrieve the top 5 chunks, then include those chunks in the prompt with instructions to answer only from the provided context.
C.Store the entire knowledge base in a single prompt and use a long-context LLM to answer questions without retrieval.
D.Configure a Vector Search index with a small embedding dimension and retrieve the top 20 chunks, then pass all chunks directly to the LLM without filtering.
AnswerB

This approach grounds the LLM by supplying only the most relevant retrieved chunks and explicitly instructing it to rely on that context. Limiting to top 5 with a relevance threshold reduces prompt size and latency while filtering out weak matches. Databricks Vector Search supports similarity search with scores, enabling this pattern, and the instruction reduces hallucination risk.

Why this answer

Grounding requires retrieving only relevant context and instructing the model to use it. A Vector Search index with a relevance threshold and a small top-k (e.g., 5) reduces prompt size and latency while filtering weak matches. Including those chunks with an explicit instruction to answer from context limits hallucination.

This combination balances performance and accuracy better than unfiltered or keyword-only retrieval.

Exam trap

The trap here is assuming that retrieving more chunks always improves grounding, when actually larger top-k without a relevance threshold increases latency and introduces noise that can cause the model to ignore the provided context.

56
MCQmedium

An engineer is designing a RAG application where the LLM must answer questions using only the retrieved context and must refuse to answer when the context is insufficient. Which prompt design approach best enforces this behavior?

A.Increase the number of retrieved chunks to ten so the model always has enough context to answer.
B.Use a system prompt that instructs the model to answer only from the provided context and to respond with a fixed refusal phrase when the context does not contain the answer.
C.Set the model's temperature to zero and rely on deterministic decoding to prevent unsupported answers.
D.Append the user's question twice in the prompt to reinforce the instruction to stay grounded.
AnswerB

A system prompt sets the model's operating constraints and is the most direct way to instruct it to rely solely on the retrieved context. Including an explicit refusal phrase gives the model a defined fallback, reducing hallucinations. This approach is deterministic in structure and works across providers that support system-level instructions.

Why this answer

The most reliable way to constrain a model to retrieved context is a system prompt that explicitly limits answers to that context and defines a refusal response when the information is missing. This gives the model a clear rule and a fallback, reducing hallucinations. Retrieval volume, temperature, and question repetition do not establish grounding boundaries.

Exam trap

The trap here is believing that deterministic decoding settings such as temperature zero eliminate hallucinations, when grounding must be enforced through prompt instructions.

57
MCQeasy

An engineer is designing a Databricks GenAI application that summarizes customer meeting notes. The notes contain personally identifiable information, and company policy requires that the summarization model never sends note text to an external provider. The team already has a fine-tuned open model registered in Unity Catalog. Which deployment choice satisfies the policy?

A.Serve the registered open model on a Databricks Model Serving endpoint and call it from the application
B.Use Databricks Foundation Model APIs with the pay-per-token option for a supported open model
C.Export the notes to a local workstation, run an open model there, and upload the summaries back
D.Call a third-party foundation model API directly from the notebook and pass the notes in the request body
AnswerA

Hosting the registered open model on a Model Serving endpoint keeps all inference inside the Databricks workspace, so note text never leaves the company's boundary. The endpoint exposes a standard API the application can call, and Unity Catalog governance continues to apply to the model artifact and its access controls.

Why this answer

The policy requires that note text never reach an external provider, so inference must run inside the company's Databricks environment. Serving the registered open model on a Databricks Model Serving endpoint keeps prompts and responses within the workspace while still offering a scalable API, and Unity Catalog continues to govern access to the model.

Exam trap

The trap here is equating a Databricks-hosted model API with in-workspace inference, when only serving the model inside the workspace guarantees the data boundary the policy demands.

Ready to test yourself?

Try a timed practice session using only Design Applications questions.