Courseiva

CCNA Application Development Questions

75 of 77 questions · Page 1/2 · Application Development · Answers revealed

1
MCQmedium

A data engineering team is using Databricks to prepare data for a RAG application. They want to ensure that document chunks are of consistent size and quality. Which tool should they use within the Databricks notebook environment to achieve this?

A.A standard Python for-loop on a single worker node.
B.Distributed Spark UDFs with LangChain text splitters.
C.Manual data extraction into a local CSV file.
D.Databricks SQL commands to manually truncate strings.
AnswerB

Using Spark UDFs allows developers to distribute text processing across the cluster, enabling high-throughput document chunking. Integrating LangChain within these UDFs provides standardized, industry-proven logic for splitting text, ensuring consistent chunk sizes that improve the quality and relevance of context retrieved during the RAG process for end users.

Why this answer

Databricks notebooks support a wide range of libraries, but for text processing, using Spark-native functions combined with libraries like LangChain allows for scalable and efficient chunking. By utilizing Spark's distributed processing capabilities, developers can process millions of documents in parallel. This is crucial for maintaining data quality in large-scale RAG systems, as it ensures that the context window of the LLM is not exceeded and that retrieval is focused and accurate.

Exam trap

Candidates tend to choose single-node Python libraries like standard LangChain text splitters without realizing that distributed Spark UDFs are required for scalable processing within notebooks.

2
MCQmedium

An AI engineer is developing a real-time customer service chatbot application using Databricks Model Serving and needs to securely store API keys and database credentials without hardcoding them into the application code or notebook. Which approach should the engineer use?

A.Store the credentials in a plain-text JSON file located in the workspace root directory and read it programmatically at runtime.
B.Hardcode the credentials directly into the scoring script deployed to the Databricks Model Serving endpoint.
C.Utilize Databricks Secrets to create a secret scope and retrieve values securely using dbutils.secrets.get during application execution.
D.Pass the credentials as plain-text arguments in the command-line interface when triggering the job cluster.
AnswerC

Databricks Secrets encrypt credentials at rest and in transit, integrating seamlessly with Unity Catalog and workspace access controls. Using dbutils.secrets ensures that sensitive tokens are masked in logs and only exposed dynamically to authorized execution contexts, safeguarding production environments effectively.

Why this answer

Databricks Secrets provide a secure mechanism for managing credentials and sensitive configuration parameters. By leveraging secret scopes and the dbutils.secrets utility within notebooks or environment variables mapped to serving endpoints, developers ensure that sensitive tokens never leak into version control systems, adhering to strict enterprise security and compliance standards for production application deployments.

Exam trap

Candidates often suggest using environment variables directly or manual file uploads, overlooking that dbutils.secrets is the standard, secure, and platform-native way to retrieve sensitive values within Databricks.

3
MCQhard

An engineer built an agent using Mosaic AI Agent Framework and wants the agent to call a Unity Catalog function that returns customer order history. The function must be invoked by the LLM at runtime without exposing raw SQL to the model. Which approach should the engineer use?

A.Expose the function through a SQL warehouse and have the agent send natural-language queries to the warehouse endpoint.
B.Embed the SQL body of the function in the system prompt so the LLM can generate the query when needed.
C.Register the Unity Catalog function as a tool in the agent and let the LLM select it by name and arguments.
D.Create a Databricks job that runs the function on a schedule and writes results to a Delta table the agent reads.
AnswerC

Mosaic AI Agent Framework supports Unity Catalog functions as tools that the LLM can invoke by name with structured arguments. The function's signature and docstring become the tool schema, so the model never sees or writes SQL directly. This is the supported pattern for giving agents governed access to governed data while keeping the interface declarative.

Why this answer

Registering the Unity Catalog function as a tool gives the agent a typed, governed interface the LLM can invoke by name with structured arguments. The function's metadata supplies the schema, so no SQL is exposed to or generated by the model. Prompt-embedded SQL, scheduled jobs, and text-to-SQL all fail the requirement of runtime, governed invocation without model-authored queries.

Exam trap

The trap here is believing the LLM must see or generate SQL to query Unity Catalog data, when the agent framework exposes functions as typed tools instead.

4
MCQhard

Refer to the exhibit. A developer is registering a model. Why is the model signature, as shown in the exhibit, considered a best practice for model registration?

A.It enables automatic data masking for sensitive strings.
B.It provides automated input validation to improve service reliability.
C.It automatically scales the number of serving nodes based on the prompt size.
D.It forces the model to use a specific version of Python for execution.
AnswerB

By defining the input schema, the model signature allows the serving infrastructure to validate every incoming request against the expected format. If the request is malformed, the system rejects it, preventing potential failures inside the model logic and ensuring that the serving endpoint maintains its stability under various load conditions.

Why this answer

Model signatures define the schema for inputs and outputs, allowing the serving layer to perform automated validation. When an incoming request doesn't match the signature, the serving infrastructure can reject it immediately, providing useful error messages. This prevents invalid data from reaching the model, reducing runtime errors and debugging time, and ensures that the application behaves predictably in production environments, which is critical for maintaining high service quality and reliability.

Exam trap

Candidates often assume model signatures are primarily used for logging metrics or tracking lineage, overlooking their critical runtime role in automated input validation.

5
MCQmedium

A team is deploying a fine-tuned LLM using Mosaic AI Model Serving. To reduce the cost of serving the model while maintaining acceptable performance, which strategy should be prioritized?

A.Use the largest available instance type to ensure zero request failures.
B.Implement model quantization to reduce memory footprint and hardware requirements.
C.Switch to a Multi-Model Serving endpoint regardless of throughput.
D.Disable all logging to save on storage costs.
AnswerB

Quantization reduces the precision of the model weights, which leads to smaller memory usage and faster inference times. This allows the model to run on smaller, cheaper instances without a significant drop in accuracy, providing a highly effective way to balance performance with operational costs in production environments.

Why this answer

Using smaller, optimized model architectures or quantization techniques significantly reduces the compute and memory footprint required for serving. By shrinking the model size, you can effectively use smaller instance types for serving, leading to direct cost savings. Furthermore, optimizing the underlying serving configuration by choosing the right hardware type ensures you are not paying for expensive GPU capacity when CPU-based serving would suffice for your throughput needs.

Exam trap

Candidates incorrectly suggest increasing GPU instance sizes to handle LLM costs, overlooking model optimization techniques that reduce hardware requirements altogether.

6
MCQmedium

A developer is deploying a custom LangChain agent application as a custom Python model to Databricks Model Serving. The application depends on a specific third-party library that is not included in the standard Databricks runtime environment. How should the developer ensure the dependency is installed when the model is loaded into the serving container?

A.Manually SSH into the serving container nodes after deployment and run pip install using terminal access.
B.Include the external library installation command inside a standard print statement within the model scoring function.
C.Specify the required packages in the extra_pip_requirements argument when logging the MLflow model artifact.
D.Rely on the serving endpoint to automatically discover and install any imported Python module dynamically from the public internet.
AnswerC

extra_pip_requirements records the third-party library in the MLflow model's dependency metadata, so Model Serving installs it into the container at load time. This satisfies the constraint that the package is absent from the standard Databricks runtime.

Why this answer

When logging custom PyFunc models to MLflow, developers must explicitly specify Python dependencies using the extra_pip_requirements parameter or a conda.yaml file. This ensures that the Databricks Model Serving environment automatically provisions and installs all necessary external libraries during container builds, preventing runtime import errors and ensuring reliable model execution.

Exam trap

Candidates often suggest manually installing packages on the cluster or using init scripts, failing to use the MLflow-native 'extra_pip_requirements' parameter that ensures environment reproducibility.

7
MCQhard

A GenAI engineer is developing an agent using Databricks Mosaic AI Agent Framework. The agent must call an external API to fetch real-time stock prices. The engineer wants to ensure the agent can be evaluated and deployed with minimal changes between development and production. Which approach should the engineer take to define the tool-calling logic?

A.Create a separate Databricks job that periodically fetches stock prices and writes them to a Delta table, then have the agent query that table.
B.Implement the API call as a Python function and register it as a tool in the agent's LangChain or similar framework, then log the agent with MLflow.
C.Use a Databricks SQL function to call the external API via a UDF, and invoke it from the agent.
D.Hardcode the API call directly in the agent's prompt as a description of the endpoint.
AnswerB

Defining tools as Python functions and registering them within the agent framework allows the logic to be logged as part of the MLflow model. This ensures the same code runs in evaluation and deployment, meeting the requirement of minimal changes between environments.

Why this answer

The Mosaic AI Agent Framework supports defining tools as Python functions that the agent can invoke. By registering the function as a tool and logging the agent with MLflow, the same code is used for evaluation and deployment, ensuring consistency. Other approaches either do not execute the call, introduce delay, or add architectural complexity that hinders portability.

Exam trap

The trap here is thinking that describing an API in the prompt enables the LLM to call it; in reality, tool execution requires code that the agent framework can invoke.

8
MCQmedium

Which approach is most effective for managing the dependencies of a custom ML model when deploying it to Mosaic AI Model Serving?

A.Install all dependencies manually on the serving cluster after deployment.
B.Include a 'requirements.txt' or 'conda.yaml' file in the model artifact.
C.Assume the serving environment has all standard ML libraries pre-installed.
D.Package the entire Python environment inside the model artifact as a zip.
AnswerB

Including dependency files in the model artifact ensures the model carries its environment definition with it. Mosaic AI Model Serving detects these files and builds the necessary environment for the model, ensuring that the production serving environment is identical to the one in which the model was validated.

Why this answer

Defining a custom environment using Conda or a 'requirements.txt' file ensures that the exact library versions required by the model are captured and recreated in the serving environment. This eliminates 'dependency hell' where a model works in development but fails in production due to library version mismatches. Ensuring consistent environments is a foundational step in robust ML engineering that guarantees reproducible performance in the serving infrastructure.

Exam trap

Candidates often assume that standard notebooks automatically bundle local Python environments into deployed endpoints, forgetting that Mosaic AI Model Serving requires explicit dependency files like requirements.txt or conda.yaml to recreate the correct libraries.

9
MCQmedium

An AI engineer is building a RAG application and wants to log the retrieval context and generated response for each request to MLflow for evaluation. They are using the `mlflow.langchain` flavor. Which method should they use to log the model so that MLflow automatically captures the necessary artifacts?

A.mlflow.sklearn.log_model
B.mlflow.tensorflow.log_model
C.mlflow.langchain.log_model
D.mlflow.pyfunc.log_model
AnswerC

`mlflow.langchain.log_model` is designed specifically for LangChain models. It automatically logs the chain's structure, prompts, and other artifacts, enabling MLflow to capture inputs and outputs for evaluation. This method simplifies logging and ensures compatibility with MLflow's evaluation tools, making it the correct choice.

Why this answer

For LangChain models, `mlflow.langchain.log_model` is the specialized logging function that captures the chain's configuration and enables MLflow to track inputs and outputs. This allows the engineer to later evaluate retrieval context and responses using MLflow's evaluation capabilities. Other logging methods do not provide this automatic integration.

Exam trap

The trap here is assuming that any logging function works, but only the flavor-specific logger for LangChain automatically captures the necessary artifacts for RAG evaluation.

10
MCQmedium

A developer is configuring a model serving endpoint as shown in the exhibit. They observe that the endpoint fails to respond quickly to the first request after a period of inactivity. What is the cause of this behavior?

A.The model version is incompatible with the CPU workload type.
B.The endpoint is entering a cold-start phase due to scale-to-zero.
C.The model name is not registered in the Unity Catalog.
D.The workload type 'CPU' is incorrectly specified as a string.
AnswerB

Setting 'scale_to_zero_enabled' to true triggers the termination of compute resources during idle periods to minimize costs. The subsequent latency is caused by the time required to provision new compute and load the model back into memory, which is the intended mechanism for this serverless configuration setting.

Why this answer

The 'scale_to_zero_enabled' flag is set to true, which instructs the Databricks infrastructure to shut down the compute resources entirely when no traffic is detected to save costs. When a new request arrives, the infrastructure must perform a 'cold start,' initializing the containers and loading the model into memory. This introduces latency, which is expected behavior when optimizing for infrastructure cost over instant availability in a serverless environment.

Exam trap

Candidates often mistake this for a network latency issue or an API rate limit. The 'scale-to-zero' configuration is a cost-saving feature that inherently introduces a cold-start delay.

11
Multi-Selectmedium

A team is developing a generative AI application that uses an LLM to answer questions based on internal documents. They want to ensure the application is robust and provides accurate responses. Which TWO practices should they implement to improve the reliability of the application? (Choose two.)

Select 2 answers
A.Implement retrieval-augmented generation (RAG) to ground the LLM's responses in the internal documents.
B.Fine-tune the LLM on the internal documents to embed knowledge directly into the model weights.
C.Implement guardrails to validate and filter the LLM's outputs before returning them to the user.
D.Use a larger LLM with more parameters to increase the likelihood of correct answers.
E.Set the LLM's temperature to a high value to encourage more creative and diverse responses.
AnswersA, C

RAG retrieves relevant documents and includes them in the prompt, so the LLM generates answers based on factual, up-to-date internal content. This reduces hallucinations and improves accuracy. It is a core practice for building reliable GenAI applications that need to answer domain-specific questions.

Why this answer

The two practices are RAG and guardrails. RAG grounds responses in internal documents, reducing hallucinations and providing accurate, up-to-date information. Guardrails validate outputs to filter out harmful or incorrect content.

Together, they significantly improve the reliability and safety of a GenAI application, especially for enterprise use cases.

Exam trap

The trap here is assuming that fine-tuning or using a larger model is the best way to improve reliability, but without retrieval and output validation, these approaches can still produce hallucinations and lack source attribution.

12
Multi-Selecthard

A company is scaling its RAG applications. Which THREE of the following are benefits of using Databricks Vector Search over managing a standalone vector database outside of the platform?

Select 3 answers
A.Automatic, continuous synchronization with Delta tables.
B.Ability to use SQL as the only way to perform embedding calculations.
C.Inheritance of Unity Catalog security and governance policies.
D.Reduction in operational overhead by using managed infrastructure.
E.Support for non-Databricks proprietary vector file formats.
AnswersA, C, D

Native integration with Delta tables allows the vector index to update automatically whenever the source table changes. This removes the need for manual ETL or complex cron jobs, drastically reducing the maintenance overhead and ensuring that the retrieved information is always fresh for the RAG application.

Why this answer

Databricks Vector Search offers benefits through tight integration, including automatic synchronization with Delta tables, which removes the need for custom ETL. Because it resides in the same security boundary, it inherits Unity Catalog's governance and access controls, ensuring data security. Finally, it uses the platform's managed infrastructure, eliminating the overhead of maintaining external database clusters, which allows developers to focus on application logic rather than infrastructure maintenance.

Exam trap

Test-takers sometimes pick features related to external model training or manual ETL pipelines, missing the native governance and synchronization benefits of Databricks Vector Search.

13
MCQmedium

Which of the following is a primary benefit of using Unity Catalog to manage models for a RAG application?

A.It automatically generates the optimal prompt for the model.
B.It provides centralized lineage and access control.
C.It eliminates the need for vector embeddings.
D.It forces all data to be stored in the cloud root.
AnswerB

Unity Catalog enables organizations to track data lineage from the source to the model, which is critical for compliance and debugging. Furthermore, it applies consistent access controls across workspaces, ensuring that sensitive data used in RAG applications is protected by the same security policies as the rest of the enterprise.

Why this answer

Unity Catalog provides a unified governance framework that manages permissions, lineage, and discovery across all data and AI assets. In a RAG context, this ensures that the entire pipeline—from the source Delta tables to the vector indexes and the final models—is traceable and governed. This consistency is essential for enterprise compliance and simplifies the operational management of the AI lifecycle by providing a centralized point for auditing and security policy application.

Exam trap

Candidates often choose 'Model Registry' or 'MLflow' without Unity Catalog. While those manage model artifacts, Unity Catalog specifically provides the enterprise-grade governance, lineage, and permissioning required for RAG pipelines.

14
Multi-Selecthard

A GenAI engineer is developing an agent on Databricks that must call a custom Python function to query an inventory database. The agent must decide when to invoke the function and must receive the result back into its reasoning loop. Which TWO actions are required to expose the function as a tool the agent can call? (Choose two.)

Select 2 answers
A.Deploy the function as a separate Mosaic AI Model Serving endpoint and hard-code its URL in the system prompt.
B.Add the function's source code to the retrieval index so the model can retrieve it as context at inference time.
C.Register the function with the agent using the framework's tool decorator or tool registration API so it appears in the tool list passed to the model.
D.Define the function with a clear docstring and type hints so the agent can infer its name, description, and parameters.
E.Convert the function's return value into an embedding and store it in Vector Search for similarity lookups.
AnswersC, D

Even a well-documented function is invisible to the agent unless it is registered. The framework's tool decorator or registration call adds the function to the set of tools advertised to the LLM, enabling the model to emit a tool call. This registration step is what connects the Python callable to the agent's reasoning loop.

Why this answer

Exposing a Python function as an agent tool requires two things: a well-formed schema the model can understand, derived from the function's name, docstring, and type hints, and explicit registration so the function is included in the tool list the agent passes to the LLM. Together these let the model decide when to call the tool and let the agent execute it and return results into the reasoning loop. The remaining options neither register the function nor execute it.

Exam trap

The trap here is thinking that documenting a function or making it retrievable is enough, when the agent also needs explicit tool registration to actually invoke it.

15
MCQhard

Refer to the exhibit. A developer wants to update this serving endpoint configuration to ensure it handles high-concurrency requests with consistent latency. Which change should be applied to the configuration?

A.Change 'workload_type' to 'GPU' and remove 'scale_to_zero_enabled'.
B.Set 'scale_to_zero_enabled' to 'false' and define 'min_provisioned_concurrency'.
C.Increase the 'model_version' to 'latest' to trigger automatic load balancing.
D.Add a 'timeout' parameter to the JSON configuration block.
AnswerB

Disabling scale-to-zero prevents the endpoint from shutting down, which eliminates cold-start latency. By defining 'min_provisioned_concurrency', the developer reserves a set number of replicas that are always running, ensuring that the system can handle concurrent requests immediately without waiting for infrastructure initialization, thus stabilizing latency under heavy load.

Why this answer

To handle high-concurrency with consistent latency, the configuration must move away from 'scale_to_zero_enabled: true' and utilize 'min_provisioned_concurrency'. Scaling to zero introduces cold-start latency that disrupts consistent response times required for high-concurrency production applications. By pinning minimum resources, the model remains active and ready to handle incoming traffic immediately, ensuring that performance remains predictable even during intermittent bursts of user activity within the production environment.

Exam trap

Candidates frequently select scale-to-zero options when high concurrency and consistent low latency are required, confusing cost savings with strict performance SLO requirements.

16
Multi-Selectmedium

Which THREE practices are recommended when using MLflow for managing LLM experiments in Databricks?

Select 3 answers
A.Log model parameters such as temperature, top_p, and system prompts.
B.Store all model weights directly in the MLflow run object.
C.Define model signatures for input and output data validation.
D.Always delete old MLflow experiments to save storage space.
E.Use MLflow tags to categorize experiments based on project or model version.
AnswersA, C, E

Logging hyperparameters and prompt configurations is essential for experiment reproducibility. It allows developers to compare how different settings impact model output, ensuring that the best configuration can be identified and replicated consistently across different environments, which is vital for maintaining model performance quality over time.

Why this answer

Effective experiment management requires traceability, reproducibility, and structured documentation. Logging parameters like temperature and system prompts ensures results can be replicated. Using signature definitions allows the system to validate inputs, reducing integration errors.

Finally, tagging models aids in lifecycle management, allowing teams to distinguish between prototypes and production-ready versions. These practices ensure that teams can maintain high standards of rigor, enabling scalable AI operations within the enterprise environment.

Exam trap

Candidates often focus only on logging metrics while forgetting to log generative hyper-parameters, system prompts, and model signatures which are essential for LLM reproducibility.

17
MCQmedium

Refer to the exhibit. A developer wants to enable monitoring for their deployed LLM endpoint. Given the current configuration, what must the developer change to ensure that request/response logs are captured for analysis?

A.Change the task to 'llm/v1/completions'.
B.Update 'auto_capture_request_payload' to true.
C.Add a new route with 0% traffic to the 'traffic_config'.
D.Increase the 'traffic_percentage' to 200%.
AnswerB

Setting 'auto_capture_request_payload' to true instructs the inference service to log incoming requests and outgoing responses to a managed Delta table. This is the direct configuration setting required to enable observability, ensuring that all data passed to the model is stored for later quality assessment and monitoring purposes.

Why this answer

The current configuration has 'auto_capture_request_payload' set to false, which prevents the logging of inference traffic. By updating this flag to true, the system will start capturing payloads into a Delta table. This is critical for monitoring model performance, drift, and quality.

Enabling this feature allows data teams to perform retrospective analysis, which is essential for iterating on model prompts and fine-tuning configurations based on real user interactions.

Exam trap

Candidates frequently attempt to create custom logging logic or external monitoring scripts, missing the built-in configuration flag 'auto_capture_request_payload' which is specifically designed for this purpose.

18
MCQhard

A developer is building an agentic workflow using Databricks and LangChain. The agent needs to decide whether to answer a user's query directly or call an external tool to retrieve additional information. The developer wants to ensure the agent's decisions are logged for debugging and auditing. Which approach should they take to achieve this?

A.Use MLflow's autologging for LangChain to automatically capture agent steps and tool calls.
B.Use Databricks Jobs to schedule the agent and rely on job run history for debugging.
C.Manually log each tool call using print statements to the driver logs.
D.Implement a custom logging function that writes agent decisions to a Delta table.
AnswerA

MLflow provides autologging for LangChain that automatically logs agent runs, including chain steps, tool invocations, and inputs/outputs. This creates a trace in MLflow, which can be viewed in the UI for debugging and auditing. It requires minimal code changes and is the recommended way to log agentic workflows on Databricks.

Why this answer

MLflow autologging for LangChain is the most efficient and integrated way to capture agent steps. It automatically logs tool calls, chain sequences, and inputs/outputs as MLflow traces, which can be inspected in the MLflow UI. This provides auditability and simplifies debugging without manual instrumentation.

Exam trap

The trap here is assuming that basic logging or job history is sufficient, but agentic workflows require detailed tracing of internal decisions, which MLflow autologging provides out of the box.

19
MCQmedium

A GenAI engineer is building a chatbot using Databricks Foundation Model APIs. They need the chatbot to maintain conversation context across multiple user turns and also allow the use of custom tools like a weather API. Which approach should they take?

A.Use the chat completion endpoint with a messages array that includes the full conversation history and define tools in the request.
B.Use the chat completion endpoint but only send the latest user message, relying on the model to remember previous turns.
C.Use the completions endpoint and concatenate all previous user inputs and model outputs into a single prompt string.
D.Use the embeddings endpoint to encode the conversation history and pass the embeddings to the model for context.
AnswerA

The chat completion endpoint accepts a messages array containing prior user and assistant messages, which provides conversational context. You can also specify tools (functions) in the request, and the model can return tool calls. This directly supports multi-turn context and custom tool integration, meeting both requirements.

Why this answer

The chat completion endpoint with a messages array is designed for multi-turn conversations and supports tool definitions. By including the full history, the model can reference prior context, and by defining tools, the model can request external API calls. This is the standard pattern for building conversational agents with Databricks Foundation Model APIs.

Exam trap

The trap here is assuming that the model retains memory between API calls, when in fact each request is stateless and must include the entire conversation history.

20
MCQmedium

A GenAI engineer is deploying a chain that calls an external LLM API. The chain must not block the serving thread while waiting on the remote API, and the endpoint must handle many concurrent requests. Which implementation approach should the engineer choose when logging the model?

A.Batch all incoming requests by writing them to a Delta table and have a separate job poll the table and call the external API.
B.Implement predict as an async function that awaits an async HTTP client, so the event loop can interleave other requests while waiting on the external API.
C.Use a synchronous HTTP client inside predict and increase the endpoint's scale-to-zero timeout so queued requests eventually complete.
D.Set the endpoint's concurrency to one and rely on the external API's own load balancing to absorb parallel requests.
AnswerB

Model Serving supports async predict, and an async HTTP client releases the event loop during network waits, allowing other requests to be handled concurrently on the same replica. This directly addresses the non-blocking requirement and improves throughput for I/O-bound chains. It is the recommended pattern for calling external LLM APIs from a served model.

Why this answer

Calling an external LLM API is I/O-bound, so the serving process should not block a thread while waiting. Writing predict as an async function with an async HTTP client lets the event loop serve other requests during those waits, improving concurrency on each replica. Synchronous clients, table-based queues, and single-concurrency settings all undermine throughput or break the synchronous API contract.

Exam trap

The trap here is treating a remote LLM call like local computation, where a synchronous client seems adequate, when in fact blocking I/O caps concurrency on the endpoint.

21
MCQmedium

A GenAI engineer is building a retrieval-augmented generation (RAG) application using Databricks Vector Search. During testing, they observe that for some queries, the retrieval step returns document chunks that are semantically similar but not actually relevant to the user's question, leading to poor answer quality. They want to improve retrieval precision without changing the embedding model. Which of the following approaches is most appropriate?

A.Reduce the number of retrieved chunks (top_k) to only the single highest-scoring chunk.
B.Switch the embedding model to a larger, more powerful model to improve semantic representation.
C.Enable hybrid search on the Vector Search index to combine keyword-based and vector similarity scores.
D.Increase the chunk size of the documents to include more context per retrieved chunk.
AnswerC

Hybrid search in Databricks Vector Search combines dense vector similarity with keyword-based (lexical) matching, which helps surface chunks that contain exact terms from the query. This improves precision when purely semantic matches drift to unrelated content. It does not require changing the embedding model and is a supported index configuration, making it the most direct and effective fix for the described issue.

Why this answer

Hybrid search enhances retrieval precision by blending vector similarity with keyword matching, ensuring that chunks containing exact query terms are ranked higher. This directly addresses the problem of semantically similar but irrelevant results without altering the embedding model. The other options either do not tackle the precision issue or violate the constraint of not changing the embedding model.

Exam trap

The trap here is assuming that any change to the retrieval pipeline, such as adjusting chunk size or top_k, will improve precision, when in fact only hybrid search directly combines lexical and semantic signals.

22
MCQhard

An engineer is troubleshooting a Vector Search index that fails to update. The index relies on a Delta table that is frequently updated. What is the most likely cause for the index failing to reflect new data?

A.The Delta table must be converted to a Parquet format.
B.The synchronization process was not triggered after the table update.
C.The embedding model version changed in the middle of a sync.
D.Unity Catalog permissions do not allow read access to the source table.
AnswerB

Vector Search indexes require an explicit trigger or a periodic sync configuration to ingest changes from the source Delta table. Without this action, the index does not automatically update, causing the RAG application to rely on outdated embeddings, which directly impacts the accuracy and relevance of the generative model's output.

Why this answer

Vector Search indexes are managed assets that require periodic synchronization with the source Delta table. If the synchronization process is not triggered automatically or via an API call, the index will remain stale. Understanding the lifecycle and sync mechanisms of Vector Search is essential for developers to ensure that RAG systems retrieve the most current information, as static indexes quickly lose relevance in dynamic data environments.

Exam trap

Test-takers often assume that updating the source Delta table automatically updates the Vector Search index in real time without needing a synchronization trigger.

23
MCQhard

A GenAI engineer is building a retrieval-augmented chatbot whose answers must cite the exact source document and page. The team wants the chatbot's responses to include structured citations that downstream UIs can render. Which Databricks feature should the engineer use to return these structured citations from the model endpoint?

A.Databricks Vector Search's built-in reranking, which automatically appends source metadata to the model's generated answer.
B.MLflow Tracing with span attributes that capture the retrieved document IDs and page numbers for each LLM call.
C.Mosaic AI Agent Framework's document citation support, where the retriever returns chunks with metadata and the agent formats them into a structured citations field in the response.
D.Mosaic AI Model Serving's payload logging, which writes each request and response to an inference table that the UI can read for citations.
AnswerC

The Mosaic AI Agent Framework supports returning retrieved documents alongside the model's answer, exposing a structured citations payload that includes source identifiers and metadata such as page numbers. This is precisely the mechanism for delivering renderable citations from a deployed agent endpoint. It pairs the retriever's chunk metadata with the agent's response schema, matching the scenario's requirement.

Why this answer

Structured citations must be part of the model endpoint's response contract, not a side channel. The Mosaic AI Agent Framework's citation support lets the retriever supply chunk metadata and the agent emit a citations field the UI can render directly. Observability traces, retrieval reranking, and payload logging serve different purposes and cannot substitute for returning structured citations to the caller at inference time.

Exam trap

The trap here is equating observability artifacts such as MLflow traces or inference tables with the response payload, when only the agent's response schema can deliver citations to the UI.

24
MCQmedium

A team needs to provide fine-grained access control to an LLM application that uses a vector index stored in a Delta table. Which Unity Catalog feature should they use to restrict access to specific rows based on user identity?

A.Column-level masking
B.Row-level security (RLS)
C.Workspace-level permissions
D.Access Control Lists (ACLs)
AnswerB

Row-level security policies filter rows at query time based on user attributes. When an application queries the vector index, the Databricks engine automatically applies the filter, ensuring the RAG model only has access to authorized context, which is the standard mechanism for implementing secure data-aware applications.

Why this answer

Row-level security (RLS) in Unity Catalog allows administrators to define policies that filter data returned from tables based on the user's identity. By applying these policies to the Delta table used as a knowledge base, the team ensures that the LLM only retrieves information the user is authorized to see. This is essential for compliance and security in multi-tenant or sensitive information application environments.

Exam trap

Candidates often confuse Row-Level Security with Column-Level Security or Attribute-Based Access Control. RLS is specifically designed to filter rows based on the current user's identity.

25
Multi-Selectmedium

A team is building a retrieval-augmented generation application on Databricks and wants to reduce hallucination by improving the quality of retrieved context before it reaches the LLM. Which TWO techniques should they apply? (Choose two.)

Select 2 answers
A.Increase the LLM temperature to encourage more diverse answers.
B.Apply a reranker model to the top-k retrieved chunks before passing them to the LLM.
C.Disable the Vector Search index and rely on keyword matching only.
D.Chunk documents with overlap and metadata so retrieval can filter by relevant attributes.
E.Store the entire document corpus in the prompt for every request.
AnswersB, D

A reranker scores retrieved candidates against the query with a cross-encoder, reordering them so the most relevant chunks appear first. This directly improves the precision of the context window and reduces the chance the LLM grounds its answer in loosely related passages. It is a standard post-retrieval step in Databricks RAG pipelines.

Why this answer

Reranking retrieved candidates with a cross-encoder and improving chunking with overlap and metadata both raise the relevance of the context supplied to the LLM, which directly reduces hallucination. Raising temperature, stuffing the full corpus, and dropping semantic search all degrade grounding rather than improve it.

Exam trap

The trap here is equating more context or more randomness with better answers, when grounding quality depends on retrieving fewer, more relevant chunks.

26
MCQeasy

A developer is using MLflow to track experiments for a generative AI application. They want to log a prompt template and its associated parameters so that they can reproduce the exact input to the model later. Which MLflow function should they use to log the prompt template as an artifact?

A.mlflow.log_artifact
B.mlflow.set_tag
C.mlflow.log_metric
D.mlflow.log_param
AnswerA

mlflow.log_artifact logs a file or directory as an artifact to the MLflow run. This is ideal for storing prompt templates as text files, allowing versioning and easy retrieval. It supports any file type and is the standard way to persist non-parameter assets like prompts, configuration files, or evaluation datasets.

Why this answer

Prompt templates are best stored as artifacts because they are files that may be reused, versioned, and referenced independently of the run. mlflow.log_artifact provides a robust way to persist such files, ensuring reproducibility and easy access later.

Exam trap

The trap here is confusing parameters with artifacts; parameters are for small configuration values, while artifacts are for files like prompt templates.

27
MCQmedium

A developer is building a RAG application. Which step is essential to prevent the model from hallucinating or providing outdated information?

A.Increase the temperature parameter of the model to 1.5.
B.Use a larger model to automatically detect and correct hallucinations.
C.Implement a real-time data ingestion pipeline for the vector database.
D.Force the model to always answer using a specific tone of voice.
AnswerC

Keeping the vector database up-to-date with the latest information is the most effective way to prevent hallucinations caused by outdated data. By ensuring that the RAG pipeline always retrieves the most current documents, the model is grounded in relevant, recent facts, which significantly improves the reliability and accuracy of the output.

Why this answer

Ensuring data freshness and providing high-quality, relevant context to the LLM during the retrieval step is critical. By implementing a robust data ingestion pipeline that updates vector indices in near real-time, the application ensures that the model is always informed by the most recent documents. This practice mitigates the risk of hallucinations by grounding the model's responses in verified, up-to-date business data, which is essential for building trustworthy GenAI solutions.

Exam trap

Candidates often choose manual static dataset uploads or model fine-tuning, overlooking that preventing hallucinations with fast-changing data specifically requires a real-time data ingestion pipeline into the vector database.

28
MCQeasy

When developing a Generative AI application in Databricks, which tool provides a collaborative environment for engineers to write code, visualize data, and document their experiments using Markdown?

A.Databricks File System (DBFS)
B.Databricks Notebooks
C.Unity Catalog Explorer
D.Databricks SQL Editor
AnswerB

Databricks Notebooks are the primary tool for collaborative data science and AI development. They support multi-language code execution, interactive visualizations, and Markdown documentation, providing the necessary features for engineers to build, document, and test GenAI models in a unified, shared environment that is accessible to the whole team.

Why this answer

Databricks Notebooks provide a versatile, collaborative environment designed for data science and AI development. They allow developers to combine code, SQL queries, visualizations, and rich text documentation (Markdown) in a single document. This makes them ideal for building GenAI applications, as engineers can document their model tuning processes, test retrieval logic, and share findings with stakeholders directly within the platform, facilitating a highly efficient and iterative development workflow.

Exam trap

Candidates often confuse Databricks Notebooks with external IDEs like VS Code or Databricks Repos, failing to recognize that Notebooks are the primary tool for integrated, collaborative experimentation with Markdown.

29
Multi-Selectmedium

A GenAI engineer is developing a conversational agent using Databricks. The agent must maintain context across multiple turns and retrieve relevant information from a knowledge base. They want to ensure the agent can handle follow-up questions that refer to previous exchanges. Which TWO techniques should they implement to manage conversation state and retrieval effectively? (Choose two.)

Select 2 answers
A.Store the conversation history in a Delta table and include it in the prompt for each turn.
B.Cache all responses from the LLM to avoid re-processing similar queries.
C.Use a separate Vector Search index for each conversation to retrieve relevant past messages.
D.Use a fixed window of the last three messages as context, ignoring older messages.
E.Rewrite follow-up questions into standalone queries using the conversation history before retrieving from the knowledge base.
AnswersA, E

Persisting conversation history in a Delta table allows the agent to retrieve past interactions and include them in the prompt, enabling context-aware responses. Delta tables provide ACID transactions and scalability, making them suitable for storing chat history. This approach ensures that follow-up questions can reference earlier turns, and the history can be trimmed or summarized to fit the context window.

Why this answer

Storing conversation history in a Delta table ensures persistence and allows inclusion of past turns in the prompt for context. Rewriting follow-up questions into standalone queries using that history improves retrieval accuracy by making the query self-contained. Together, these techniques enable the agent to handle multi-turn conversations effectively, where follow-up questions depend on earlier exchanges, and ensure the retrieval step has the necessary context.

Exam trap

The trap here is thinking that a separate Vector Search index per conversation or a fixed context window is sufficient, when scalable conversation state and query rewriting are the key techniques.

30
MCQmedium

An organization requires that all GenAI models deployed in Databricks be tracked and managed with a unified registry for compliance. Which feature should the developer use?

A.Use Git branches to manage different model versions.
B.Unity Catalog Model Registry
C.Store model artifacts as Pickle files in a shared workspace folder.
D.Create a custom Python class to track model metadata in a Delta table.
AnswerB

Unity Catalog Model Registry allows for central management, governance, and deployment of models. It enforces consistent access policies, tracks the lineage of model artifacts, and manages the lifecycle stages of models, ensuring that only approved models reach production while maintaining a clear, auditable trail for compliance requirements.

Why this answer

Unity Catalog Model Registry serves as the centralized hub for governing, versioning, and deploying machine learning models across an organization. By using the Unity Catalog, developers ensure that every model has a documented lineage, access control, and deployment status. This is critical for regulatory compliance and enterprise security, as it prevents unvetted models from being deployed and provides a transparent audit trail of every model version currently in use.

Exam trap

Candidates frequently confuse legacy MLflow Model Registry workspaces with the organization-wide compliance and governance features provided by the Unity Catalog Model Registry.

31
Multi-Selecthard

A team is designing an LLM application that requires strict data privacy. Which TWO approaches ensure that sensitive data is not leaked during the model inference process or training?

Select 2 answers
A.Redact PII from prompts using local libraries before sending data to the model.
B.Enable public access on all input data buckets for faster retrieval.
C.Restrict training data access using Unity Catalog fine-grained permissions.
D.Store all API keys as plain text in notebook variables.
E.Use the model provider's default logging for all user queries.
AnswersA, C

Sanitizing prompts locally ensures that sensitive information is stripped away before it enters the model inference pipeline. This prevents the model from processing or accidentally memorizing PII, providing a critical layer of security that mitigates the risks associated with data leakage in third-party or internal LLM deployments.

Why this answer

Data privacy in GenAI necessitates both input sanitization and secure environment configuration. By using PII redaction libraries before sending tokens to the model and enforcing Unity Catalog access controls on training datasets, developers create a robust defense-in-depth posture. These techniques protect against prompt injection and unauthorized data exposure, which are critical security considerations when deploying large language models within enterprise-grade environments where compliance and data sovereignty are top priorities.

Exam trap

Candidates assume that server-side data encryption alone is sufficient for privacy, missing that client-side PII redaction and granular access controls are necessary to prevent leaks during inference and training.

32
MCQmedium

When integrating an external LLM via a Databricks Model Serving endpoint, how should the API credentials be managed to ensure they are not exposed in the application code?

A.Store credentials as environment variables in the notebook directly.
B.Use Databricks Secrets to reference credentials at runtime.
C.Hardcode the credentials in a hidden Python module.
D.Encrypt credentials and store them in a JSON file within the repo.
AnswerB

Databricks Secrets are designed to securely store and manage sensitive credentials. Using the secret utility API, developers can inject keys into their code at runtime without the keys ever being written to the source code or persisted in plain text, maintaining high security standards for application integrations.

Why this answer

Databricks Secrets provide a secure way to reference sensitive information like API keys without hardcoding them in notebooks or source files. By using the 'dbutils.secrets.get' function, the application pulls the key at runtime from a secure vault. This is a best practice for developers to ensure security and prevent credentials from being accidentally committed to version control systems or visible to unauthorized users within the workspace.

Exam trap

Candidates often suggest environment variables or hardcoded config files. Neither is secure in a collaborative Databricks workspace; Secrets are the only approved way to manage sensitive credentials.

33
MCQmedium

A developer is creating a Databricks notebook to orchestrate a GenAI pipeline that includes data ingestion, vector index refresh, and model inference. They want to ensure that the pipeline can be easily tested and deployed across different environments. Which Databricks feature should they use to define the pipeline as code and manage deployments?

A.Databricks Jobs with notebook tasks and parameters.
B.Delta Live Tables with expectations.
C.MLflow Projects with a conda environment.
D.Databricks Asset Bundles (DABs).
AnswerD

Databricks Asset Bundles allow you to define jobs, notebooks, and other resources as code in YAML files. They support parameterization for different environments and can be deployed via CLI or CI/CD. This makes them ideal for managing GenAI pipelines across development, staging, and production, ensuring consistency and reproducibility.

Why this answer

Databricks Asset Bundles provide a declarative way to define and deploy Databricks resources, including jobs, notebooks, and configurations, as code. They support environment-specific parameters and CI/CD integration, making them the best choice for managing a GenAI pipeline across multiple environments.

Exam trap

The trap here is confusing orchestration tools like Jobs or Delta Live Tables with infrastructure-as-code solutions; Asset Bundles are specifically designed for packaging and deploying resources.

34
MCQmedium

A developer is using Databricks AI Functions to extract structured information from a large set of customer reviews stored in a Delta table. They want to apply a prompt to each review and store the results in a new column. Which function should they use?

A.ai_query()
B.ai_summarize()
C.ai_extract()
D.ai_classify()
AnswerA

ai_query() is a Databricks AI Function that allows you to run a prompt against a column of data in a SQL query. It sends the text to a specified model endpoint and returns the model's response, which can be stored in a new column. This is exactly suited for applying a prompt to each review in a Delta table.

Why this answer

Databricks AI Functions provide SQL-native access to generative AI models. ai_query() is the general-purpose function that accepts a prompt and a model endpoint, making it ideal for custom extraction tasks. It can be used in a SELECT statement to process each row and write the output to a new column. Other functions like ai_summarize and ai_classify are specialized and do not offer the same flexibility for structured extraction.

Exam trap

The trap here is assuming there is a dedicated ai_extract() function, when in fact ai_query() is the correct general-purpose function for custom prompts.

35
Multi-Selecthard

When designing an agentic workflow in Databricks, which TWO practices are essential to ensure the application remains observable and maintainable?

Select 2 answers
A.Enable MLflow tracing for all agent execution steps.
B.Avoid unit testing to speed up the development process.
C.Use automated evaluation suites to check for performance regressions.
D.Store all agent logic in a single, massive notebook file.
E.Disable logging to minimize storage consumption costs.
AnswersA, C

Tracing provides a detailed view of the agent's execution, including tool calls, reasoning steps, and model inputs/outputs. This visibility is essential for debugging complex agentic behaviors and ensures that developers can pinpoint where failures occur, ultimately making the agentic system easier to maintain and troubleshoot during production operations.

Why this answer

Observability and maintainability are critical for enterprise-grade AI agents. Using MLflow to log trace data provides visibility into the agent's decision-making process, which is necessary for debugging multi-step reasoning. Simultaneously, implementing robust evaluation suites ensures that agent updates do not introduce regressions in task performance.

These two practices form the foundation of a reliable AI operations strategy, enabling teams to confidently deploy and iterate on agentic workflows in production environments.

Exam trap

Test-takers often assume standard application logs are sufficient for complex multi-step agentic workflows, overlooking the necessity of deep execution tracing.

36
MCQhard

An AI engineer is deploying a RAG application using Databricks Model Serving. They need to ensure the endpoint can handle high traffic with low latency and automatically scale based on demand. Which configuration should they use?

A.Configure autoscaling with a minimum of 1 and maximum of 10 replicas, and set appropriate concurrency per replica.
B.Enable scale-to-zero and set a maximum concurrency of 1.
C.Disable autoscaling and manually provision 20 replicas to handle peak load.
D.Use a single large replica with maximum concurrency set to 100.
AnswerA

Autoscaling dynamically adjusts the number of replicas based on load, ensuring low latency during traffic spikes and cost savings during idle periods. Setting a minimum of 1 avoids cold starts, while a maximum of 10 caps resource usage. Concurrency per replica should be tuned to the model's throughput. This meets the requirements for high traffic and automatic scaling.

Why this answer

Autoscaling with a minimum and maximum replica count allows the endpoint to scale out during high traffic and scale in during low traffic, optimizing both latency and cost. Setting appropriate concurrency per replica ensures efficient utilization. This configuration provides the elasticity required for high traffic with low latency, while avoiding the pitfalls of fixed provisioning or overly restrictive concurrency limits.

Exam trap

The trap here is focusing only on scale-to-zero for cost savings without considering the need for sufficient concurrency and autoscaling to handle high traffic.

37
MCQeasy

Which tool in the Databricks ecosystem is best suited for developers to experiment with prompt engineering and tool-calling logic iteratively?

A.Databricks SQL Editor.
B.Databricks Notebooks.
C.Unity Catalog explorer.
D.Cluster Management UI.
AnswerB

Notebooks offer the most flexible and interactive environment for iterating on LLM prompts and testing code-based logic. The ability to run segments of code, visualize results, and maintain a history of experiments makes it the preferred tool for developers building and refining their generative AI application logic.

Why this answer

Databricks notebooks provide an interactive environment that is perfectly suited for iterative experimentation with prompts and chains. Developers can easily test different variations, observe the model's output in real-time, and integrate various tools, making it the ideal workbench for rapid prototyping in the early stages of the AI development lifecycle before committing to a final, production-ready implementation of their generative AI application.

Exam trap

Candidates often select heavy CI/CD deployment pipelines or MLflow registry tools for initial prompt experimentation, ignoring the need for rapid interactive testing.

38
MCQeasy

An AI engineer is using MLflow to track experiments for a generative AI application. They want to log parameters, metrics, and artifacts for each run, and later compare runs to select the best model. Which MLflow component should they use to organize runs into a named group for a specific project?

A.MLflow Model Registry
B.MLflow Run
C.MLflow Experiment
D.MLflow Tracking Server
AnswerC

An MLflow Experiment is a logical grouping of runs for a specific project or objective. It allows you to organize and compare runs, view metrics, and manage artifacts. By creating an experiment, you can log multiple runs and later query them to find the best model based on metrics. This is the core organizational unit in MLflow.

Why this answer

MLflow Experiments are designed to group runs for a specific project. They allow you to log and compare multiple runs, making it easy to track progress and select the best model. The experiment is the primary organizational unit in MLflow, and it is where runs are created and stored.

Exam trap

The trap here is confusing the organizational unit (Experiment) with the execution unit (Run) or the model management component (Model Registry), which serve different purposes.

39
MCQeasy

Which Databricks feature allows developers to track experiment parameters, model artifacts, and evaluation metrics in a structured way?

A.Unity Catalog
B.MLflow Tracking
C.Delta Lake
D.Databricks SQL
AnswerB

MLflow Tracking provides a robust API for logging experiments, including parameters, metrics, and models. This functionality allows developers to systematically track their progress, reproduce results, and compare different iterations, which is essential for managing the complexity of model development and ensuring that model quality is documented and verifiable.

Why this answer

MLflow Tracking is the industry-standard component integrated within Databricks for recording experiments. It allows developers to log parameters (e.g., learning rate), metrics (e.g., accuracy), and artifacts (e.g., model weights) during the development process. This structured approach is vital for reproducibility and comparing model versions, ensuring that the most effective model is identified before promotion to production, which is a key requirement for any serious ML engineering project.

Exam trap

Candidates often confuse MLflow with general project management tools. MLflow Tracking is the specific, structured component for recording parameters and metrics necessary for model reproducibility and comparison.

40
MCQeasy

Which Databricks asset is best suited for scheduling and orchestrating a multi-step GenAI pipeline that includes data ingestion, vector index updating, and model evaluation?

A.Databricks Feature Store
B.Databricks SQL Warehouse
C.Databricks Workflows
D.Unity Catalog
AnswerC

Databricks Workflows provides a robust orchestration engine to schedule and run multi-step pipelines. It supports various task types, including notebooks, JARs, and SQL queries, and allows for complex branching and dependencies, making it the ideal tool for orchestrating the end-to-end lifecycle of a GenAI data and model pipeline.

Why this answer

Databricks Workflows is the unified tool for orchestrating multi-step pipelines. It allows developers to define dependencies between tasks, manage retries, and monitor job execution. For GenAI, this is crucial because pipelines often involve sequential dependencies—such as ensuring vector database updates complete before a model is evaluated.

Leveraging Workflows ensures reliability and observability in automated production pipelines, which are essential for maintaining the integrity of data and models in an enterprise environment.

Exam trap

Candidates often select MLflow or Unity Catalog for pipeline scheduling, confusing artifact tracking and governance with the actual task orchestration capabilities provided by Databricks Workflows.

41
MCQmedium

A developer is creating a custom model serving endpoint that requires an external API call for data enrichment. What is the recommended way to handle sensitive API keys within the Databricks environment?

A.Store keys as environment variables in the notebook code.
B.Use the Databricks Secrets API to manage and retrieve keys.
C.Hardcode the keys directly into the model serving inference function.
D.Use the DBFS root directory to store key-value text files.
AnswerB

The Databricks Secrets API provides a secure, centralized location for managing sensitive information. It allows for role-based access control, ensuring that only authorized users or services can access the secrets, which protects the application from credential leakage and simplifies secret rotation and management across different deployment environments.

Why this answer

Secrets management is the cornerstone of secure application development in Databricks. By using the Databricks Secrets API, developers can decouple sensitive credentials from source code. This practice prevents the accidental disclosure of keys in version control systems and ensures that credentials are injected into the runtime environment securely, which is mandatory for maintaining a secure and audit-compliant AI development lifecycle in enterprise cloud environments.

Exam trap

Test-takers sometimes hardcode API credentials or configuration parameters directly inside the notebook or model artifact, compromising security and compliance standards.

42
MCQhard

An engineer is using Mosaic AI Agent Evaluation to score a conversational agent that calls tools. The agent sometimes answers correctly but with fabricated citations. Which evaluation approach best surfaces this specific failure mode?

A.Measure only token usage and latency to detect when the agent is hallucinating citations.
B.Run the agent without tools so it cannot produce citations at all.
C.Enable the groundedness judge and supply the retrieved context so the judge can verify claims against sources.
D.Rely solely on the correctness judge, since a correct answer cannot contain fabricated citations.
AnswerC

The groundedness judge checks whether statements in the response are supported by the provided context. Fabricated citations are exactly the failure mode it detects, because the cited source either does not exist in the context or does not contain the claim. Supplying the retrieved context is required for the judge to perform this verification accurately.

Why this answer

The groundedness judge compares response claims against the retrieved context, which is precisely how fabricated citations are detected. A correctness judge evaluates answers against references and misses unsupported citations, operational metrics say nothing about grounding, and removing tools avoids the scenario rather than evaluating it.

Exam trap

The trap here is conflating answer correctness with grounding, assuming that a factually right answer cannot also contain a fabricated citation.

43
MCQeasy

When building an application that retrieves context from Databricks Vector Search, what is the recommended data format for storing the document chunks?

A.JSON files in an S3 bucket
B.Delta tables
C.Parquet files on DBFS root
D.SQL Server database
AnswerB

Delta tables are the standard and recommended source for Vector Search indexes. They offer optimized performance, built-in change detection, and native compatibility with Databricks infrastructure. This ensures that the vector search index is always synchronized with the source data, which is crucial for building reliable and accurate AI-driven applications.

Why this answer

Delta tables are the native, optimized storage format for Databricks. When using Vector Search, the index is automatically maintained by Delta tables. This provides a performant and reliable way to sync data between the source Delta table and the vector index.

Storing data in Delta ensures that the vector search index can benefit from CDC (Change Data Capture) patterns, keeping the RAG knowledge base automatically updated and consistent.

Exam trap

Many candidates mistakenly select standalone vector databases or generic file formats like CSV/JSON instead of native Databricks storage formats that automatically sync indices.

44
MCQmedium

A developer needs to monitor the performance of an LLM application in production. They want to track the latency of their model endpoint. Where can they find this metric in the Databricks workspace?

A.In the Unity Catalog table lineage view.
B.In the 'Monitoring' tab of the Model Serving endpoint UI.
C.In the Databricks SQL Warehouse query history.
D.In the workspace audit logs.
AnswerB

The 'Monitoring' tab within the Model Serving endpoint interface is the dedicated location for viewing operational metrics. It displays auto-generated graphs for latency, request volume, and error rates, enabling developers to quickly assess the health and performance of their deployed models without needing additional configuration.

Why this answer

Databricks Model Serving provides built-in monitoring dashboards for every endpoint. These dashboards automatically track key performance indicators such as request latency, throughput, and error rates. Monitoring these metrics is vital for understanding the user experience and detecting potential performance bottlenecks.

By providing this information natively, Databricks helps developers maintain high service quality without requiring external observability tools, streamlining the monitoring and optimization process for deployed models.

Exam trap

Candidates frequently look for external APM tools or cluster logs, failing to recognize that Model Serving endpoints feature a dedicated built-in Monitoring tab.

45
MCQmedium

When building a RAG application, a developer wants to ensure that the retrieved context is strictly limited to documents the user has access to. Where should this security logic be enforced?

A.Inside the prompt engineering layer.
B.At the Vector Search retrieval layer.
C.Within the final response generation phase.
D.At the client-side browser application level.
AnswerB

Enforcing access control during the retrieval phase is the most effective approach. By filtering the documents returned from the vector index based on user identity or group membership, you ensure that the generative model only receives content that the specific user is authorized to view and process.

Why this answer

Access control must be enforced at the retrieval layer, ideally within the vector search engine or via a data-filtering mechanism that respects user identity. Relying on the model to enforce security is ineffective, as models cannot reliably manage authorization. Proper integration with Unity Catalog ensures that data retrieval is governed by the same permissions as the underlying tables, maintaining consistent security across the entire data-to-AI lifecycle.

Exam trap

Many candidates incorrectly assume that security logic should be handled by the LLM prompt or the application layer, ignoring that LLMs are prone to prompt injection and cannot reliably enforce data access.

46
Multi-Selectmedium

An AI engineer is designing a scalable customer support application on Databricks that integrates custom vector search indexes with a fine-tuned LLM. Which TWO architectural components are essential for enabling efficient similarity search and low-latency retrieval within the Databricks ecosystem? (Choose TWO)

Select 2 answers
A.Databricks Vector Search
B.An external third-party Hadoop cluster managed via SSH tunnels
C.Databricks Model Serving
D.A manual cron job running on a local developer laptop
E.A static CSV file stored on local driver disk storage
AnswersA, C

Databricks Vector Search is a serverless vector database natively integrated into the Databricks platform. It automatically synchronizes with Delta tables, manages embedding indexes, and provides high-performance similarity search APIs essential for retrieving contextually relevant documents in RAG applications.

Why this answer

Databricks Vector Search provides serverless vector database capabilities to store and query embeddings efficiently without managing external infrastructure. Combined with Databricks Model Serving for hosting the LLM, these native services ensure high throughput, enterprise-grade security, and low-latency inference required for modern retrieval-augmented generation applications built on the Databricks platform.

Exam trap

Candidates often select generic options like 'Delta Tables' or 'MLflow' without identifying the specific services (Vector Search and Model Serving) required for low-latency RAG performance.

47
MCQmedium

A team is deploying a GenAI application using Mosaic AI Model Serving. They want to ensure that the endpoint can handle sudden spikes in traffic without dropping requests. Which feature should they configure?

A.Request batching
B.Autoscaling
C.Model versioning
D.Provisioned throughput
AnswerB

Autoscaling dynamically adjusts the number of concurrent requests the endpoint can handle based on traffic. It scales resources up during spikes and down during lulls, ensuring requests are not dropped. This is the correct feature to handle sudden traffic increases while optimizing cost.

Why this answer

Autoscaling is designed to automatically adjust the serving endpoint's capacity based on incoming traffic. It ensures that the endpoint can handle sudden spikes by adding more resources and scales down when traffic decreases. This maintains availability and performance without manual intervention, making it the correct choice for handling variable loads.

Exam trap

The trap here is confusing provisioned throughput with autoscaling; provisioned throughput is fixed and does not automatically scale during unexpected spikes.

48
MCQeasy

A developer is packaging a GenAI chat application as an MLflow model that will be deployed to a Mosaic AI Model Serving endpoint. The application needs to load a retrieval index and a prompt template at startup so the first request is not slowed by initialization. Which MLflow logging pattern should the developer use?

A.Log the application with mlflow.pyfunc.log_model and perform the heavy initialization in the model's load_context method so resources are ready before predict is called.
B.Use mlflow.pyfunc.log_model with a signature only, and rely on the serving environment's default warm-up to populate the index and template automatically.
C.Log only the prompt template as an artifact and let the Model Serving endpoint fetch the index from Unity Catalog on each request.
D.Log the application with mlflow.pyfunc.log_model and load the index and template inside the predict function on every request.
AnswerA

MLflow's PyFunc flavor invokes load_context when the model is loaded into the serving process, before any predict call. Placing index and template initialization there ensures the endpoint pays the cost once per replica, and subsequent requests reuse the loaded objects. This is the documented pattern for stateful resources in MLflow models deployed to Mosaic AI Model Serving.

Why this answer

MLflow's PyFunc flavor calls load_context when the model is loaded, which happens before the endpoint begins serving requests. Initializing the retrieval index and prompt template there ensures they are ready once per replica and reused across calls. Loading inside predict or fetching resources per request would add latency, and a signature alone cannot prepare application state.

Exam trap

The trap here is assuming that serving platforms automatically warm up custom resources, when in fact the model author must initialize them in load_context.

49
Multi-Selectmedium

A team is evaluating their RAG application using Mosaic AI Model Evaluation. Which TWO metrics are most relevant for assessing the quality of the generated responses?

Select 2 answers
A.Faithfulness.
B.CPU utilization percentage.
C.Answer relevance.
D.Network latency in milliseconds.
E.Total storage cost of the index.
AnswersA, C

Faithfulness measures the degree to which the generated answer is derived exclusively from the retrieved context. This is the primary metric for detecting hallucinations, ensuring that the model does not invent information, which is a critical requirement for enterprise applications that prioritize truthfulness and reliability in their generative outputs.

Why this answer

Model evaluation requires quantitative metrics to measure both the accuracy of retrieved context and the quality of the final output. Faithfulness ensures the model sticks to the provided context, while answer relevance checks if the model actually addresses the user's intent. Using these metrics allows developers to iterate on their prompt engineering and retrieval strategy, ensuring the application remains accurate and useful for end-users while minimizing the risk of hallucinations.

Exam trap

Candidates often select traditional ML metrics like accuracy or F1-score, forgetting that LLM-based RAG evaluation requires specialized generative metrics like faithfulness and answer relevance.

50
MCQmedium

A GenAI engineer is building a retrieval-augmented generation application on Databricks. They want to store document embeddings and perform fast approximate nearest-neighbor search without managing a separate vector database. They have already created a source Delta table with columns: id (string), text (string), and embedding (array<float>). Which Databricks feature should they use to create a Vector Search index that automatically syncs with the Delta table?

A.Databricks Feature Store with a training set
B.MLflow Model Registry with a custom PyFunc model
C.Databricks Vector Search with a Delta Sync index
D.Delta Live Tables with a materialized view
AnswerC

Databricks Vector Search natively supports Delta Sync indexes that automatically keep embeddings in sync with a source Delta table. You can create an index using the embedding column and specify the source table; Databricks manages the underlying vector database and synchronization. This is the intended serverless vector search capability for RAG applications on Databricks.

Why this answer

Databricks Vector Search is the managed service designed for storing embeddings and performing fast similarity search. A Delta Sync index automatically syncs with a source Delta table, so when new documents or embeddings are added, the index updates without manual intervention. This eliminates the need to manage a separate vector database and integrates natively with Databricks workflows.

Exam trap

The trap here is assuming that any Databricks storage feature like Delta tables or Feature Store can serve as a vector index, but only Vector Search provides the required similarity search and automatic synchronization.

51
MCQhard

Refer to the exhibit. An AI engineer is configuring a Databricks Asset Bundle (DAB) to deploy a generative AI application. When executing 'databricks bundle deploy --target prod', which workspace host will the bundle resources be deployed to, and why?

A.The deployment will fail because the default target conflicts with the production target definition during workspace authentication initialization.
B.It will deploy to the development workspace because the dev target is explicitly marked with 'default: true'.
C.It will deploy to the production workspace at https://adb-987654321.azuredatabricks.net because the target flag explicitly selects the prod configuration block.
D.It will deploy to both workspaces simultaneously to maintain synchronization between development and production environments.
AnswerC

The '--target prod' CLI argument instructs the bundle deployment engine to evaluate the corresponding YAML section under targets. Consequently, the workspace host URL defined in the prod block is utilized for provisioning all associated jobs, pipelines, and endpoints.

Why this answer

Databricks Asset Bundles allow developers to define multiple deployment targets within the databricks.yml configuration file. When the user explicitly passes the --target prod flag to the CLI command, the bundle execution context switches from the default target to the specified target configuration, overriding default settings and directing deployment to the production workspace host URL defined under the prod block.

Exam trap

Candidates assume Databricks Asset Bundles always default to production configurations or local workspaces, ignoring the explicit role of the command-line target flag.

52
MCQeasy

A developer needs to store prompt templates, model parameters, and evaluation results for a GenAI application so that each iteration can be compared and reproduced later. Which Databricks capability should they use?

A.Delta Live Tables pipelines
B.MLflow Tracking with experiments and runs
C.Databricks SQL dashboards
D.Unity Catalog volumes
AnswerB

MLflow Tracking records parameters, metrics, artifacts, and tags per run within an experiment, which is exactly what is needed to compare prompt templates, model settings, and evaluation metrics across iterations. It is the standard mechanism in Databricks for reproducible GenAI development and integrates with Mosaic AI evaluation outputs.

Why this answer

MLflow Tracking provides experiments and runs that capture parameters, metrics, and artifacts, making it the right tool to version prompts, record model settings, and store evaluation results for comparison and reproduction. Data pipelines, volumes, and dashboards serve different purposes and do not offer run-based experiment tracking.

Exam trap

The trap here is assuming that any storage location for prompts, such as a Unity Catalog volume, also provides experiment tracking and run comparison.

53
MCQeasy

Which component in the Databricks GenAI stack is responsible for orchestrating the flow between data retrieval, prompt construction, and model invocation?

A.Unity Catalog.
B.MLflow Tracking.
C.Mosaic AI Agent Framework.
D.Databricks SQL.
AnswerC

The Mosaic AI Agent Framework provides the tools and abstractions needed to build, evaluate, and deploy agentic AI applications. It acts as the orchestration layer that connects data retrieval tools with model endpoints and manages the prompt engineering lifecycle, making it the correct choice for defining complex AI application logic.

Why this answer

Mosaic AI Agent Framework is designed to manage the complex orchestration required for RAG and agentic workflows. By providing a structured way to define tools, chains, and prompt strategies, it allows developers to build sophisticated applications that dynamically retrieve data and interact with LLMs. This orchestration layer is vital for building robust, maintainable AI applications where the logic needs to be clearly separated from the underlying model serving infrastructure.

Exam trap

Candidates incorrectly attribute prompt construction and retrieval orchestration to basic model serving or MLflow, ignoring the specialized role of the Mosaic AI Agent Framework.

54
MCQmedium

A developer is building a RAG application using Mosaic AI Model Serving. They need to ensure that the model endpoint logs inference requests and responses for audit purposes. Which configuration parameter should they enable?

A.Enable 'auto_capture_logs' in the endpoint environment variables.
B.Configure the 'inference_table_config' block in the endpoint request JSON.
C.Set the 'log_level' to 'DEBUG' in the Serving Endpoint settings.
D.Enable 'external_logging' in the Unity Catalog schema settings.
AnswerB

The 'inference_table_config' parameter is the specific configuration key required to enable automated request and response logging in Databricks Model Serving. This maps incoming requests directly to a specified Delta table, providing a structured, queryable record of all model interactions, which is essential for ongoing performance monitoring.

Why this answer

Enabling 'inference_table_config' in the model serving endpoint configuration is the standard Databricks approach for capturing telemetry data. This feature automatically writes request and response payloads to a Delta table, enabling compliance auditing, model monitoring, and drift detection. Understanding this integration is critical for production-grade AI deployments where transparency, debugging, and regulatory logging are mandatory requirements for enterprise-scale machine learning operations.

Exam trap

Candidates often assume logging is automatic or handled by the workspace. You must explicitly configure the 'inference_table_config' to capture request/response data into a Delta table for auditing.

55
MCQmedium

A developer needs to deploy a custom Python model that requires non-standard library dependencies. Which MLflow feature should the developer use to specify these environment requirements during model logging?

A.Model Schema definition
B.MLflow environment requirements (pip_requirements)
C.Global workspace library settings
D.Custom model signatures
AnswerB

By explicitly providing a list of required libraries via the `pip_requirements` argument during the logging process, you guarantee that the Model Serving environment will install them before loading the model. This is the industry-standard way to ensure that complex, custom model code runs successfully in production.

Why this answer

When logging a custom model with MLflow, the developer should use the `pip_requirements` or `conda_env` parameter in the `mlflow.pyfunc.log_model` function. This ensures that the environment is correctly packaged and reproducible. During deployment, the Databricks Model Serving service reads these requirements to recreate the identical software environment, preventing runtime errors caused by missing dependencies when the model is loaded and served in the production environment.

Exam trap

Candidates often assume the environment is automatically captured from the local machine. They fail to explicitly define 'pip_requirements' or 'conda_env', causing deployment failures when the server lacks local dependencies.

56
MCQmedium

A GenAI engineer has a RAG application whose retrieval step uses Databricks Vector Search. Users report that answers are sometimes irrelevant because the retriever pulls chunks from documents the user is not authorized to see. The engineer must enforce per-user document ACLs at query time without re-indexing the corpus. Which approach should the engineer take?

A.Post-filter the retrieved chunks in the application layer by querying Unity Catalog for the user's group memberships and dropping chunks that fail the check.
B.Create a separate Vector Search index per user and have the application route each request to the user-specific index.
C.Add a metadata column that stores the document's ACL principals, then pass a filter string with the caller's identity to the Vector Search query.
D.Enable Delta Sharing on the source Delta table and rely on the workspace's existing table ACLs to filter Vector Search results.
AnswerC

Databricks Vector Search supports metadata columns on the index and accepts a filters argument at query time. Storing allowed principals (users or groups) in a metadata column and passing the caller's identity as a filter ensures only authorized chunks are returned, with no re-indexing. This is the documented pattern for per-user authorization in RAG retrieval on Databricks.

Why this answer

Vector Search indexes can carry metadata columns that are queryable through the filters parameter, letting the application pass the caller's identity so only permitted chunks are returned. This satisfies per-user ACL enforcement at query time without rebuilding the index when permissions change, and it avoids leaking unauthorized content into the application tier. The other approaches either do not affect index results or are operationally impractical at scale.

Exam trap

The trap here is assuming that Unity Catalog or Delta Sharing ACLs automatically propagate to a Vector Search index's query results, when in fact the index must be filtered explicitly via metadata.

57
MCQmedium

Refer to the exhibit. A developer is deploying a model using the provided JSON configuration. What is the primary benefit of setting 'scale_to_zero_enabled' to true in this production RAG application?

A.It increases the throughput of the model by optimizing memory allocation during inference.
B.It ensures that the model always stays warm to minimize latency for every user request.
C.It reduces operational costs by shutting down compute resources when the endpoint is not serving traffic.
D.It enables high availability by automatically replicating the model across multiple availability zones.
AnswerC

The primary purpose of enabling scale-to-zero is cost reduction. In environments where request patterns are sporadic or non-continuous, this feature ensures that the company is only billed for the compute resources actually consumed during active inference periods, rather than paying for constant uptime when no requests are present.

Why this answer

Setting 'scale_to_zero_enabled' to true allows the Databricks Model Serving endpoint to automatically shut down compute resources when no requests are being processed. This is highly beneficial for cost optimization, as it eliminates idle runtime costs during periods of inactivity. When a new request arrives, the service automatically initializes the model, ensuring cost efficiency without requiring manual intervention to scale the underlying infrastructure up or down.

Exam trap

Candidates often confuse 'scale_to_zero_enabled' with model accuracy improvements or latency reduction, forgetting that its primary purpose is exclusively cost optimization by eliminating idle compute charges during inactive periods.

58
MCQmedium

A GenAI engineer is building an agent with LangChain on Databricks. The agent must call a Unity Catalog function `catalog.schema.get_weather` to fetch current weather. They want the LLM to decide when to invoke this function. Which LangChain component should they use to expose the Unity Catalog function to the LLM?

A.Databricks Vector Search retriever
B.Databricks Unity Catalog function as a LangChain tool
C.MLflow pyfunc model wrapper
D.Databricks SQL Connector
AnswerB

LangChain provides a `UCFunctionToolkit` that wraps Unity Catalog functions as LangChain tools. This allows the LLM to see the function's metadata and invoke it when needed. The toolkit handles authentication and execution via the Databricks SDK, making it the correct choice for integrating Unity Catalog functions into an agent.

Why this answer

To let an LLM decide when to call a Unity Catalog function, the function must be presented as a LangChain tool. The `UCFunctionToolkit` from Databricks integrates with LangChain, automatically generating tool definitions from Unity Catalog functions. This enables the agent to invoke the function dynamically based on the conversation, which is essential for building responsive GenAI applications.

Exam trap

The trap here is confusing data retrieval components like Vector Search with tool-calling mechanisms, overlooking that Unity Catalog functions require a specific toolkit to be exposed as tools.

59
MCQhard

A GenAI engineer is building a multi-step agent that uses Databricks Foundation Model APIs. The agent must decide when to call a weather tool and when to answer directly. The engineer wants to ensure the agent's decision-making is reliable and that failures in tool calls are handled gracefully. Which design approach should the engineer use?

A.Chain multiple LLM calls where the first call always invokes the weather tool and the second call formats the answer.
B.Implement a ReAct-style loop where the LLM outputs a thought, action, and action input, and the agent executes the tool and feeds back the observation.
C.Fine-tune the LLM on examples of weather queries and answers, then deploy it without tools.
D.Use a single prompt that instructs the LLM to output a JSON with either an answer or a tool call, and parse it once.
AnswerB

A ReAct-style loop structures the agent's reasoning and tool use, allowing it to decide when to call the weather tool and incorporate the result. It also provides a clear place to handle tool errors by catching exceptions and feeding error messages back to the LLM for recovery.

Why this answer

A ReAct-style loop enables the agent to reason about whether a tool is needed, execute it, and observe the result before deciding the next step. This iterative process supports graceful error handling because exceptions can be captured and returned as observations, allowing the LLM to adjust. Single-pass or fixed-chain approaches lack this adaptability and feedback.

Exam trap

The trap here is assuming that a single LLM call can both decide to use a tool and produce a final answer incorporating the tool's output, which is not possible without a loop.

60
MCQmedium

A GenAI engineer has registered a RAG chain in Unity Catalog as a model and now needs to deploy it for real-time inference with per-request token usage and latency captured automatically. Which Databricks capability should they enable on the serving endpoint?

A.Attach an MLflow experiment to the endpoint and rely on the run history for telemetry.
B.Enable Unity Catalog lineage on the registered model version.
C.Configure the endpoint with autoscaling enabled to log per-request metrics.
D.Enable inference tables on the Mosaic AI Model Serving endpoint.
AnswerD

Inference tables on Mosaic AI Model Serving automatically capture the request payload, response, and metadata for each served request, enabling monitoring of token usage and latency without custom logging code. For a RAG chain registered in Unity Catalog, this gives immediate observability into production traffic and is the supported mechanism for capturing real-time inference telemetry.

Why this answer

Inference tables are the native Mosaic AI Model Serving feature that persists request and response payloads plus metadata to a Delta table, giving automatic capture of token usage and latency for real-time endpoints. Autoscaling, MLflow experiments, and Unity Catalog lineage address capacity, development tracking, and governance respectively, but none record per-request serving telemetry.

Exam trap

The trap here is assuming that any monitoring-adjacent feature, such as autoscaling or lineage, will automatically capture request-level telemetry for a served model.

61
Multi-Selecthard

Which TWO of the following are mandatory requirements for developing an AI application using the Databricks Mosaic AI Model Serving environment?

Select 2 answers
A.The model must be stored in a legacy Databricks Workspace folder.
B.The model must be registered as a model version within a Unity Catalog schema.
C.The model must be served using a shared-access mode interactive cluster.
D.The serving endpoint must be configured with a defined compute resource.
E.The application code must perform manual model sharding across nodes.
AnswersB, D

Unity Catalog acts as the central repository for model artifacts and versions in Databricks. Registering the model here provides the necessary metadata, lineage, and access control required by the serving infrastructure to deploy the model securely and ensure it remains reachable by authorized internal or external applications.

Why this answer

Mosaic AI Model Serving requires specific configurations to ensure secure and performant access. Firstly, the model must be registered in Unity Catalog to maintain governance and lineage. Secondly, the serving endpoint requires defined compute resources, typically GPU-accelerated for LLMs, to handle inference requests.

These requirements are essential for productionizing models, as they ensure that models are discoverable, governed, and have the necessary hardware to meet low-latency performance targets in real-time scenarios.

Exam trap

Candidates often focus on model training parameters or API authentication methods, missing the foundational requirements of Unity Catalog registration and defined compute resources for serving.

62
MCQeasy

A developer is creating a Databricks notebook to prototype a GenAI application. They need to install the `databricks-langchain` library to use LangChain integrations with Databricks. Which command should they use in the notebook?

A.%conda install databricks-langchain
B.dbutils.library.installPyPI("databricks-langchain")
C.!pip install databricks-langchain
D.%pip install databricks-langchain
AnswerD

The `%pip install` magic command is the standard way to install Python packages in Databricks notebooks. It ensures the package is installed in the notebook's environment and available for import. This is the correct command to install the `databricks-langchain` library, which provides LangChain integrations for Databricks.

Why this answer

In Databricks notebooks, the `%pip install` magic command is the recommended way to install Python libraries. It ensures the package is installed in the notebook's isolated environment and is immediately available for import. Using `%pip` with `databricks-langchain` correctly sets up the LangChain integrations needed for the GenAI application.

Exam trap

The trap here is using outdated installation methods like `dbutils.library.installPyPI` or shell commands, which are not recommended in current Databricks runtimes.

63
Multi-Selectmedium

When designing a production-ready Databricks notebook for model inference, which TWO practices improve maintainability and performance?

Select 2 answers
A.Embed all logic including data preprocessing in a single, large notebook.
B.Use Databricks Widgets for parameterizing inputs like model paths.
C.Hardcode all file paths and configuration settings for consistency.
D.Refactor reusable code into libraries imported in the notebook.
E.Always run inference on the driver node to avoid network latency.
AnswersB, D

Widgets allow developers to pass parameters to notebooks at runtime, which is essential for making notebooks reusable. By parameterizing critical values like model paths, input table names, or configuration flags, developers can trigger the same notebook in different workflows without modifying the code, significantly increasing flexibility and maintainability in production.

Why this answer

Modular code design allows for easier testing and unit-level debugging, which is crucial for complex inference logic. Using parameterization enables the notebook to be reused across different environments and models without changing the core code. These practices reduce technical debt and simplify the lifecycle of ML models, making the transition from development to production much smoother and ensuring consistent performance across different stages of the CI/CD pipeline.

Exam trap

Test-takers mistakenly believe that hardcoding file paths or keeping all logic inside a single monolithic notebook improves execution speed, ignoring best practices for code maintainability.

64
MCQeasy

A developer is using MLflow to track experiments for a RAG application. They want to log the retrieval step's parameters, such as the number of documents retrieved (k) and the embedding model used. Which MLflow API should they use?

A.mlflow.log_artifact()
B.mlflow.set_tag()
C.mlflow.log_param()
D.mlflow.log_metric()
AnswerC

mlflow.log_param() is used to log a single parameter (key-value pair) for a run. Parameters are typically configuration settings like k or model name. This is the correct API for logging retrieval parameters such as the number of documents and the embedding model, as they are scalar values that define the run's configuration.

Why this answer

MLflow parameters are meant for logging configuration settings that are constant for a run. The number of documents retrieved and the embedding model are such settings. mlflow.log_param() records them as key-value pairs, enabling easy comparison across runs in the MLflow UI. Metrics, artifacts, and tags serve different purposes and are not suitable for this use case.

Exam trap

The trap here is confusing parameters with metrics; parameters are configuration inputs, while metrics are output measurements that can vary during training.

65
MCQhard

A developer deploys a new model version as shown in the exhibit. What is the purpose of this configuration?

A.To increase compute capacity for the primary model.
B.To conduct A/B testing or a canary rollout.
C.To load balance between two different cloud regions.
D.To force all traffic to model-v1 when model-v2 errors.
AnswerB

Configuring traffic percentages allows for controlled, incremental rollouts, which are essential for A/B testing or canary deployments. This allows developers to validate the new model's performance on live traffic while minimizing the impact if the new model version exhibits unexpected behavior or suboptimal accuracy metrics.

Why this answer

This configuration implements a canary deployment strategy by splitting traffic between two model versions. By routing 90% of requests to the stable version and 10% to the new version, the team can monitor performance and accuracy in a real-world scenario with minimal risk. This is a standard practice for safely rolling out model improvements, allowing teams to collect data on the new version's behavior before a full-scale transition.

Exam trap

Candidates often misinterpret traffic splitting as a load balancing or scaling technique, rather than identifying its primary purpose: testing new models against live traffic to minimize risk.

66
MCQmedium

When developing a feature engineering pipeline using Feature Store, which practice ensures maximum code reusability across training and inference?

A.Calculate features in the application code and pass them as raw inputs.
B.Hardcode feature calculation logic inside the training notebook.
C.Use the Feature Store client to log feature definitions and retrieve them.
D.Export features to a static CSV file after every training run.
AnswerC

Utilizing the Feature Store API allows developers to define features once and publish them to a Feature Table. This centralized repository acts as a single source of truth, allowing both training pipelines and online serving endpoints to fetch consistent, pre-computed feature values using the same lookup key.

Why this answer

Defining features as code using the Databricks Feature Store ensures that the same logic is applied during both batch training and real-time inference. By encapsulating feature calculations in Feature Tables, developers avoid the 'training-serving skew' where features are calculated differently in production. This practice is critical in GenAI development to ensure model performance consistency and reduce technical debt caused by disjointed preprocessing pipelines across different stages of the ML lifecycle.

Exam trap

Candidates often suggest using local Python functions or manual SQL scripts. These approaches do not track lineage or ensure consistency, leading to 'training-serving skew' in production.

67
Multi-Selectmedium

Which TWO actions are necessary to ensure that a model serving endpoint in Databricks remains available and performant during peak traffic hours?

Select 2 answers
A.Disable the serving endpoint's logging to save on compute cycles.
B.Enable auto-scaling for the serving endpoint.
C.Configure the endpoint to run on a single, fixed-size node to minimize latency.
D.Select appropriate GPU-accelerated compute types for the workload.
E.Manually restart the endpoint every hour to clear the cache.
AnswersB, D

Auto-scaling allows the serving endpoint to dynamically adjust its instance count based on current request volume. This ensures that the system handles spikes in traffic effectively without manual intervention, maintaining consistent performance and avoiding service degradation when user demand increases suddenly during peak hours or application usage cycles.

Why this answer

Maintaining performance requires both vertical and horizontal scaling strategies. First, configuring auto-scaling on the endpoint allows the infrastructure to adjust to fluctuating demand automatically. Second, ensuring the endpoint uses appropriate instance types—such as GPU-accelerated instances for LLMs—ensures the compute capacity is sufficient for inference tasks.

These two actions are foundational for achieving high availability and low latency, preventing request timeouts and ensuring a smooth experience during heavy usage periods.

Exam trap

Students often select only one correct action or mistakenly choose manual scaling scripts, forgetting that both auto-scaling and GPU acceleration are required for high performance.

68
MCQeasy

Which Databricks feature is specifically designed to facilitate the rapid development and deployment of LLM applications by providing a managed environment for hosting and testing prompts?

A.Databricks SQL Warehouse
B.Databricks Mosaic AI Playground
C.Unity Catalog Volumes
D.Delta Live Tables
AnswerB

Mosaic AI Playground is the dedicated UI for interacting with models, testing prompts, and comparing responses in real-time. It is the primary tool for early-stage development, allowing developers to see how different model configurations and prompts affect output quality without writing code, effectively accelerating the initial application development cycle.

Why this answer

Databricks Mosaic AI Playground provides a low-code interface for developers to experiment with different LLMs and system prompts. It allows for quick iteration and testing of model behavior before moving into full production deployment. This environment is crucial because it bridges the gap between experimentation and application development, ensuring that developers can validate prompt engineering strategies within the same governed environment where their data and production models reside.

Exam trap

Test-takers frequently guess generic cloud tools or external third-party playgrounds instead of the native Databricks-specific environment explicitly designed for rapid prompt engineering and LLM testing.

69
MCQmedium

A developer is building a RAG application using Mosaic AI Model Serving. They need to ensure that the embedding model endpoint is strictly accessed only by specific service principals within the workspace. Which feature should the developer configure to enforce this security requirement?

A.Network Security Groups
B.Unity Catalog External Locations
C.Model Serving Permissions
D.Cluster Access Control Lists
AnswerC

Model Serving Permissions allow administrators to define precise access control lists for each inference endpoint. By assigning 'Can query' privileges only to the required service principals, the developer effectively secures the endpoint, ensuring that unauthorized entities cannot perform inference operations against the deployed embedding model in the production environment.

Why this answer

To restrict access to Mosaic AI Model Serving endpoints, developers should use Model Serving Permissions. By navigating to the Permissions tab of the specific endpoint, they can grant 'Can query' access exclusively to authorized service principals. This ensures that only authenticated and authorized applications can retrieve embeddings, protecting the model from unauthorized inference calls while maintaining a secure development lifecycle in Databricks.

Exam trap

Candidates often confuse workspace-level access controls or IAM roles with endpoint-specific permissions. They mistakenly select broad settings like 'Workspace Admin' or 'Can View' instead of the specific 'Can query' permission required for inference.

70
MCQmedium

A developer is building a retrieval-augmented generation (RAG) application on Databricks. They need to ensure that embeddings are updated automatically when the underlying Delta table changes. Which approach is the most efficient and scalable?

A.Write a manual PySpark job that runs every hour to scan the entire Delta table and recompute all embeddings.
B.Use a Databricks Job with a notebook task triggered by a cron schedule to process changes using a watermark.
C.Implement a Delta Live Tables pipeline using streaming tables and a Python UDF to generate embeddings on arrival.
D.Deploy a Unity Catalog volume to store the embeddings and trigger an external API call from an event-driven function.
AnswerC

Delta Live Tables streaming tables automatically handle incremental data ingestion and state management. By applying a transformation function during the stream, the pipeline computes embeddings only for new or updated records, significantly reducing compute overhead and ensuring the vector index is always current without manual maintenance or scheduling logic.

Why this answer

Delta Live Tables (DLT) with streaming tables allows for continuous data processing and incremental updates. By integrating embedding generation directly into the pipeline, the developer ensures that the vector database stays in sync with the source data without manual intervention or complex scheduling. This architecture minimizes latency and improves data consistency, which is a foundational requirement for production-grade RAG applications within the Databricks ecosystem.

Exam trap

Candidates often suggest manual batch jobs or triggers, which are less efficient and harder to scale than Delta Live Tables' native streaming capabilities for continuous data processing.

71
MCQmedium

A developer is building a RAG application and notices that the retrieval step often returns irrelevant context. Which step in the pipeline should be improved to address this?

A.Increase the number of API calls to the LLM.
B.Refine the chunking strategy and embedding model quality.
C.Enable auto-scaling on the serving endpoint.
D.Switch to a larger, more expensive LLM.
AnswerB

The quality of retrieval is heavily dependent on how the source data is chunked and represented as vectors. Refining these parameters—such as using smaller, more meaningful segments or a higher-quality embedding model—is the standard approach to ensuring that the search engine finds the most relevant information for any given query.

Why this answer

Improving retrieval accuracy often involves enhancing the quality of the embeddings or the text chunking strategy. By adjusting how data is partitioned into chunks (e.g., overlapping chunks, semantic chunking) or refining the embedding model, the vector search index can better capture the semantic meaning of the documents. This is a common and critical improvement path in RAG development, as the quality of the retrieval is fundamentally limited by the index's representational capability.

Exam trap

Candidates often jump to increasing the number of retrieved chunks or changing the LLM model. These address symptoms rather than the root cause of poor semantic retrieval quality.

72
MCQmedium

A developer is configuring a RAG application and needs to ensure that the LLM response is based on specific, trusted document snippets. Which technique, when implemented correctly, helps mitigate hallucination by grounding the response in provided context?

A.Increasing the model's temperature parameter
B.Retrieval Augmented Generation (RAG)
C.Fine-tuning the base model on all internal documents
D.Using a larger foundation model without context injection
AnswerB

RAG grounds the model's output by providing relevant, factual information from trusted sources within the prompt. This context-based approach limits the model's tendency to hallucinate by forcing it to answer based on the provided document snippets, which are retrieved via similarity search before the generation step occurs.

Why this answer

Retrieval Augmented Generation (RAG) is the primary technique for grounding LLM responses. By retrieving relevant, trusted documents from a vector store based on a user's query and injecting those documents into the LLM's prompt, the developer forces the model to synthesize an answer based on specific retrieved context rather than relying solely on its internal training data. This significantly reduces hallucinations and increases the accuracy and relevance of the generated responses.

Exam trap

Candidates often confuse RAG with Fine-tuning or Prompt Engineering. They assume the model's internal weights are being updated, when RAG is strictly about providing external context at inference time.

73
MCQhard

Refer to the exhibit. A developer encounters this error when trying to call a Model Serving endpoint from a job. Which action should the developer take to resolve this authorization failure?

A.Increase the 'workload_size' in the model configuration.
B.Grant 'Can query' permissions to the service principal in the endpoint settings.
C.Update the workspace token expiration policy.
D.Restart the serving endpoint to refresh security policies.
AnswerB

This is the correct approach to fix an authorization error. By updating the access control list (ACL) for the specific Model Serving endpoint to include the service principal with 'Can query' rights, the security policy is satisfied, allowing the job to successfully execute inference calls against the LLM.

Why this answer

The error indicates a lack of 'Can query' permission for the service principal or user attempting to access the endpoint. To resolve this, the developer must update the permissions for the specific endpoint within the Databricks workspace. By granting the required 'Can query' privilege to the service principal running the job, the application will regain authorization to perform inference against the model, satisfying the security policy requirements.

Exam trap

Test-takers often look for cluster-level spark configurations or IAM role modifications when encountering endpoint authorization errors, overlooking workspace access control lists.

74
MCQmedium

A GenAI engineer has built a retrieval-augmented generation (RAG) application using Databricks Vector Search and a Databricks-hosted LLM served via Mosaic AI Model Serving. Users report that responses are sometimes irrelevant or cite incorrect document passages. The engineer wants to systematically improve answer quality by identifying which retrieved chunks are actually being used by the LLM. Which approach should the engineer take to capture the relationship between retrieved context and the generated response for later evaluation?

A.Configure the Model Serving endpoint to log all requests and responses to a Delta table for offline analysis.
B.Increase the Vector Search index's embedding dimension to improve semantic matching.
C.Add a reranker model after Vector Search to reorder retrieved chunks by relevance score.
D.Enable MLflow Tracing on the RAG chain to log each retrieval and generation step with inputs and outputs.
AnswerD

MLflow Tracing captures spans for retrieval and LLM calls, including retrieved chunks and the final response, enabling correlation analysis. This directly addresses the need to see which context influenced the answer and supports systematic evaluation and debugging of the RAG pipeline.

Why this answer

MLflow Tracing instruments each stage of a RAG pipeline, recording retrieved chunks and the LLM's output as linked spans. This gives the engineer the data needed to see which context was passed and potentially used, enabling targeted improvements and evaluation. Other options either change retrieval behavior or log only endpoint-level data, lacking the granular linkage required.

Exam trap

The trap here is assuming that endpoint request logging alone provides enough detail to know which retrieved chunks influenced the answer, when it only captures the final prompt and response.

75
MCQhard

A team is deploying a LLM-based application using Databricks Model Serving. They want to implement robust observability and monitoring for their endpoint. Which TWO features should they utilize to track performance and quality metrics? (Select TWO)

A.Model Inference Tables
B.Model Serving Monitoring Tab
C.Unity Catalog Lineage
D.Databricks Delta Sharing
E.Workspace-level Audit Logs
AnswerA, B

Inference tables provide a structured way to capture all requests and responses sent to a serving endpoint. This data is written directly into Unity Catalog, allowing developers to perform SQL analysis on input prompts and model outputs to evaluate quality, detect drift, and perform offline debugging of model performance.

Why this answer

To effectively monitor a LLM application, teams must integrate both system-level performance metrics and application-level quality telemetry. Inference tables automatically log input and output data for analysis, while the native 'Monitoring' tab in the Model Serving UI provides latency and throughput metrics. Combined, these features provide a comprehensive view of how the model is performing, identifying bottlenecks in latency or degradation in response quality over time.

Exam trap

Candidates tend to confuse generic cluster monitoring tools with LLM-specific telemetry features, missing that Model Inference Tables and the Serving Monitoring Tab are specifically built for tracking endpoint inputs, outputs, and quality metrics.

Page 1 of 2 · 77 questions totalNext →

Ready to test yourself?

Try a timed practice session using only Application Development questions.