Databricks · Free Practice Questions · Last reviewed May 2026
36real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
A Data Engineer needs to ensure that PII data in a Delta table is masked before serving it to non-privileged users. Which Databricks feature provides the most efficient, centralized control for this requirement?
Apply Row-Level Security filters using Delta table constraints.
Create materialized views with hard-coded redacted values.
Utilize Unity Catalog column-level masking functions.
Unity Catalog enables defining dynamic masks on columns using SQL functions. This centralizes security policy enforcement, applying masking rules consistently across all users and compute clusters. It is the most robust method for PII protection because it prevents unauthorized visibility without altering the underlying raw data storage files.
Implement Spark UDFs to perform in-memory data masking.
You need to ingest data from an external JSON source into a Delta table. The source schema is inconsistent. Which strategy is most effective for preparing this data?
Use inferSchema = true in the read configuration.
Define a rigid schema at the ingestion point.
Implement a bronze-silver medallion pattern.
The bronze-silver pattern enables ingestion of raw, unstructured data into a staging area, followed by rigorous cleaning and validation in a downstream silver table. This approach prevents pipeline failures during ingestion and ensures that cleaning logic is decoupled from the data acquisition layer, allowing for better maintainability.
Convert the JSON to CSV before loading into Delta.
Which TWO of the following are benefits of using Delta Lake for data preparation over standard Parquet files on cloud storage?
ACID compliance for reliable write operations.
ACID compliance ensures that concurrent reads and writes are managed correctly, preventing data corruption during simultaneous operations. In standard Parquet, a failing write operation could leave orphan files or partial data, whereas Delta Lake ensures an 'all-or-nothing' consistency model that is critical for production pipelines.
Automatic conversion of JSON to XML format.
Time travel capability for auditing and debugging.
Time travel allows users to query older versions of a table using snapshot IDs or timestamps. This is invaluable for debugging data preparation issues, recovering from accidental overwrites, or maintaining audit trails, which is not natively possible with standard Parquet files on object storage.
Native support for cross-cloud streaming ingestion.
Automatic hardware upgrades for compute clusters.
Which Databricks feature is best suited for maintaining the lineage of data used during the preparation of training sets for Generative AI?
Delta Lake Time Travel
Unity Catalog
Unity Catalog captures fine-grained lineage information as data moves from raw ingestion to final training sets. It offers a unified view of dependencies, which is essential for auditing the training process. This visibility is crucial for ensuring that training data meets safety and compliance standards in enterprise environments.
Cluster Policies
MLflow Experiments
When preparing a dataset for fine-tuning an LLM, you need to ensure the data is representative of the target domain. What is the most effective approach to detect and mitigate sampling bias in your training set using Databricks?
Increase the batch size during the fine-tuning process.
Perform statistical profiling on feature distributions and re-sample if necessary.
Statistical profiling reveals imbalances, allowing engineers to apply techniques like oversampling or undersampling to correct them. This quantitative approach is objective and standard for ensuring high-quality machine learning training sets. It prevents the model from favoring specific patterns simply because they were more frequent in the dataset.
Change the model architecture to a larger parameter count.
Use a random seed to shuffle the data before splitting.
Which of the following describes the purpose of using a 'Feature Store' when preparing data for Generative AI applications?
To store raw text files for long-term archival.
To provide a unified interface for consistent feature computation during training and inference.
The primary goal of a feature store is to eliminate inconsistencies in feature logic. By centralizing this, organizations ensure that the features the model was trained on are computed in exactly the same way when deployed. This is foundational for building stable, repeatable machine learning pipelines in Databricks.
To serve as a high-performance database for LLM model weights.
To replace the need for data cleaning and preprocessing entirely.
Want more Data Preparation practice?
Practice this domainA developer is building a RAG application using Mosaic AI Model Serving. They need to ensure that the model endpoint logs inference requests and responses for audit purposes. Which configuration parameter should they enable?
Enable 'auto_capture_logs' in the endpoint environment variables.
Configure the 'inference_table_config' block in the endpoint request JSON.
The 'inference_table_config' parameter is the specific configuration key required to enable automated request and response logging in Databricks Model Serving. This maps incoming requests directly to a specified Delta table, providing a structured, queryable record of all model interactions, which is essential for ongoing performance monitoring.
Set the 'log_level' to 'DEBUG' in the Serving Endpoint settings.
Enable 'external_logging' in the Unity Catalog schema settings.
Refer to the exhibit. A developer wants to update this serving endpoint configuration to ensure it handles high-concurrency requests with consistent latency. Which change should be applied to the configuration?
Change 'workload_type' to 'GPU' and remove 'scale_to_zero_enabled'.
Set 'scale_to_zero_enabled' to 'false' and define 'min_provisioned_concurrency'.
Disabling scale-to-zero prevents the endpoint from shutting down, which eliminates cold-start latency. By defining 'min_provisioned_concurrency', the developer reserves a set number of replicas that are always running, ensuring that the system can handle concurrent requests immediately without waiting for infrastructure initialization, thus stabilizing latency under heavy load.
Increase the 'model_version' to 'latest' to trigger automatic load balancing.
Add a 'timeout' parameter to the JSON configuration block.
When developing a feature engineering pipeline using Feature Store, which practice ensures maximum code reusability across training and inference?
Calculate features in the application code and pass them as raw inputs.
Hardcode feature calculation logic inside the training notebook.
Use the Feature Store client to log feature definitions and retrieve them.
Utilizing the Feature Store API allows developers to define features once and publish them to a Feature Table. This centralized repository acts as a single source of truth, allowing both training pipelines and online serving endpoints to fetch consistent, pre-computed feature values using the same lookup key.
Export features to a static CSV file after every training run.
Which Databricks asset is best suited for scheduling and orchestrating a multi-step GenAI pipeline that includes data ingestion, vector index updating, and model evaluation?
Databricks Feature Store
Databricks SQL Warehouse
Databricks Workflows
Databricks Workflows provides a robust orchestration engine to schedule and run multi-step pipelines. It supports various task types, including notebooks, JARs, and SQL queries, and allows for complex branching and dependencies, making it the ideal tool for orchestrating the end-to-end lifecycle of a GenAI data and model pipeline.
Unity Catalog
An organization requires that all GenAI models deployed in Databricks be tracked and managed with a unified registry for compliance. Which feature should the developer use?
Use Git branches to manage different model versions.
Unity Catalog Model Registry
Unity Catalog Model Registry allows for central management, governance, and deployment of models. It enforces consistent access policies, tracks the lineage of model artifacts, and manages the lifecycle stages of models, ensuring that only approved models reach production while maintaining a clear, auditable trail for compliance requirements.
Store model artifacts as Pickle files in a shared workspace folder.
Create a custom Python class to track model metadata in a Delta table.
When integrating an external LLM via a Databricks Model Serving endpoint, how should the API credentials be managed to ensure they are not exposed in the application code?
Store credentials as environment variables in the notebook directly.
Use Databricks Secrets to reference credentials at runtime.
Databricks Secrets are designed to securely store and manage sensitive credentials. Using the secret utility API, developers can inject keys into their code at runtime without the keys ever being written to the source code or persisted in plain text, maintaining high security standards for application integrations.
Hardcode the credentials in a hidden Python module.
Encrypt credentials and store them in a JSON file within the repo.
Want more Application Development practice?
Practice this domainAn organization is deploying a GenAI application using Databricks Model Serving. Which TWO steps are required to ensure the deployment environment handles model governance and observability effectively?
Enable Unity Catalog for the registered model and its versions.
Unity Catalog acts as the central governance layer for all data and AI assets. Enabling it for registered models provides unified access control, lineage tracking, and auditability, ensuring that only authorized users can deploy or modify models, which is a foundational requirement for robust enterprise AI security and governance.
Use MLflow to enable inference logging for model endpoints.
Enabling inference logging via MLflow captures request and response data, which is vital for monitoring model performance and identifying potential hallucinations. This data helps engineers analyze drift, debug production failures, and optimize the generative AI model, providing the necessary visibility into how the model behaves when deployed in production.
Hardcode API credentials within the model inference script.
Disable automatic scaling to maintain consistent latency.
Manually deploy the model via the standard cluster user interface.
Which Databricks feature is primary for managing the lifecycle, versioning, and deployment readiness of custom Generative AI models?
Unity Catalog Volumes
MLflow Model Registry
The MLflow Model Registry provides a centralized store for managing the full lifecycle of a model. It allows teams to version models, transition them between lifecycle stages like 'Staging' and 'Production', and maintain a comprehensive history of model development, which is essential for reliable and reproducible deployments in production environments.
Databricks SQL Warehouse
Delta Live Tables
You are deploying a RAG application. You need to ensure the model uses the most recent vector data without redeploying the model. What should you use?
Fine-tune the model with new data daily.
Implement a RAG pattern with a vector database index.
The RAG pattern retrieves relevant context from a vector database and feeds it into the model's prompt. Since the vector store can be updated independently of the model, this provides a scalable way to ensure the application always has access to the latest data without needing to redeploy models.
Re-register the model in MLflow with every update.
Embed all data into the model weights at deployment.
A team has deployed a model that is experiencing high latency. How should they identify if the bottleneck is the model inference or the preprocessing code?
Check the cluster cost in the Databricks billing console.
Use MLflow Tracing to inspect the execution pipeline.
MLflow Tracing provides detailed visibility into the duration of every step in the request flow, from input preprocessing to model inference and output post-processing. This allows engineers to pinpoint exactly where the latency is occurring, enabling them to optimize the specific components that are causing the performance delays.
Re-train the model on a larger dataset.
Increase the memory limit for the inference cluster.
Refer to the exhibit. What is the most likely cause for this error in a production RAG application?
The model is too large for the GPU memory.
Network connectivity between Model Serving and the vector store is misconfigured.
The timeout error explicitly points to a failure in establishing a connection to the vector store. This suggests that the network routing, VPC peering, or security group rules are blocking the communication path, which is a common deployment issue that must be addressed to restore RAG service functionality.
The model weights are corrupted.
The input prompt is too long for the model.
When deploying a Generative AI application, why is it recommended to use a dedicated Serving Endpoint rather than a shared interactive cluster?
Shared clusters are cheaper to run for production.
Serving endpoints provide better isolation and predictable performance.
Serving endpoints ensure that inference requests are not competing with interactive workloads or data science tasks for CPU and GPU resources. This isolation is critical for maintaining predictable, low-latency performance in production, which is a non-negotiable requirement for high-quality generative AI user experiences.
Shared clusters cannot be used for Python-based models.
Serving endpoints automatically train the model on new data.
Want more Assembling and Deploying Apps practice?
Practice this domainA data engineering team is deploying a RAG application using Mosaic AI Model Serving. They need to monitor the quality of the model's responses in production. Which Databricks feature should they use to capture and analyze inference data, such as requests, responses, and latency metrics?
Databricks SQL Alerts
Mosaic AI Inference Tables
Inference tables automatically log inference requests and responses to a Delta table in Unity Catalog. This enables seamless integration with monitoring tools for assessing model performance, latency, and throughput. It is the standardized method for capturing production data required to perform comprehensive model evaluation and drift detection.
Delta Live Tables Audit Logs
MLflow Experiment Tracking
Which of the following metrics is most effective for evaluating a RAG-based chatbot's ability to retrieve relevant context from a vector database during production monitoring?
Model Training Loss
Context Relevance
Context relevance measures how much of the retrieved information is actually necessary to answer the user query. Low scores indicate a failure in the embedding or vector search configuration, providing actionable insights for tuning the retriever component, which is a critical part of RAG evaluation.
Inference Latency
Parameter Count
A Generative AI engineer is configuring a Mosaic AI Model Serving endpoint for a production-grade LLM. Which TWO of the following tasks are necessary to ensure effective monitoring and evaluation of the endpoint? (Choose two)
Enable Inference Tables on the serving endpoint.
Enabling inference tables is the mandatory first step to store production request/response logs. Without this, you lack the raw material needed for post-hoc evaluation, analysis of failure modes, or detection of data drift. It creates a persistent data asset in Unity Catalog for ongoing oversight.
Set up an automatic retraining job on the endpoint.
Implement an automated evaluation pipeline to score inference logs.
An automated evaluation pipeline uses the captured logs to run metrics such as faithfulness or relevance. This process transforms raw logs into actionable insights about model performance. By scoring these logs, the team can identify regressions or quality degradation before they severely impact end-user experience.
Increase the GPU count to maximize inference throughput.
Delete all request logs after 24 hours to save storage.
When evaluating a generative AI model using the Mosaic AI Model Evaluation tool, what is the primary purpose of providing a 'baseline' dataset?
To increase the number of tokens the model can generate.
To provide a reference for comparative performance analysis.
A baseline allows developers to quantitatively compare current model performance against known good results. This comparison is vital for validating that new model versions or updated RAG configurations maintain or exceed performance standards, facilitating data-driven decisions about whether to promote a new version to production.
To act as a secondary training set for fine-tuning.
To optimize the vector database search latency.
Which THREE of the following are common challenges when monitoring LLM applications in production that differ significantly from traditional ML monitoring? (Choose three)
Non-deterministic output behavior.
LLMs can produce different outputs for the same prompt due to temperature settings or inherent stochasticity. This makes simple equality checks useless for evaluation. Monitoring must account for this variance, requiring probabilistic or semantic similarity metrics rather than static, exact-match validation common in traditional classification tasks.
Difficulty in defining a single 'ground truth' for responses.
In generative tasks, there are often many valid ways to phrase a correct answer. Unlike a classification model where a label is either right or wrong, LLM outputs require qualitative judgment, which is inherently more subjective and difficult to automate without robust evaluation frameworks.
Lack of high-throughput API endpoints.
High cost of manual evaluation at scale.
Human evaluation is the gold standard but is prohibitively expensive and slow at production scale. Monitoring strategies must therefore balance automated proxy metrics (like model-based evaluation) with periodic human spot-checks to ensure the automated systems remain aligned with human-perceived quality over the long term.
The model's inability to connect to internet data.
When monitoring a RAG application, you notice a high discrepancy between the retrieved context and the generated answer. Which metric would specifically help identify if the model is ignoring the provided context?
Context Precision
Faithfulness
Faithfulness specifically assesses whether the answer is logically derived from the provided context. High faithfulness indicates the model is respecting the context; low faithfulness suggests the model is generating responses based on its own training data, which leads to hallucinations and incorrect information in RAG systems.
Retrieval Recall
Semantic Similarity
Want more Evaluation and Monitoring practice?
Practice this domainA data engineer needs to ensure that sensitive PII columns are masked for specific groups while remaining visible to analysts. Which Unity Catalog feature should be used to implement this requirement?
Row-level security filters
Dynamic data masking
Dynamic data masking is specifically designed to redact or transform sensitive column data in real-time based on the user's identity or group. By attaching a masking function to a column in Unity Catalog, data engineers can ensure that analysts see masked results without modifying the underlying data files.
Credential passthrough
Attribute-based access control (ABAC)
Refer to the exhibit. The user is a member of the 'finance_team'. Why might the user encounter an access error when executing this join query?
The user lacks the SELECT privilege on the schema object.
The user lacks the USAGE privilege on the 'sales' schema.
Accessing any object in Unity Catalog requires the USAGE privilege on all containing objects, including the schema. Even if a user has explicit SELECT rights on the tables, the query will fail if they have not been granted USAGE on the parent schema container.
The user requires the MODIFY privilege to perform joins.
The user needs to be an owner of the tables to perform joins.
Which Unity Catalog object is used to link a specific cloud storage path to a catalog, schema, or table, allowing users to create tables without managing individual storage credentials?
Storage Credential
External Location
The external location object defines the specific path in cloud storage and associates it with a storage credential. It is the primary mechanism for accessing data in Unity Catalog that is stored in user-managed cloud accounts, providing a governed interface for data engineers to register data assets.
Managed Table
Unity Catalog Volume
Which THREE conditions must be met for a user to successfully create a new table in a Unity Catalog schema?
The user must have the USAGE privilege on the catalog.
Access to any object within a catalog starts with the USAGE privilege at the catalog level. Without this fundamental permission, a user cannot navigate into the catalog to access schemas or perform any operations like creating new tables, even if they have other schema-level permissions.
The user must have the USAGE privilege on the schema.
The USAGE privilege on a schema is mandatory for any interaction with objects contained within that schema. It acts as the gateway to the schema's contents, and without it, a user cannot perform DDL operations, including table creation, regardless of other privileges assigned to them.
The user must have the CREATE TABLE privilege on the schema.
Beyond simple navigation, the CREATE TABLE privilege on the schema specifically grants the user the right to define new table objects. This granular control is a core feature of Unity Catalog, allowing administrators to delegate table management responsibilities to specific users or groups securely.
The user must have the MODIFY privilege on the catalog.
The user must have the OWNER role for the entire metastore.
Which action must an administrator perform to allow a user to use Databricks SQL to query a table that is stored in an external storage location?
Grant the user the OWNER role for the storage credential.
Grant the user the READ FILES privilege on the external location.
Querying an external table requires the user to have the READ FILES privilege on the external location object that manages the underlying storage path. This is in addition to the standard catalog, schema, and table privileges, ensuring that storage access is explicitly governed.
Grant the user the ALL PRIVILEGES privilege on the metastore.
Add the user to the workspace-level admin group.
Refer to the exhibit. Why did the analyst group lose access after the table was recreated?
The catalog owner needs to refresh the table metadata.
The analyst group needs to be re-added to the schema permissions.
The new table is a different object, so the GRANT statement must be repeated.
Unity Catalog treats the newly created table as a entirely new resource with a fresh identity. Any permissions applied to the previous version of the table do not carry over to the replacement, requiring administrators to re-issue GRANT commands for the new object.
The user who recreated the table is not the catalog owner.
Want more Governance practice?
Practice this domainAn organization needs to build a RAG application on Databricks that minimizes data egress and maximizes security by keeping all data within the workspace perimeter. Which architectural pattern best satisfies this requirement?
Call external LLM APIs from a standard Python notebook without VPC constraints.
Export data to an external vector database and use a cloud-hosted LLM.
Deploy an embedding model on Mosaic AI Model Serving and utilize Databricks Vector Search.
This approach keeps all data processing, storage, and inference within the Databricks workspace perimeter. By using internal serving for embeddings and native vector search capabilities, the architecture minimizes egress traffic and simplifies security policy enforcement, ensuring that data never leaves the protected environment during the RAG retrieval process.
Use a public LLM endpoint with a public bucket to store the vector index.
Which TWO factors should be prioritized when selecting an embedding model for a domain-specific RAG application on Databricks?
The total number of parameters in the model regardless of the domain.
The semantic relevance of the model to the target domain's terminology.
Embedding models must map domain-specific terms to accurate vector representations to ensure relevant document retrieval. A model that has not been trained or fine-tuned on the specific domain’s jargon will produce poor vector alignments, leading to inaccurate RAG responses, regardless of the model's performance on general-purpose benchmarks.
The availability of the model on the public Hugging Face repository.
The computational resource requirements for inference latency.
Inference latency directly impacts the user experience and the scalability of the RAG application. Choosing a model that fits within the allotted compute resources ensures that the retrieval step does not become a bottleneck, allowing for high-throughput interactions while maintaining cost efficiency during periods of high user demand.
The color scheme of the model's documentation page.
Which THREE strategies improve the quality of retrieval in a Databricks Vector Search-based RAG application?
Implementing hybrid search using both semantic vectors and keyword-based filtering.
Hybrid search combines the strengths of semantic understanding with the precision of keyword matching. This ensures that specific technical terms or IDs are found accurately while still capturing the intent behind user queries, significantly improving retrieval quality for complex domain-specific datasets where semantic similarity alone might lead to irrelevant results.
Increasing the chunk size to include the entire dataset in a single vector.
Adding relevant metadata tags to documents to enable targeted filtering.
Metadata filtering allows the search process to narrow down the candidate document set before or during vector similarity matching. This reduces the noise in retrieval results, ensuring that only information relevant to a specific department, date, or document category is considered, thereby enhancing the relevance of retrieved content.
Optimizing the chunking strategy to maintain context boundaries.
Maintaining context boundaries, such as headers or paragraph breaks, ensures that retrieved chunks are self-contained and semantically coherent. Poor chunking can split a sentence or concept, leading to fragments that lack the necessary information for the LLM to provide a correct answer, directly undermining the quality of RAG.
Removing all stop words from the documents during the ingestion phase.
What is the primary benefit of using Unity Catalog when designing generative AI applications in Databricks?
It automatically generates Python code for model training.
It provides centralized governance, lineage, and access control for data assets.
Unity Catalog enables secure, governed access to all data used in the AI lifecycle. By tracking data lineage from the source to the model, it ensures transparency and compliance. This centralized approach simplifies security management and audit readiness, which are essential when handling proprietary data in generative AI applications.
It converts unstructured text into vectors automatically.
It eliminates the need for data preprocessing before RAG.
When designing a production RAG application, which technique is most effective for preventing the LLM from hallucinating based on outdated information?
Increasing the temperature parameter of the model to maximum.
Enforcing a streaming data pipeline to keep the vector index updated.
Keeping the vector index synchronized with the source data via a streaming pipeline ensures that the context retrieved during RAG is always current. This minimizes the risk of the model using outdated information, which is a major source of hallucinations in production systems that rely on rapidly changing business data.
Restricting the LLM to a specific list of keywords for its output.
Adding a long system prompt instructing the model not to lie.
Which Databricks feature is specifically designed to monitor model quality and drift in production?
Unity Catalog audit logs.
MLflow Model Monitoring.
MLflow Model Monitoring is the dedicated Databricks component for tracking the health of deployed models. It allows engineers to monitor metrics, detect performance degradation, and identify drift in model outputs. This is essential for maintaining high-quality generative AI applications and responding effectively to changes in data or user behavior.
The Databricks SQL query history.
Cluster event logs.
Want more Design Applications practice?
Practice this domainThe Databricks-GenAI-Assoc exam has 60–90 questions and must be completed in 120 minutes. The passing score is 700/1000.
Scenario-based questions covering exam objectives with detailed answer explanations.
The exam covers 6 domains: Data Preparation, Application Development, Assembling and Deploying Apps, Evaluation and Monitoring, Governance, Design Applications. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official Databricks Databricks-GenAI-Assoc exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.