Databricks · Free Practice Questions · Last reviewed May 2026
24real exam-style questions organised by domain, each with the correct answer highlighted and a plain-English explanation of why it's right — and why the others are wrong.
A data scientist is building a pipeline using MLflow. Which TWO of the following tasks should be performed to ensure experiment reproducibility and auditability?
Log the git commit hash using mlflow.set_tag().
Logging the git commit hash directly links the model to the exact state of the source code. This is a best practice for tracking changes over time, as it allows developers to revert to specific training configurations and understand the lineage of the model artifacts within the MLflow Tracking server.
Manually copy the training data into the MLflow model folder.
Capture the execution environment using log_model(conda_env=...).
Logging the environment dependencies ensures that when the model is loaded in a different inference environment, the correct library versions are present. This prevents runtime errors and unexpected behavior due to library version drift, which is critical for consistent model performance in production and staging environments.
Disable the automatic logging feature to save memory.
Use a global variable for all experiment parameters.
When performing hyperparameter tuning using Hyperopt with SparkTrials on Databricks, what is the primary advantage of using SparkTrials over the standard Trials object?
SparkTrials automatically selects the best hyperparameters.
SparkTrials allows for the distribution of training jobs across multiple workers.
SparkTrials distributes the trials across the Spark cluster, allowing multiple hyperparameter configurations to be tested concurrently. This dramatically shortens the search time for complex models, making it a critical tool for scaling machine learning experiments in environments where compute resources are available but time-to-market is the primary constraint.
SparkTrials provides built-in visualization of the parameter space.
SparkTrials forces the use of a GPU-enabled cluster.
A model has been trained on a dataset containing categorical features with high cardinality. Which technique is most effective for preparing these features for a linear model in a Databricks environment?
One-hot encode all categorical variables.
Use Target Encoding to map categories to their mean target value.
Target encoding is highly effective for high-cardinality features, as it transforms categorical levels into a continuous numerical representation based on the target variable. This reduces dimensionality while retaining meaningful information, which is ideal for linear models that perform best when features are appropriately scaled and numeric.
Drop the categorical features entirely.
Convert all features into a single string representation.
Which Databricks feature allows data scientists to automatically track parameters, metrics, and models during the training process?
Delta Lake.
MLflow Tracking.
MLflow Tracking is specifically designed to record and query experiment results, including parameters, metrics, and artifacts. It serves as the single source of truth for the model development history, allowing data scientists to identify the best-performing models easily and maintain a clean audit trail for deployment.
Databricks Jobs.
Unity Catalog.
Refer to the exhibit. Why is including the `signature` and `input_example` in the `log_model` call considered a professional best practice?
It forces the model to run on a GPU during inference.
It speeds up the model training process.
It enables schema validation and improves model serving usability.
Signatures provide a clear contract for the model, enabling automatic validation of input data. Input examples help serving tools generate accurate API definitions, making it easier for external applications to call the model. This reduces integration friction and ensures that only valid data is passed to the model.
It reduces the storage size of the model artifact.
A data scientist is using Databricks to train a deep learning model. They need to monitor training loss and accuracy in real-time. Which tool is best suited for this task?
Databricks File System (DBFS) logs.
MLflow log_metric API.
The MLflow log_metric API allows for step-wise tracking of performance metrics. This is the standard way to monitor deep learning models in Databricks, as it provides a clean, web-based UI to plot the training curves and compare the performance of different runs in real-time during the development process.
Spark UI metrics tab.
Unity Catalog lineage.
Want more Model Development practice?
Practice this domainA data scientist is training a model using MLflow on Databricks and needs to ensure that all parameters and metrics are logged for every training run. Which approach ensures the most reliable logging of artifacts and metrics during model training?
Manually call mlflow.log_metric for every individual iteration inside the training loop.
Configure the MLflow tracking URI to point to an external database before initiating the training.
Invoke mlflow.autolog() at the beginning of the notebook cell prior to training.
Invoking mlflow.autolog() enables automatic logging for supported libraries like Scikit-learn, PyTorch, or XGBoost. This captures parameters, metrics, and models without manual intervention. It is the best practice for ensuring full visibility into experiment runs, supporting the Databricks requirement for reliable and reproducible machine learning experimentation.
Use the model.save() method instead of MLflow tracking for better persistence.
A machine learning engineer is troubleshooting a Model Registry issue where models are not being transitioned correctly. Which TWO actions should the engineer take to ensure proper governance and automated testing in the Registry?
Implement Webhooks to trigger external CI/CD validation pipelines upon status transitions.
Webhooks allow Databricks to trigger external services like Jenkins or GitHub Actions whenever a model version changes state. This enables automated testing and validation workflows, ensuring that models meet performance benchmarks and quality gates before being approved for staging or production, which is a critical governance requirement.
Delete all older versions of the model to keep the registry clean and performant.
Manually update the model stage via the UI for every version to ensure maximum control.
Use Model Registry tags to perform metadata-based filtering for downstream automated testing.
Tags provide a powerful mechanism to annotate model versions with metadata such as 'test_passed', 'model_owner', or 'data_version'. Automated pipelines can query these tags to determine if a model is ready for transition, providing a scalable and metadata-driven approach to model governance and lifecycle management.
Disable the model versioning feature to save storage space in the DBFS.
Refer to the exhibit. A data scientist is attempting to deploy a model using the MLflow client. The error above occurs during the deployment script. What is the most likely cause of this failure?
The cluster does not have the necessary MLflow libraries installed.
The model has not been registered in the Model Registry with the specified name and version.
The RestException explicitly confirms that the registry lookup failed for the provided name and version. This indicates a mismatch between the deployment configuration and the actual state of the registry. The user must verify the registration status via the UI or list_model_versions to confirm existence.
The user lacks permissions to read the model version from the Registry.
The model version is still in the 'Archived' state and cannot be accessed.
When preparing data for machine learning in Databricks, which feature of Delta Lake is most beneficial for managing large-scale datasets during the training process?
Compaction (Auto Optimize) for improved query speed during data reading.
Time Travel (versioning) to access previous snapshots of training data.
Time Travel enables access to specific versions of the data, which is critical for model reproducibility. By referencing a specific timestamp or version, data scientists can guarantee that their training experiments are conducted on the exact same data state, facilitating auditability and consistency across different model training iterations.
Schema enforcement to prevent corrupted data from entering the training set.
Z-Ordering to speed up filtering on specific columns in the dataset.
Which Databricks feature should be used to provide a managed, secure, and scalable endpoint for real-time inference of models logged in the Model Registry?
Databricks Jobs running the model in a scheduled notebook.
Databricks Model Serving endpoints.
Model Serving endpoints are purpose-built for low-latency, real-time model inference. They manage the containerization and infrastructure deployment automatically, ensuring that models are accessible via secure REST APIs. This is the optimal Databricks-native solution for deploying models into production environments that require immediate, scalable prediction capabilities.
A standard interactive cluster running a Flask server.
Delta Live Tables pipelines.
A data scientist needs to perform hyperparameter tuning using Hyperopt. Which THREE components are essential to successfully implement an automated tuning run on a Databricks cluster?
A defined objective function that minimizes or maximizes a metric.
The objective function is the core of the tuning process. It takes parameters as input, trains the model, and returns a scalar value (e.g., loss or accuracy) that Hyperopt aims to optimize. Without this function, the algorithm cannot evaluate the effectiveness of different parameter configurations during the search.
A search space defining the range and distribution of hyperparameters.
The search space defines the bounds and types of hyperparameters (e.g., uniform, log-uniform, or choice) that Hyperopt will explore. It provides the necessary structure for the algorithm to sample candidates. Without a clearly defined search space, the optimization algorithm has no domain in which to search for better model parameters.
A pre-trained model checkpoint to initialize the search.
An optimization algorithm (e.g., fmin, tpe.suggest).
The optimization algorithm dictates how Hyperopt selects the next set of hyperparameters to test. Choosing an appropriate algorithm, such as Tree-structured Parzen Estimator (TPE), is crucial for efficiently exploring the search space and converging on the optimal parameter set faster than random search methods would allow.
A dedicated GPU cluster for every single trial run.
Want more Databricks Machine Learning practice?
Practice this domainA data scientist needs to deploy a model to Databricks Model Serving. Which component is strictly required to be logged in MLflow to enable the 'Model Serving' feature?
The model's training accuracy metrics
The model signature defining input and output schema
The model signature provides the necessary schema metadata for the serving endpoint. This allows Databricks to enforce input validation for all incoming REST API requests. By defining the signature during the log_model call, you ensure the serving container knows how to translate JSON payloads into the correct format.
A Unity Catalog registered function
A dedicated high-concurrency cluster
When deploying a model using Model Serving, how does Databricks ensure that the environment remains consistent between the training workspace and the serving environment?
By requiring the user to provide a Dockerfile
By capturing the model's dependencies during log_model
MLflow automatically logs the environment dependencies, including Python packages and versions, when log_model is called. The serving infrastructure reads this metadata to rebuild the environment in the container. This ensures that the code runs in an environment identical to the one used during training and testing phases.
By strictly enforcing the use of the latest stable libraries
By running the model on the same training cluster
Which strategy is most effective for managing model drift in a production Databricks environment?
Manually check the model accuracy once a year
Re-deploy the model with new data every hour
Log inference data to Delta tables for analysis
Logging inference inputs and predictions to Delta tables creates an audit trail that enables monitoring. This data can be analyzed to measure drift in input features or model outputs. This is the industry-standard approach in Databricks for building a feedback loop that informs when a model needs to be updated.
Use a static model that never changes
Which THREE factors should be considered when choosing the 'workload size' (e.g., Small, Medium, Large) for a Databricks Model Serving endpoint?
The memory requirements of the model artifact
The model artifact size and its runtime memory consumption directly dictate the minimum compute resources needed. If the model requires more RAM than the chosen workload size provides, the endpoint will fail to load or experience frequent crashes. Assessing memory usage during the testing phase is critical for size selection.
The total number of users in the workspace
Expected request latency targets
Different workload sizes provide varying CPU and memory resources, which directly impact the inference latency. If a business requirement demands low latency for a model that is computationally heavy, selecting a larger workload size can help meet these targets by providing more compute power for the inference operation.
The volume of expected incoming traffic
High request throughput necessitates higher compute allocation to ensure the queue does not grow and cause latency spikes. The workload size defines the capacity per container, and autoscaling will then manage the number of replicas. Choosing the right size ensures that each replica handles its share of traffic efficiently.
The color scheme of the MLflow UI
When deploying a model to a production endpoint, what is the best practice for handling dependencies?
Manually install dependencies in the cluster terminal
Include a requirements.txt file in the model artifact
Including a requirements.txt file or letting MLflow capture the environment ensures that the serving environment mirrors the training environment. This is the standard method to maintain consistency, allowing the Databricks serving infrastructure to install the correct package versions during the initial container build and deployment process.
Use the 'latest' tag for all library imports
Only use libraries that come pre-installed in the Databricks Runtime
Which Databricks feature allows you to manage the lifecycle of a model, including transitions from 'Staging' to 'Production'?
Databricks SQL
MLflow Model Registry
The MLflow Model Registry is the purpose-built service in Databricks for managing the entire model lifecycle. It allows for versioning, metadata tracking, and stage transitions, providing an organized approach to moving models from initial experimentation through staging to final deployment in a production serving endpoint.
Delta Live Tables
Unity Catalog
Want more Model Deployment practice?
Practice this domainA data scientist is training a model using MLflow on Databricks. They need to ensure that the model artifacts and metrics are logged automatically without adding manual logging code to the training script. Which approach should they use?
Call mlflow.set_tracking_uri() inside the training function to redirect logs to the workspace.
Wrap the training logic within an mlflow.start_run() block without any additional configuration.
Execute mlflow.autolog() at the start of the notebook cell before running the model training code.
Calling mlflow.autolog() enables the library-specific hooks that automatically record metrics, parameters, and artifacts during the execution of supported model training methods. This is the standard practice for Databricks ML workflows to reduce manual instrumentation overhead while ensuring that every model iteration is fully documented and tracked automatically.
Configure the Databricks cluster environment variable MLFLOW_TRACKING_ENABLED to true.
Refer to the exhibit. A Databricks job failed to start, returning the error shown. The job depends on MLflow for tracking. What is the most likely cause of this failure?
The Databricks Runtime version is too recent and contains a breaking change.
The cluster library configuration is missing the required MLflow package.
This error occurs when the driver or worker nodes lack the MLflow library. Since the job fails during initialization, it indicates that the environment definition for the job cluster does not include MLflow, preventing the code from accessing the tracking client at runtime during the start-up sequence.
The MLflow tracking server is currently unreachable due to network security policies.
The job cluster has insufficient memory to load the MLflow dependency.
When designing an ML workflow, what is the primary benefit of using MLflow Projects over executing raw scripts?
They automatically scale the underlying compute cluster size based on the task.
They allow for the automatic versioning and tracking of data snapshots.
They facilitate consistent execution by encapsulating environment and dependency definitions.
MLflow Projects use a project specification file to define dependencies, allowing the environment to be recreated reliably. This ensures that the same code runs identically on any environment, which is vital for professional ML workflows where consistency between development, staging, and production environments is mandatory for reliable model results.
They replace the need for unit testing individual functions within the pipeline.
Which THREE actions are best practice when deploying a machine learning model using Databricks Model Serving?
Hardcode API credentials directly into the model inference script.
Ensure the model signature is defined to enable input validation.
Defining a model signature allows the serving endpoint to validate incoming request data against the expected schema. This prevents runtime errors and unexpected model behavior by rejecting malformed input, which is a crucial safeguard for stable and reliable production-grade ML inference services in a distributed environment.
Utilize the Model Registry to manage the versioning of the deployed model.
The Model Registry acts as the source of truth for Model Serving. By pointing the serving endpoint to a specific registered model version or alias, you ensure that deployments are reproducible and traceable, allowing for seamless rollbacks and controlled updates to production models without manual configuration changes.
Perform inference on the same cluster used for training to save costs.
Implement logging within the inference function to monitor performance.
Logging within the inference function is critical for observability. It allows teams to monitor latency, error rates, and data drift, which are necessary to detect performance degradation in production. Proper logging enables proactive maintenance and ensures that models continue to meet accuracy and performance requirements over time.
You are building a pipeline in Databricks and need to ensure that a training job only runs after the upstream data preparation job has successfully completed. Which Databricks feature should you use?
Use the mlflow.end_run() command at the end of the data prep script.
Use Delta Live Tables (DLT) expectations to trigger the job.
Use Databricks Workflows task dependencies.
Databricks Workflows enables users to build multi-task jobs where tasks are linked by success dependencies. This is the native and most efficient way to ensure that a training task waits for the data preparation task to complete, maintaining a strict and reliable execution order within the pipeline.
Use a cron schedule with a 30-minute delay for the training job.
Your team is experiencing 'data drift' in production where the model's accuracy drops over time. What is the most recommended Databricks-native approach to address this?
Increase the size of the serving cluster to handle more data points.
Implement a retraining pipeline that is triggered when performance metrics drop.
A robust ML workflow includes continuous monitoring and an automated retraining loop. When performance metrics drop, a pipeline should be triggered to retrain the model on the most recent data. This effectively mitigates data drift by keeping the model updated with the current characteristics of the production environment.
Hard-code the input feature ranges in the inference function to filter out outliers.
Switch to a more complex model architecture to better fit the production data.
Want more ML Workflows practice?
Practice this domainThe Databricks-ML-Assoc exam has 60–90 questions and must be completed in 120 minutes. The passing score is 700/1000.
Scenario-based questions covering exam objectives with detailed answer explanations.
The exam covers 4 domains: Model Development, Databricks Machine Learning, Model Deployment, ML Workflows. Questions are weighted by domain — higher-weight domains appear more on your actual exam.
No. These are original exam-style practice questions written against the official Databricks Databricks-ML-Assoc exam objectives. They are not copied from the real exam. Courseiva focuses on genuine understanding, not memorisation of braindumps.
Courseiva tracks your accuracy per domain and routes you toward weak areas automatically. Free, no account required.