Databricks-ML-Assoc · domain
Databricks Machine Learning
This domain covers operational ML on Databricks: MLflow experiment tracking (params, metrics, artifacts), the Model Registry, model deployment to Model Serving endpoints, and serving logs for debugging. Questions present exhibits—error messages, logs, or client tracebacks—and ask you to diagnose the root cause and select the correct resolution using Databricks and MLflow tooling.
Focused practice
Practice Databricks Machine Learning questions
Scored sessions drawing only from this domain — pick a length below.
Start 20-question practice test →What this domain covers
What to know about Databricks Machine Learning
Be able to trace a deployment or inference failure from the exhibit to its cause using MLflow and Model Serving concepts. The most important thing: read the exact error or log message and match it to the correct Databricks component and fix.
Using MLflow tracking to log parameters, metrics, and model artifacts during training runs
Registering models in the MLflow Model Registry and managing model versions and stages
Deploying MLflow models to Databricks Model Serving endpoints and configuring the endpoint
Reading Model Serving endpoint logs to diagnose connection and inference failures
Watch out for
Common Databricks Machine Learning exam traps
- ▸Assuming an inference error is a model-quality problem when the exhibit actually shows a request schema, dependency, or environment mismatch
- ▸Confusing MLflow tracking (logging runs) with the Model Registry (versioning and deployment), and choosing the wrong component
- ▸Ignoring the endpoint or serving logs and guessing at causes instead of reading the actual error text shown in the exhibit
Question index
All Databricks Machine Learning questions (79)
Click any question to see the full explanation, or start a practice session above.
A data scientist is using MLflow to log a custom PyTorch model in a Databricks notebook. They want to register the model in the Databricks Model Registry and later serve it with MLflow model serving. Which function should they call within their MLflow run to log the model with the necessary signature and dependencies?
Medium2A machine learning engineer is troubleshooting a Model Registry issue where models are not being transitioned correctly. Which TWO actions should the engineer take to ensure proper governance and automated testing in the Registry?
Hard3A data science team is transitioning from local model development to Databricks. They want to ensure their models are portable across different Databricks workspaces. What is the recommended practice for managing models in this environment?
Medium4Which Databricks tool is primarily used for organizing and documenting experiments, tracking parameters, and versioning models during the machine learning lifecycle?
Medium5A machine learning engineer wants to register a model in the Databricks Model Registry using MLflow. They have already trained a model and logged it with MLflow. Which method should they use to register the model programmatically?
Easy6A data science team is using Databricks Repos to manage a machine learning project. They want to ensure that their notebooks and supporting modules are version-controlled and that they can collaborate without overwriting each other's changes. Which TWO practices should they follow? (Choose two.)
Hard7A data scientist has trained a scikit-learn model and logged it with MLflow. They now want to register the model in the Databricks Model Registry and transition it to the 'Production' stage. Which sequence of MLflow API calls should they use?
Medium8When sharing a Databricks ML experiment with another team, what is the best way to ensure they can reproduce your results exactly?
Medium9Which Databricks artifact should be used to encapsulate a model, its environment dependencies, and the required code to ensure consistent model behavior across different deployment environments?
Hard10A team wants to compare multiple model runs for a fraud detection project, view their metrics side by side, and identify which run produced the best area under the ROC curve. They have already logged each run with MLflow. Which Databricks capability should they use to perform this comparison?
Easy11A machine learning team is using Databricks Feature Store to serve features for online inference. They need to ensure that the online store remains consistent with the offline store and supports low-latency lookups. Which two practices should they follow? (Choose two.)
Hard12A data scientist has trained a model using scikit-learn and wants to log it to MLflow for deployment. They need to ensure that the model can be served with the correct dependencies. Which MLflow function should they use to log the model?
Easy13A data scientist is using MLflow to track experiments on Databricks. They want to log a custom metric that is computed during model training and later compare it across runs using the MLflow UI. Which MLflow API call should they use?
Medium14What is the primary benefit of using 'AutoML' in Databricks for a machine learning project?
Easy15A team is building a Feature Store in Databricks. What is the primary advantage of using the Feature Store for training models compared to using raw Delta tables?
Medium16Which approach is most efficient for deploying a high-throughput, low-latency model in Databricks?
Medium17A data scientist wants to track the performance of a model training run in Databricks. They use MLflow to log parameters and metrics. After the run, they need to view all runs for the experiment in a web-based interface. Which URL should they navigate to?
Easy18A machine learning engineer is using Databricks Feature Store to create a training dataset for a model that predicts customer lifetime value. The feature table includes a timestamp key. Which TWO statements are true regarding point-in-time correctness when creating the training set? (Choose two.)
Hard19Refer to the exhibit. A model is failing to deploy to Databricks Model Serving. The error log indicates 'Schema Mismatch'. Based on the JSON signature provided, what is the most likely cause of the deployment failure?
Hard20A machine learning engineer needs to track model parameters, metrics, and artifacts across distributed training runs executed on Databricks. Which component of Databricks Machine Learning should they use to manage and organize this experiment metadata?
Medium21A data scientist needs to perform hyperparameter tuning using Hyperopt. Which THREE components are essential to successfully implement an automated tuning run on a Databricks cluster?
Medium22A data scientist needs to prepare a model for deployment in a highly regulated environment. Which TWO tasks must they complete to ensure the model meets auditability requirements?
Medium23Refer to the exhibit. A machine learning engineer deployed an MLflow model to Databricks Model Serving, but inference requests are failing with the error shown in the exhibit. How should the engineer resolve this issue?
Hard24A data science team is preparing to deploy a custom machine learning model to Databricks Model Serving. Which TWO steps are required to ensure the model can successfully load and serve predictions using MLflow? Choose 2 answers.
Medium25Refer to the exhibit. A user attempts to load a model using the MLflow Python API, but the load fails. Based on the JSON snippet, what is the most likely issue?
Hard26A data scientist is using Databricks AutoML to train a classification model on a dataset with a highly imbalanced target variable. They want to ensure the model evaluation focuses on the minority class. Which evaluation metric should they prioritize when interpreting AutoML results?
Hard27A data scientist is training a scikit-learn model on Databricks and wants to automatically log parameters, metrics, and models without writing explicit MLflow logging code. Which approach should they use?
Easy28A data scientist is using MLflow to track experiments on Databricks. They notice that the metrics logged during a run are not appearing in the MLflow UI. The run is part of an experiment with many runs. What is the most likely cause for the missing metrics?
Hard29Which Databricks feature should be used to provide a managed, secure, and scalable endpoint for real-time inference of models logged in the Model Registry?
Medium30A data scientist is training a model using MLflow on Databricks. They need to ensure that the model artifacts, environment dependencies, and signature are automatically captured to facilitate seamless deployment to Databricks Model Serving. Which command should they use within the training script?
Medium31A data scientist is using Databricks Feature Store to create a feature table for a recommendation model. They want to ensure that the same feature computation logic is used both during training and at inference time to avoid training-serving skew. Which Feature Store capability directly addresses this requirement?
Easy32A data scientist is training a model using MLflow on Databricks and needs to ensure that all parameters and metrics are logged for every training run. Which approach ensures the most reliable logging of artifacts and metrics during model training?
Medium33Refer to the exhibit. The logs indicate a persistent connection failure for a Databricks Model Serving endpoint. What is the most likely cause?
Hard34A machine learning engineer registers a model in the Databricks Model Registry and wants to serve it with low-latency online inference. The model's Python dependencies include a custom private library that is not publicly available. Which deployment approach should the engineer use to ensure the private library is available at inference time?
Hard35Which component in Databricks is used to manage the lineage of machine learning data, ensuring that you can trace a model back to the exact version of the data it was trained on?
Medium36A data scientist has trained a scikit-learn model and wants to log it to MLflow with a custom signature that includes input and output schema. Which MLflow method should they use to log the model along with the signature?
Medium37A machine learning engineer is using MLflow on Databricks to track an experiment. They want to record the model's hyperparameters and evaluation metrics, but they do not want to save the trained model artifact. Which MLflow API calls should they use?
Medium38A machine learning engineer is orchestrating a multi-step training workflow on Databricks using Databricks Jobs. The workflow includes data preprocessing, model training, and evaluation. The engineer needs to ensure that the evaluation step runs only if the training step completes successfully, and that the preprocessing step runs first. Which feature of Databricks Jobs should be used to define these dependencies?
Medium39Which Databricks ML component is best suited for managing access control for machine learning experiments and models across different teams?
Easy40Which THREE of the following are supported methods for serving machine learning models in Databricks?
Hard41A data scientist is working in a Databricks notebook and wants to use MLflow to log a trained scikit-learn model. They want to ensure that the model can be loaded later for inference. What is the correct MLflow function to log the model?
Easy42A data scientist is training a machine learning model on Databricks and needs to ensure that every experiment run is automatically tracked, including parameters, metrics, and model artifacts. Which component should the scientist use to achieve this with minimal code changes?
Medium43A data scientist has registered a model in the Databricks Model Registry. They want to transition the model from 'Staging' to 'Production' but need to ensure that only specific users can perform this transition. Which Databricks feature should they use to enforce this access control?
Medium44A data scientist needs to track parameters, metrics, and model artifacts during training on Databricks. Which component is the primary tool for managing the entire lifecycle of these ML experiments?
Medium45Why is it important to use a 'Feature Store' rather than joining raw tables directly in the training notebook?
Medium46A data scientist is monitoring model drift in Databricks. Which TWO approaches are recommended to detect performance degradation in a production model?
Hard47A data scientist is using MLflow to track experiments on Databricks. They want to compare multiple runs and identify the best performing model based on a custom metric. Which TWO features of MLflow can be used to achieve this? (Choose two.)
Hard48A data scientist is using MLflow on Databricks to log a scikit-learn model. They call mlflow.sklearn.log_model(model, 'model') and then inspect the run. They notice the model artifact is stored, but the run does not appear in the Models page of the workspace. They did not call any model registration function. What is the most likely reason the model is not listed in the Models page?
Medium49What is the primary function of the 'Model Signatures' in MLflow?
Easy50When evaluating a machine learning model, what is the main purpose of creating a separate evaluation dataset in Databricks?
Medium51Which TWO of the following are primary benefits of using the Databricks Feature Store for machine learning workflows?
Medium52A data scientist is using Databricks Feature Store to build training sets and wants to ensure the features used at training time are consistent with those served at inference time. Which TWO practices help guarantee this consistency? (Choose two.)
Medium53Which Databricks component should be used to track parameters, code versions, metrics, and output files when running machine learning experiments?
Easy54A machine learning engineer is training a model using Databricks AutoML. They notice that the generated notebook includes a step that uses Hyperopt for hyperparameter tuning, but the tuning process is taking too long. They want to reduce the search space without sacrificing model performance significantly. Which Hyperopt configuration change should they make?
Hard55Which Databricks feature provides a managed environment specifically optimized for machine learning libraries like TensorFlow, PyTorch, and XGBoost?
Easy56A user is experiencing 'Out of Memory' (OOM) errors during the evaluation phase of a large XGBoost model on Databricks. What is the most effective way to address this while utilizing the distributed nature of Databricks?
Medium57A data scientist is using MLflow to train a scikit-learn model on Databricks. They call mlflow.sklearn.autolog() before fitting the model. After the run completes, they need to retrieve the automatically logged model and load it for batch inference in a separate notebook. Which approach correctly retrieves the logged model for loading?
Medium58A data engineer is working with a large Delta table and wants to optimize it for machine learning feature engineering. They frequently filter data by a column named 'event_date' and join on a column named 'user_id'. Which Delta Lake feature should they use to improve query performance?
Easy59A machine learning engineer is using Databricks Feature Store to create a training dataset. They want to ensure that the features used during training are exactly the same as those served at inference time. Which Feature Store capability should they rely on?
Hard60Refer to the exhibit. A data scientist is attempting to deploy a model using the MLflow client. The error above occurs during the deployment script. What is the most likely cause of this failure?
Medium61A data scientist registered a model in Unity Catalog and now wants to serve it as a low-latency REST endpoint for an application. They need automatic scaling, a secure endpoint URL, and the ability to update the served model version without redeploying infrastructure. Which Databricks capability should they use?
Hard62When a data scientist needs to share an MLflow experiment with a team member, what is the best practice for ensuring collaborative access within Databricks?
Medium63A data scientist wants to speed up the process of finding the optimal hyperparameters for a machine learning model. Which Databricks-supported library is optimized for distributed hyperparameter tuning?
Medium64A data scientist is using Databricks Feature Store to create a feature table for a recommendation model. They want to ensure that the feature table can be used for both training and batch scoring, and that it supports point-in-time correctness. Which two actions must they take when creating the feature table? (Choose two.)
Hard65What is the primary technical limitation when deploying an MLflow model that has custom Python dependencies not included in the standard Databricks Runtime?
Hard66A machine learning engineer needs to track hyperparameter tuning runs and log artifacts using MLflow inside a Databricks Notebook. Which approach should be used to ensure runs are automatically nested under a parent run?
Medium67A machine learning engineer is deploying a model using MLflow Model Serving on Databricks. They want to ensure the endpoint can handle bursts of traffic and automatically scale. Which configuration should they set when creating the served model?
Medium68A machine learning engineer is using Databricks Feature Store to create a feature table for a model that predicts customer churn. The feature table includes customer demographics and transaction history. The engineer wants to ensure that the model can access the latest feature values during online inference. What should the engineer do?
Medium69Which approach is most efficient for handling high-cardinality categorical features in a machine learning model while maintaining compatibility with standard Databricks model serving?
Hard70A data scientist is using MLflow on Databricks to track experiments. After running several training jobs, they notice that the run metrics are recorded but the model artifact is missing when they view the run details. They logged the model using the default MLflow API. What is the most likely cause?
Medium71What is the primary function of the 'Model Registry' in the Databricks ML ecosystem?
Medium72A machine learning engineer is using MLflow Tracking on Databricks to compare multiple runs. They want to programmatically retrieve the best run based on a metric called 'rmse' from an experiment. Which MLflow API call should they use?
Medium73A data scientist has registered a model in the Databricks Model Registry and wants to deploy it as a REST API endpoint for real-time inference. Which Databricks feature should they use?
Easy74Refer to the exhibit. A user attempts to transition a model to production, but the code fails with the provided error. What is the most likely cause?
Hard75Which approach is most efficient for tracking hyperparameter tuning metrics across thousands of runs on Databricks?
Medium76A data scientist is performing hyperparameter tuning for a scikit-learn model using MLflow on a Databricks cluster. They want to parallelize the tuning to reduce total runtime. Which Databricks-supported method allows them to run multiple trials concurrently with minimal code changes?
Medium77A machine learning engineer is training a scikit-learn model on a Databricks cluster and wants every hyperparameter, evaluation metric, and the fitted estimator itself to be captured automatically without writing custom logging code. The engineer uses MLflow with autologging enabled. Which statement best describes what MLflow autologging records for this scikit-learn run?
Medium78A machine learning engineer is orchestrating a multi-step training workflow with Databricks Jobs. The workflow must retrain a model, evaluate it, and only register it if a quality threshold is met, otherwise stop without registering. Which approach best implements this conditional logic within the job?
Hard79When preparing data for machine learning in Databricks, which feature of Delta Lake is most beneficial for managing large-scale datasets during the training process?
EasyOther domains
All Databricks-ML-Assoc exam domains
Frequently asked questions
- What does the Databricks Machine Learning domain cover on the Databricks-ML-Assoc exam?
- Be able to trace a deployment or inference failure from the exhibit to its cause using MLflow and Model Serving concepts. The most important thing: read the exact error or log message and match it to the correct Databricks component and fix.
- How many questions are in this domain?
- This page lists all 79 Databricks Machine Learning questions in the Databricks-ML-Assoc question bank. The actual exam draws from this domain proportionally to its weighting in the official exam blueprint.
- What is the best way to practise this domain?
- Start with a short focused session (10 questions) to identify gaps, then work through explanations. Repeat with a longer session once the weak areas feel solid.
- Can I practise only Databricks Machine Learning questions?
- Yes — the session launcher on this page filters questions to this domain only. Choose any session length for inline explanations and scoring.