Sample questions
Databricks Certified Machine Learning Associate practice questions
Your team is experiencing 'data drift' in production where the model's accuracy drops over time. What is the most recommended Databricks-native approach to address this?
A machine learning engineer is training a model on a Databricks cluster and wants the training code to run inside a container that they control, with the same Python libraries avai…
When logging a model, what is the significance of the 'code_path' parameter in mlflow.log_model?
A machine learning engineer has an existing Databricks Model Serving endpoint named churn-endpoint serving version 3 of a model. The team has validated version 5 and wants to direc…
Refer to the exhibit. What is the most likely cause of this error in a deployed MLflow model?
Refer to the exhibit. A machine learning engineer deployed an MLflow model to Databricks Model Serving, but inference requests are failing with the error shown in the exhibit. How…
What is the primary advantage of using Databricks Model Serving over deploying a model on a standalone web server?
Refer to the exhibit. A Databricks job failed to start, returning the error shown. The job depends on MLflow for tracking. What is the most likely cause of this failure?
A data scientist is preparing a feature table in Databricks Feature Store. To ensure the feature table can be used for online inference with low latency, which step is mandatory?
Refer to the exhibit. The logs indicate a persistent connection failure for a Databricks Model Serving endpoint. What is the most likely cause?
A data scientist is using Databricks to train a deep learning model. They need to monitor training loss and accuracy in real-time. Which tool is best suited for this task?
A data scientist needs to track parameters, metrics, and model artifacts during training on Databricks. Which component is the primary tool for managing the entire lifecycle of the…
A data scientist is training a model using MLflow on Databricks. They need to ensure that the model artifacts and metrics are logged automatically without adding manual logging cod…
A data scientist is comparing multiple hyperparameter configurations for a model and wants to view the resulting metrics side by side in a single interface, sort runs by accuracy,…
Refer to the exhibit. A data scientist is attempting to deploy a model using the MLflow client. The error above occurs during the deployment script. What is the most likely cause o…
A machine learning engineer is using MLflow to track experiments. They want to compare multiple runs and identify the run that produced the best model based on a custom metric call…
A data scientist is training a deep learning model on Databricks using Horovod for distributed training. They find that the model is converging slowly. What is the most likely caus…
A data scientist registered a model in Unity Catalog and now wants to serve it as a low-latency REST endpoint for an application. They need automatic scaling, a secure endpoint URL…
A data scientist is using Spark MLlib and wants to perform feature scaling on a large dataset. Which transformer should they use within a Pipeline to ensure that the scaling logic…
A data scientist is using Databricks Feature Store to create a feature table for a machine learning model. They want to ensure that the features used during training are consistent…
When a data scientist needs to share an MLflow experiment with a team member, what is the best practice for ensuring collaborative access within Databricks?
An ML engineer wants to package a training project so it can be run reproducibly from a Databricks Job across environments, with its Python dependencies and entry point defined. Wh…
What is the primary technical limitation when deploying an MLflow model that has custom Python dependencies not included in the standard Databricks Runtime?
A data scientist is using Databricks AutoML to train a classification model on a dataset with a binary target. They want to understand which features contributed most to the model'…