Courseiva

Databricks-ML-Assoc · topic practice

Databricks Machine Learning practice questions

This domain covers operational ML on Databricks: MLflow experiment tracking (params, metrics, artifacts), the Model Registry, model deployment to Model Serving endpoints, and serving logs for debugging. Questions present exhibits—error messages, logs, or client tracebacks—and ask you to diagnose the root cause and select the correct resolution using Databricks and MLflow tooling.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Databricks Machine Learning

What the exam tests

What to know about Databricks Machine Learning

Be able to trace a deployment or inference failure from the exhibit to its cause using MLflow and Model Serving concepts. The most important thing: read the exact error or log message and match it to the correct Databricks component and fix.

Using MLflow tracking to log parameters, metrics, and model artifacts during training runs

Registering models in the MLflow Model Registry and managing model versions and stages

Deploying MLflow models to Databricks Model Serving endpoints and configuring the endpoint

Reading Model Serving endpoint logs to diagnose connection and inference failures

Watch out for

Common Databricks Machine Learning exam traps

  • ▸Assuming an inference error is a model-quality problem when the exhibit actually shows a request schema, dependency, or environment mismatch
  • ▸Confusing MLflow tracking (logging runs) with the Model Registry (versioning and deployment), and choosing the wrong component
  • ▸Ignoring the endpoint or serving logs and guessing at causes instead of reading the actual error text shown in the exhibit

Practice set

Databricks Machine Learning questions

20 questions · select your answer, then reveal the explanation

Refer to the exhibit. An administrator is configuring a Unity Catalog access policy to restrict experiment deletion. Given the JSON snippet provided, what is the impact of this policy configuration?

Exhibit

JSON: { "policy": "deny", "actions": ["mlflow:DeleteExperiment"], "effect": "allow" }

When logging a model with MLflow, which THREE of the following are essential components of a model signature?

An ML engineer wants to compute features with point-in-time correctness to prevent data leakage when training a fraud detection model. Which Databricks Feature Store capability should be utilized?

A team is designing a feature engineering pipeline in Databricks. Which TWO considerations are essential when using Feature Store to ensure data consistency between training and inference?

A machine learning engineer needs to tune a deep learning model using Hyperopt on Databricks. Which THREE steps are required to implement distributed hyperparameter tuning effectively?

A team is building a real-time churn prediction system. Which THREE requirements are critical for the Online Store in this Databricks Feature Store architecture?

During a model deployment test, you notice that the model's inference performance is inconsistent. What is the most likely cause, and how can you mitigate it?

A data scientist wants to share their notebook with a team member, but they want to restrict access so the team member can only view the results without modifying the code. What is the correct sharing permission setting?

An organization is preparing to deploy a machine learning model into production using Databricks Model Serving. Which TWO actions must the user perform to ensure the model is ready for high-availability inference?

A machine learning engineer has a scikit-learn model registered in Unity Catalog as `prod_ml.churn.final_model` with version 3 tagged `"champion"`. The batch scoring job reads the model with `mlflow.pyfunc.spark_udf(spark, model_uri="models:/prod_ml.churn.final_model/3")`. When a data scientist registers version 4 and sets the alias `champion` to it, the batch job starts scoring with version 4 without any code change. The engineer must guarantee the batch job always scores with the reviewed version 3 until a new version is explicitly approved. What should the engineer do?

A machine learning engineer is using Databricks Feature Store to create a training dataset. They define a feature table with a primary key and a timestamp column. They then use create_training_set to generate a training set. They notice that the training set includes features as of the latest timestamp, not as of the timestamp of each label event. What is the most likely cause of this issue?

A data scientist has registered a model in the Databricks Model Registry and wants to deploy it as a real-time endpoint using MLflow Model Serving. They need to ensure that the endpoint automatically scales based on traffic and that the model version is updated without downtime when a new version is registered. Which combination of features should they use?

A machine learning engineer is using Databricks Feature Store to create a feature table. They want to ensure that the feature table can be used for both model training and online serving. Which of the following is a required step when creating a feature table in Databricks Feature Store?

A data scientist is training a model using Databricks AutoML. They notice that the generated notebooks include a step that performs cross-validation using a specific metric. They want to change the primary metric used for model selection to area under the ROC curve (AUC). Where should they specify this in the AutoML experiment configuration?

A data scientist is using Hyperopt with SparkTrials on Databricks to tune a model. They notice that some trials fail due to out-of-memory errors on the driver. Which adjustment should they make to improve the tuning process?

A machine learning engineer is using MLflow Tracking on Databricks to log experiments. They want to ensure that the experiments are reproducible and that all necessary information is captured. Which TWO of the following should they log to achieve this? (Choose two.)

A machine learning engineer is using Databricks AutoML to train a model on a dataset with a binary target. They want to ensure the model is optimized for recall. Which AutoML configuration should they use?

A machine learning engineer is deploying a model to a Databricks model serving endpoint. They need to enable auto-scaling to handle variable traffic. Which configuration parameter should they set when creating the endpoint?

A data scientist is building a Feature Store in Databricks. They need to ensure that features are available for both model training and online serving. They have created a feature table and want to publish it to an online store. Which TWO actions must they perform to make the features available for low-latency online inference? (Choose two.)

A machine learning engineer is deploying an MLflow model to a Databricks model serving endpoint. The model was trained and logged with a custom Python function registered in the model's signature. When the endpoint is invoked, the request fails with an error indicating the custom function is not found. The engineer confirms the model artifact and its dependencies are correctly logged in the MLflow run. Which step should the engineer take to resolve this issue?

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Databricks Machine Learning sessions

Start a Databricks Machine Learning only practice session

Every question in these sessions is drawn from the Databricks Machine Learning domain — nothing else.

Related practice questions

Related Databricks-ML-Assoc topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-ML-Assoc exam test about Databricks Machine Learning?
Be able to trace a deployment or inference failure from the exhibit to its cause using MLflow and Model Serving concepts. The most important thing: read the exact error or log message and match it to the correct Databricks component and fix.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Databricks Machine Learning questions in a focused session?
Yes — the session launcher on this page draws every question from the Databricks Machine Learning domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-ML-Assoc topics?
Use the topic links above to move to related areas, or go back to the Databricks-ML-Assoc question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-ML-Assoc exam covers. They are not copied from any real exam or dump site.