Courseiva

Databricks-ML-Pro · topic practice

Model Development practice questions

This domain covers building and training models on Databricks: experiment tracking with MLflow, distributed training, hyperparameter tuning, and lifecycle management via Model Registry. Questions test whether you can choose the right Databricks tool for cross-validation, model versioning, drift monitoring, and artifact storage, and reason about how these integrate across the workspace.

Courseiva uses original exam-style practice questions designed for learning and revision. The goal is to understand the concepts, recognise exam patterns, and improve through explanations — not memorise copied exam dumps.

Editorial oversight:Johnson Ajibi· MSc IT Security, IEEE Senior Member
20 questionsDomain: Model Development

What the exam tests

What to know about Model Development

You must be able to log experiments with MLflow, run scalable cross-validation on Spark, register and transition model versions, and set up drift monitoring. The single most important thing: know where MLflow artifacts are stored and how the Model Registry tracks lifecycle stages.

Using MLflow Tracking to log parameters, metrics, and artifacts during model training runs.

Performing cross-validation on large data with Spark ML or spark-sklearn wrappers.

Managing model versions and stage transitions with the MLflow Model Registry.

Monitoring production feature drift using Databricks Lakehouse Monitoring or model serving metrics.

Watch out for

Common Model Development exam traps

  • ▸Assuming MLflow artifacts are stored only in the workspace filesystem; they actually go to the configured artifact store (DBFS, S3, ADLS).
  • ▸Confusing Model Registry stage transitions with deployment; registering a model does not automatically serve it.
  • ▸Using default scikit-learn cross-validation on Spark DataFrames, which collects data to the driver and fails on large datasets.

Practice set

Model Development questions

20 questions · select your answer, then reveal the explanation

A data scientist is training a Random Forest model on Databricks. They want to ensure that the hyperparameter tuning process is efficient while preventing overfitting during the training phase. Which approach best achieves this goal using MLflow and Hyperopt?

A machine learning engineer needs to handle imbalanced datasets for a fraud detection model using PySpark and MLlib. Which TWO techniques are valid for mitigating this issue within the Databricks environment?

Refer to the exhibit. A data scientist is deploying a model to the Model Registry. What is the impact of the registered_model_name parameter in the log_model function?

Exhibit

import mlflow.sklearn
with mlflow.start_run():
    model = RandomForestClassifier()
    model.fit(X_train, y_train)
    mlflow.sklearn.log_model(model, 'my_model', registered_model_name='ProductionModel')

A data scientist is training a deep learning model using TensorFlow on Databricks. They observe that training speed is slow despite using a GPU cluster. Which action should they prioritize to optimize performance?

A machine learning engineer needs to deploy a custom model that uses a non-standard pre-processing library. How should they package the model to ensure the environment is correctly replicated during inference?

When sharing a Databricks notebook that contains sensitive model training logic, which feature should be used to provide users access without exposing the underlying source code?

A data scientist is training a PyTorch model on Databricks using MLflow. They need to ensure that the model architecture and training parameters are automatically logged without explicitly calling mlflow.pytorch.autolog() in every notebook. Which approach is the most efficient?

Which TWO of the following steps are required to properly register a custom Python model in the Unity Catalog Model Registry using the MLflow Fluent API?

When training a model using Scikit-Learn in a Databricks notebook, which MLflow component should a data scientist use to group multiple related training runs under a single experiment?

Which THREE of the following are valid ways to improve the performance of a model during the development phase in Databricks?

Question 11mediummultiple choice
Read the full Model Development explanation →

Which of the following describes the correct relationship between an MLflow Run and a Model Version?

Question 12mediummultiple choice
Read the full Model Development explanation →

A data scientist is training a Random Forest model on a 500GB dataset in Databricks. They notice extreme memory pressure on the driver node during the final aggregation phase. Which approach best resolves this while maintaining model performance?

Which TWO of the following configurations are required to effectively track and reproduce a deep learning experiment using MLflow in Databricks?

Which THREE actions are required when migrating a custom Scikit-learn model to the Databricks Model Registry to ensure it can be served via MLflow Model Serving?

Refer to the exhibit. A data engineer attempts to transition a model version to 'Staging' but receives this error. What is the most likely cause?

Exhibit

MLflow Traceback:
mlflow.exceptions.RestException: RESOURCE_DOES_NOT_EXIST: Model version with name 'DemandForecast' and version '5' not found.

A data scientist is using Databricks to train a model with a very large vocabulary, leading to high-dimensional sparse inputs. Which storage format for the feature table is most efficient for both storage and retrieval?

Which TWO of the following are valid ways to register a model in the Databricks Model Registry?

Question 18mediummultiple choice
Read the full Model Development explanation →

When designing a custom MLflow model flavor for a proprietary framework, what is the most critical method to implement to ensure compatibility with `mlflow.pyfunc`?

Refer to the exhibit. You are attempting to log a model to the MLflow Model Registry, but receive the serialization error shown. What is the cause of this error?

Exhibit

MLflow error: 'Could not serialize model: Model contains references to local file system path /dbfs/mnt/data/model.pkl'
Question 20mediummultiple choice
Read the full Model Development explanation →

Refer to the exhibit. You have multiple experimental runs in your MLflow tracking server. You want to retrieve the best-performing model based on accuracy. Which command should you use?

Exhibit

MLflow run: training_run_123, metrics: {accuracy: 0.85}, params: {n_estimators: 100}

Free account

Track your progress over time

Create a free account to save your results and see which topics improve across sessions.

Focused Model Development sessions

Start a Model Development only practice session

Every question in these sessions is drawn from the Model Development domain — nothing else.

Related practice questions

Related Databricks-ML-Pro topic practice pages

Move into related areas when this topic feels solid.

Frequently asked questions

What does the Databricks-ML-Pro exam test about Model Development?
You must be able to log experiments with MLflow, run scalable cross-validation on Spark, register and transition model versions, and set up drift monitoring. The single most important thing: know where MLflow artifacts are stored and how the Model Registry tracks lifecycle stages.
How should I use these practice questions?
Select your answer before revealing the explanation. Then read why each option is right or wrong — this active recall approach builds retention far faster than re-reading notes.
Can I practise just Model Development questions in a focused session?
Yes — the session launcher on this page draws every question from the Model Development domain. Use a 10-question session first to gauge your baseline, then move to 20 or 30 once the weak spots are clear.
Where can I practise other Databricks-ML-Pro topics?
Use the topic links above to move to related areas, or go back to the Databricks-ML-Pro question bank to see all topics.
Are these real exam questions or dumps?
These are original practice questions written to test the same concepts the Databricks-ML-Pro exam covers. They are not copied from any real exam or dump site.