Courseiva
Model Development →hardMultiple Choice

Databricks-ML-Pro Model Development Practice Question

Exhibit

import mlflow
from sklearn.ensemble import RandomForestRegressor

# Configure model
model = RandomForestRegressor()

# Training logic
with mlflow.start_run():
    mlflow.sklearn.log_model(model, "model")
    # Further code...

Refer to the exhibit. A developer wants to ensure the Random Forest model can be used for automated inference at scale. Based on the provided code, what is missing to enable the model to support the 'predict' method within the Databricks Model Serving environment?

⚠ Common exam trap

Candidates often focus on the model type (e.g., Random Forest) or the logging library, missing the fundamental requirement that a model must be trained (fitted) to be serializable.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The model.fit(X_train, y_train) method must be called before log_model.

The provided code logs an unfitted model instance. MLflow cannot serialize a model's 'predict' capabilities until the model has actually been trained (fitted) on a dataset. Attempting to deploy or run an unfitted model results in an error, as there is no learned logic to execute. The model must be fitted with data prior to being passed to log_model.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The model must be registered to the model registry before logging.

    Why it's wrong here

    Registration is a separate step that occurs after logging. While registration is required for deployment, it does not fix the issue of a non-fitted model. An unfitted model will still fail to produce predictions, regardless of whether it is registered in the registry or not.

  • ✓

    The model.fit(X_train, y_train) method must be called before log_model.

    Why this is correct

    Scikit-learn models must be trained via the .fit() method before they contain the internal state necessary for inference. Logging an unfitted instance saves an empty shell, which cannot perform predictions. The model must be trained on representative data to ensure the internal weights are populated correctly.

  • ✗

    The model must be wrapped in an mlflow.pyfunc.PythonModel class.

    Why it's wrong here

    While creating a custom pyfunc is useful for complex logic, it is not required for standard Scikit-Learn models. Scikit-learn has a native MLflow flavor that handles serialization automatically. The issue here is the lack of training, not the choice of the model wrapper or framework integration.

  • ✗

    The code must explicitly define the input signature in log_model.

    Why it's wrong here

    A signature is recommended for validation but is not strictly required to enable the 'predict' method. A model can technically function without a signature, though it is not a production best practice. The fundamental failure in this scenario is that the model has no learned logic to execute.

About these practice questions

Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.