Courseiva
ML Workflows →mediumMultiple Choice

Databricks-ML-Assoc ML Workflows Practice Question

You are training a scikit-learn model with MLflow in a Databricks notebook. The model's preprocessing includes a custom Python function that you wrote in the notebook. You need to register the model to the Databricks Model Registry and later deploy it with Model Serving, ensuring the preprocessing is applied automatically at inference. Which approach should you use?

⚠ Common exam trap

The trap here is assuming that logging a function as an artifact or passing it to the signature parameter makes MLflow execute it during inference, when only code packaged inside the model artifact is actually run.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Wrap the preprocessing and model in an mlflow.pyfunc.PythonModel subclass, and log it with mlflow.pyfunc.log_model.

To ensure custom preprocessing runs at inference time, the logic must be packaged inside the logged model artifact. Subclassing mlflow.pyfunc.PythonModel and logging with mlflow.pyfunc.log_model bundles the preprocessing code with the trained model, so Model Serving executes it automatically. Other approaches either misuse parameters, store code without wiring it into scoring, or assume unsupported endpoint hooks.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Register the model with the sklearn flavor, and configure a pre-processing hook in the Model Serving endpoint configuration.

    Why it's wrong here

    Databricks Model Serving does not expose a pre-processing hook in the endpoint configuration. Any transformation logic must be part of the logged model artifact itself, such as a pyfunc wrapper or an sklearn Pipeline. Relying on an endpoint-level hook is not a supported feature and would leave the custom function unexecuted during inference.

  • ✗

    Log the preprocessing function as a separate artifact using mlflow.log_artifact, and reference it from the model's conda environment.

    Why it's wrong here

    log_artifact stores arbitrary files in the run's artifact location, and the conda environment file only declares package dependencies. Neither mechanism causes MLflow or Model Serving to execute the preprocessing function when scoring. The function would not be loaded or invoked, so predictions would be made on raw, untransformed input and likely produce incorrect results.

  • ✗

    Log the model with mlflow.sklearn.log_model, passing the custom preprocessing function in the signature parameter.

    Why it's wrong here

    The signature parameter of log_model expects a ModelSignature object describing input and output schema, not a callable. Passing a function there will raise an error or be ignored, and the preprocessing will not be captured, so Model Serving will not apply it. The custom logic must instead be packaged with the model artifact, for example via a custom pyfunc or a pipeline that includes the transformation.

  • ✓

    Wrap the preprocessing and model in an mlflow.pyfunc.PythonModel subclass, and log it with mlflow.pyfunc.log_model.

    Why this is correct

    A custom PythonModel encapsulates the preprocessing code and the trained model in a single artifact. When logged with mlflow.pyfunc.log_model, MLflow serializes the class and its dependencies into the model directory, so Model Serving loads and executes the preprocessing automatically at inference. This is the supported pattern for custom inference logic in Databricks.

About these practice questions

Courseiva writes every Databricks-ML-Assoc question from scratch — 319 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.