Courseiva
ML Workflows →hardMultiple Choice

Databricks-ML-Assoc ML Workflows Practice Question

You are deploying a model using MLflow, and you want to log the custom pre-processing logic alongside the model so it is automatically applied during inference. How should you achieve this?

⚠ Common exam trap

Test-takers often select generic artifact logging functions instead of the specialized pyfunc wrapper required to bundle custom pre-processing logic with the model.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the mlflow.pyfunc.log_model() function with a custom code wrapper.

MLflow allows you to log custom Python functions or entire pipelines as part of the model using the 'pyfunc' flavor. By wrapping your pre-processing steps and model into a custom pyfunc class and logging it, you encapsulate both the transformation and the prediction logic. This is a critical best practice for production deployments, as it ensures consistent pre-processing between training and inference, eliminating training-serving skew and simplifying the deployment process for the serving endpoint.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Create a separate inference job that runs the pre-processing code before the model.

    Why it's wrong here

    Splitting logic between an inference job and a model introduces complexity and potential drift. If the pre-processing code is updated, the inference job must also be updated, leading to synchronization issues. Encapsulating the logic within the model artifact itself is a safer and more robust approach to production deployment.

  • ✓

    Use the mlflow.pyfunc.log_model() function with a custom code wrapper.

    Why this is correct

    The pyfunc flavor is the standard way to package custom logic with models in MLflow. It allows users to define a wrapper that executes arbitrary code, including pre-processing, at inference time. This ensures that the model is self-contained, simplifying deployment and guaranteeing that input data is processed correctly.

  • ✗

    Store the pre-processing logic in a shared Delta table and call it at inference.

    Why it's wrong here

    Storing code in a database table is not a standard practice and is generally discouraged. It makes version control and testing difficult and creates a dependency that could fail at runtime. Code should be stored in version-controlled repositories and packaged with the model, not stored as data.

  • ✗

    Log the pre-processing parameters as tags in the MLflow run and re-apply them.

    Why it's wrong here

    Tags are intended for metadata and search, not for storing complex logic or transformation code. Attempting to store logic as tags is ineffective, as the inference runtime would still need the code to interpret those tags. This approach is not scalable and violates the principles of clean model packaging.

About these practice questions

One of 319 original Databricks-ML-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.