Courseiva
Model Deployment →hardMultiple Choice

Databricks-ML-Pro Model Deployment Practice Question

An ML engineer is deploying a model to Databricks Model Serving that uses a custom Python function as a pre-processing step. The function relies on a global variable defined in a separate module. After deployment, the endpoint returns errors indicating the global variable is not defined. The engineer confirmed the module is included in the model's conda environment. What is the most likely cause?

⚠ Common exam trap

The trap here is assuming that any module in the conda environment will have its global state preserved, when in fact MLflow serializes only the model object and its direct code dependencies.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The global variable is not serialized with the model because MLflow only saves the model's predict method and its direct dependencies.

The error arises because MLflow's serialization does not capture global variables from external modules unless they are explicitly included in the model's code. When the model is loaded for serving, the pre-processing function may reference a global variable that was not saved, resulting in an undefined variable error. To fix this, the engineer should ensure all necessary state is encapsulated within the model or use `code_path` to include the module and initialize variables properly.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Databricks Model Serving runs each inference in a separate process, so global variables are not shared across requests.

    Why it's wrong here

    While it's true that serverless serving may isolate requests, the error occurs on the first request, indicating a serialization issue rather than cross-request state. Global variables defined at module load time would be initialized in each process if the module is loaded, but if the variable is not defined in the serialized model, it won't exist. The problem is not about sharing across requests but about the variable not being available at all.

  • ✗

    The model was logged with `mlflow.sklearn.log_model` instead of `mlflow.pyfunc.log_model`, so custom pre-processing code is ignored.

    Why it's wrong here

    Using `mlflow.sklearn.log_model` does not ignore custom pre-processing if the model is a scikit-learn pipeline that includes the pre-processing step. However, if the pre-processing is a separate custom function not part of the pipeline, it might not be captured. But the error specifically mentions a global variable not defined, which points to serialization of state, not the logging function. `mlflow.pyfunc.log_model` would allow more control but is not the root cause here.

  • ✓

    The global variable is not serialized with the model because MLflow only saves the model's predict method and its direct dependencies.

    Why this is correct

    MLflow's default model saving mechanism serializes the model object and its immediate dependencies, but it does not automatically capture global variables or module-level state from custom modules unless they are explicitly referenced within the model's class or function. If the pre-processing function relies on a global variable from another module, that state may not be preserved during serialization, leading to a NameError or undefined variable at serving time.

  • ✗

    The conda environment does not include the module because MLflow only captures packages installed via pip, not local modules.

    Why it's wrong here

    MLflow can capture local modules if they are part of the model's code path and properly referenced. However, the error is about an undefined variable, not a missing module. If the module were missing, the error would be a ModuleNotFoundError. The conda environment typically includes local modules when using `mlflow.pyfunc.log_model` with the `code_path` parameter, but even without it, the issue here is about variable state, not module availability.

About these practice questions

This Databricks-ML-Pro question is part of Courseiva's 300-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.