Courseiva
Model Development →hardMultiple Choice

Databricks-ML-Assoc Model Development Practice Question

Exhibit

MLflow error: 'Cannot find module: sklearn.utils.extmath'

Refer to the exhibit. A model was successfully logged but fails to load in a production environment with the error shown in the exhibit. What is the most likely cause of this issue?

⚠ Common exam trap

Candidates often blame the model code itself, missing the most common cause: environment drift where the production library versions differ from the training environment.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The production environment has a different version of scikit-learn than the training environment.

The error indicates a missing dependency or version mismatch in the inference environment. When a model relies on specific sub-modules, these must be available in the underlying Python environment. This issue highlights the importance of correctly specifying dependencies during the logging phase. If the environment is not captured precisely, the inference cluster will fail to load the model because the necessary shared libraries are missing or incompatible.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The model artifact is corrupted during the transfer to the production environment.

    Why it's wrong here

    A corrupted artifact would typically result in a pickle error or a manifest mismatch error, not a 'Cannot find module' error. The module error specifically points to a Python environment configuration issue where the required library (scikit-learn) is either not installed or not accessible in the current path.

  • ✓

    The production environment has a different version of scikit-learn than the training environment.

    Why this is correct

    Module errors in Python usually arise when the code attempts to import a module that does not exist in the current version of the library. If the training environment had a newer scikit-learn version and the production environment is running an older one, sub-modules may not be present.

  • ✗

    The model was trained using an incompatible version of the Spark runtime.

    Why it's wrong here

    The Spark runtime version does not directly impact the ability to import standard Python libraries like scikit-learn. While Spark runtime differences can cause issues with Java/Scala dependencies, a 'Cannot find module' error is almost exclusively a Python-level environment or library path configuration problem.

  • ✗

    The model requires an external network connection to download the sklearn library during loading.

    Why it's wrong here

    MLflow model loading should be self-contained. Attempting to download libraries during model loading is not supported and would lead to timeouts or security errors, not module-specific errors. The environment must be fully provisioned before the model loading process starts, typically using the dependencies defined in the MLmodel file.

About these practice questions

This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.