Courseiva
Model Development →hardMultiple Choice

Databricks-ML-Assoc Model Development Practice Question

When logging a PyTorch model to the MLflow Model Registry, which component must be explicitly defined to allow the model to be loaded in an environment where the original code structure might not exist?

⚠ Common exam trap

Test-takers mistakenly think deep learning model architectures are fully self-contained, overlooking the need to explicitly define pip requirements or conda environments for portable PyTorch inference.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The environment dependencies, often captured via a requirements.txt or conda.yaml file.

When saving a model, especially deep learning models like PyTorch, the 'conda_env' or 'pip_requirements' must be explicitly captured. This ensures that the environment (dependencies and versions) is portable. Without defining these, the model might fail to load in other environments due to version mismatches or missing packages, making the model unusable for inference in production or shared staging environments.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The hardware specification of the training cluster.

    Why it's wrong here

    Hardware specifications like CPU vs GPU are environment-dependent but are not part of the MLflow model logging process. While inference performance may depend on hardware, the model artifact itself must be platform-agnostic to be portable, and MLflow handles the environment dependencies, not the specific underlying infrastructure hardware configurations.

  • ✓

    The environment dependencies, often captured via a requirements.txt or conda.yaml file.

    Why this is correct

    Capturing dependencies is vital for reproducibility. MLflow serializes the current Python environment configuration, allowing the target environment to recreate the exact package versions used during training. This prevents 'it works on my machine' scenarios by guaranteeing that the inference service has the exact library versions required by the model.

  • ✗

    A list of all users who have access to the model.

    Why it's wrong here

    Access control is handled by Databricks Unity Catalog permissions, not by the model logging process itself. Including user information in the model artifact is unnecessary and potentially a security risk. MLflow focuses on artifact integrity, reproducibility, and metadata, keeping it separate from the platform's authentication and authorization mechanisms.

  • ✗

    The full raw dataset used for training.

    Why it's wrong here

    Saving the entire training dataset within the model artifact is inefficient and generally not done. Instead, models should log the data URI or lineage metadata. Storing raw data in the artifact storage leads to massive overhead, slow logging times, and unnecessary duplication of data across the environment.

About these practice questions

This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.