Databricks-ML-Assoc Model Deployment Practice Question
A team is deploying a batch scoring pipeline that loads a registered MLflow model and runs predictions over a large Delta table using Spark. They want the scoring job to reuse the model's training-time preprocessing and to remain reproducible months later. Which two practices should they follow? (Choose two.)
⚠ Common exam trap
The trap here is treating an alias like a version pin, when aliases intentionally move and therefore cannot provide the immutability that reproducible batch scoring requires.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Reference the model by its numeric version when loading it in the batch job, so the exact artifact used in validation is always retrieved.
Reproducible batch scoring requires an immutable artifact reference and self-describing metadata. Logging the model with a signature and input example embeds schema and preprocessing, while loading a specific numeric version pins the exact artifact so reruns match validation. Moving aliases, external notebook imports, and raw pickles all introduce drift or lose the metadata needed to reproduce training-time behavior.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Reference the model by its numeric version when loading it in the batch job, so the exact artifact used in validation is always retrieved.
Why this is correct
Loading a specific numeric version pins the job to an immutable artifact, guaranteeing that reruns months later use the identical weights and preprocessing. This is the standard way to achieve reproducibility for batch scoring, and it complements the signature metadata. Relying on a moving alias would instead silently change behavior whenever the alias is reassigned, breaking reproducibility.
- ✗
Load the model using the `latest` alias so the batch job always benefits from the most recent improvements automatically.
Why it's wrong here
A `latest` alias resolves to a moving target, so a rerun of the same job at a later date could load a different model and produce different scores. That directly conflicts with the reproducibility requirement. Aliases are useful for serving endpoints that should follow promotions, but batch jobs needing stable, auditable results should pin an immutable numeric version instead.
- ✓
Log the model with `mlflow.pyfunc.log_model` (or framework flavor) including the model signature and input example so the schema and preprocessing are captured with the artifact.
Why this is correct
Logging the model with a signature and input example records the expected input schema and a sample, letting the serving or batch loader validate and coerce columns. When the pyfunc wrapper encapsulates preprocessing, batch scoring applies the same transformations as training. This is the mechanism that makes the artifact self-describing and reproducible across environments and time.
- ✗
Store the preprocessing code in a separate notebook and import it at scoring time so the batch job and training job share the same functions.
Why it's wrong here
Sharing notebook code couples the scoring job to an external, mutable source that can drift independently of the logged model. If someone edits the notebook, historical reruns change behavior even though the model artifact is unchanged. Encapsulating preprocessing inside the logged pyfunc model, rather than importing live code, is what guarantees the batch job reproduces training-time transformations.
- ✗
Convert the model to a pickle file, store it in a Delta table as a binary column, and load it via Spark at scoring time.
Why it's wrong here
Storing a pickle in a Delta table discards MLflow's flavor metadata, signature, and dependency tracking, and pickles are fragile across library versions. It also bypasses registry lineage, making audits and rollbacks harder. This approach does not capture preprocessing or schema, so it cannot ensure the batch job reproduces training-time behavior, and it introduces security risks from deserializing arbitrary objects.
About these practice questions
One of 319 original Databricks-ML-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.