Databricks-ML-Assoc ML Workflows Practice Question
An ML engineer trains a scikit-learn model on a Spark DataFrame in a Databricks notebook using MLflow autologging. The run logs parameters and metrics, but the engineer later opens the MLflow run and cannot find any input dataset lineage. Which action should the engineer take to ensure the training dataset is recorded with the run in the MLflow UI?
⚠ Common exam trap
The trap here is assuming MLflow autologging automatically captures input dataset lineage, when in fact dataset tracking must be added explicitly with mlflow.log_input().
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Wrap the training data with mlflow.data.from_spark() and pass it to mlflow.log_input().
MLflow records input dataset lineage only when a Dataset object is created and logged to the active run. The mlflow.data.from_spark() constructor builds that object from a Spark DataFrame, and mlflow.log_input() attaches it, after which the run UI shows dataset source, digest, and schema. Autologging captures params, metrics, and the model, but not input data lineage, so an explicit call is required.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Set the MLFLOW_TRACKING_URI environment variable to the Unity Catalog metastore and rerun training.
Why it's wrong here
MLFLOW_TRACKING_URI controls where runs are stored, not what metadata is captured about a run. Pointing tracking at a different backend will not add dataset lineage to the existing run or to future runs. Dataset logging must be invoked explicitly in code; changing the tracking URI is orthogonal to the missing lineage problem.
- ✗
Call mlflow.log_artifact() on the Spark DataFrame object directly after training.
Why it's wrong here
log_artifact expects a local file or directory path, not a DataFrame object, so passing a Spark DataFrame will fail or produce an unusable artifact. Even if the data were written to disk first, that artifact would not populate the Dataset lineage panel of the MLflow run; dataset tracking requires the dedicated dataset logging API, not a generic artifact upload.
- ✗
Enable Delta table time travel on the source table and reference the table path in the run tags.
Why it's wrong here
Time travel and tags do not create MLflow dataset lineage entries. Manually tagging a run with a table path is free-text metadata and will not render in the Datasets panel or support dataset comparison across runs. While Delta time travel is useful for reproducibility, it is not the API that records an input dataset on an MLflow run.
- ✓
Wrap the training data with mlflow.data.from_spark() and pass it to mlflow.log_input().
Why this is correct
mlflow.data.from_spark() constructs a Dataset object from a Spark DataFrame, and mlflow.log_input() records it on the active run, which is exactly what populates the Datasets section of the MLflow run UI. This is the supported mechanism for capturing input data lineage in MLflow on Databricks, and it works alongside autologging rather than replacing it.
About these practice questions
This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.