Databricks-ML-Assoc Model Development Practice Question
A data scientist is training a gradient boosting model using Spark MLlib on a large dataset in Databricks. They notice that the model's performance on a validation set is significantly worse than on the training set, and they suspect overfitting. They want to use MLflow to track hyperparameters and metrics to diagnose the issue. Which combination of MLflow logging practices will best help them identify overfitting across multiple runs?
⚠ Common exam trap
The trap here is assuming that MLflow autologging automatically captures validation metrics, when in fact validation metrics must be explicitly computed and logged by the user.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Log training and validation metrics separately for each run, and log hyperparameters such as maxDepth and minInstancesPerNode.
To diagnose overfitting, it is essential to compare training and validation performance across runs. Logging both metrics separately, along with relevant hyperparameters that control model complexity, allows the data scientist to visualize the gap and tune accordingly. MLflow's tracking and comparison features then make it straightforward to identify runs where validation performance lags.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Log only the training metric and the model artifact, and rely on the MLflow UI to automatically compute validation metrics.
Why it's wrong here
MLflow does not automatically compute validation metrics from a logged model artifact. Without explicitly logging validation metrics, there is no way to compare training and validation performance. Relying on automatic computation is incorrect and would leave the data scientist unable to diagnose overfitting, as only training performance would be visible.
- ✗
Log hyperparameters and the final model, but log metrics only at the end of training to reduce clutter.
Why it's wrong here
Logging metrics only at the end still provides a single point of comparison, but without separate training and validation metrics, overfitting cannot be detected. The key is to log both types of metrics. Additionally, logging metrics only at the end misses the opportunity to track learning curves, which can reveal overfitting earlier. This approach is insufficient for diagnosis.
- ✓
Log training and validation metrics separately for each run, and log hyperparameters such as maxDepth and minInstancesPerNode.
Why this is correct
Logging both training and validation metrics allows direct comparison to detect overfitting, where training performance improves while validation performance degrades. Logging key hyperparameters like maxDepth and minInstancesPerNode enables correlation of model complexity with overfitting. This combination provides the necessary data to diagnose and tune the model effectively using MLflow's comparison features.
- ✗
Use MLflow autologging for Spark MLlib, which automatically logs training and validation metrics and hyperparameters.
Why it's wrong here
MLflow autologging for Spark MLlib logs parameters and metrics, but it does not automatically log validation metrics unless they are explicitly computed and logged. Autologging captures training metrics from the evaluator, but validation metrics require a separate evaluator on validation data. Therefore, relying solely on autologging would not provide the validation metrics needed to diagnose overfitting.
About these practice questions
This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.