Courseiva
Model Development →hardMultiple Choice

Databricks-ML-Assoc Model Development Practice Question

A data scientist is training a gradient boosting model using Spark MLlib on a large dataset in Databricks. They notice that the model's performance on a validation set is significantly worse than on the training set, and they suspect overfitting. They want to use MLflow to track hyperparameters and metrics to diagnose the issue. Which combination of MLflow logging practices will best help them identify overfitting across multiple runs?

⚠ Common exam trap

The trap here is assuming that MLflow autologging automatically captures validation metrics, when in fact validation metrics must be explicitly computed and logged by the user.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Log training and validation metrics separately for each run, and log hyperparameters such as maxDepth and minInstancesPerNode.

To diagnose overfitting, it is essential to compare training and validation performance across runs. Logging both metrics separately, along with relevant hyperparameters that control model complexity, allows the data scientist to visualize the gap and tune accordingly. MLflow's tracking and comparison features then make it straightforward to identify runs where validation performance lags.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Log only the training metric and the model artifact, and rely on the MLflow UI to automatically compute validation metrics.

    Why it's wrong here

    MLflow does not automatically compute validation metrics from a logged model artifact. Without explicitly logging validation metrics, there is no way to compare training and validation performance. Relying on automatic computation is incorrect and would leave the data scientist unable to diagnose overfitting, as only training performance would be visible.

  • ✗

    Log hyperparameters and the final model, but log metrics only at the end of training to reduce clutter.

    Why it's wrong here

    Logging metrics only at the end still provides a single point of comparison, but without separate training and validation metrics, overfitting cannot be detected. The key is to log both types of metrics. Additionally, logging metrics only at the end misses the opportunity to track learning curves, which can reveal overfitting earlier. This approach is insufficient for diagnosis.

  • ✓

    Log training and validation metrics separately for each run, and log hyperparameters such as maxDepth and minInstancesPerNode.

    Why this is correct

    Logging both training and validation metrics allows direct comparison to detect overfitting, where training performance improves while validation performance degrades. Logging key hyperparameters like maxDepth and minInstancesPerNode enables correlation of model complexity with overfitting. This combination provides the necessary data to diagnose and tune the model effectively using MLflow's comparison features.

  • ✗

    Use MLflow autologging for Spark MLlib, which automatically logs training and validation metrics and hyperparameters.

    Why it's wrong here

    MLflow autologging for Spark MLlib logs parameters and metrics, but it does not automatically log validation metrics unless they are explicitly computed and logged. Autologging captures training metrics from the evaluator, but validation metrics require a separate evaluator on validation data. Therefore, relying solely on autologging would not provide the validation metrics needed to diagnose overfitting.

About these practice questions

This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.