Courseiva
Model Development →mediumMultiple Choice

Databricks-ML-Assoc Model Development Practice Question

A data scientist is training a deep learning model on Databricks using Horovod for distributed training. They find that the model is converging slowly. What is the most likely cause related to the distributed configuration?

⚠ Common exam trap

Candidates often assume that distributed training automatically scales the learning rate. They fail to realize that increasing the global batch size requires a corresponding increase in the learning rate to maintain stable convergence.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The learning rate was not adjusted for the distributed batch size.

In Horovod distributed training, the learning rate must often be scaled relative to the number of workers, because each worker processes a batch of data. A common mistake is failing to adjust the learning rate or optimizer settings to compensate for the aggregate batch size across the cluster. This results in ineffective gradient updates, slowing down convergence significantly compared to a single-node setup.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The cluster has too much memory per worker.

    Why it's wrong here

    Excessive memory per worker is generally beneficial for training and would not cause slow convergence. Slow convergence in distributed training is almost always related to optimization parameters like learning rate, batch size, or synchronization frequency, rather than the amount of RAM available on the individual worker nodes in the cluster.

  • ✓

    The learning rate was not adjusted for the distributed batch size.

    Why this is correct

    Distributed training increases the effective batch size by multiplying the per-worker batch size by the number of workers. If the learning rate is not adjusted to account for this increase, the model will likely exhibit unstable or slow convergence. This is a fundamental concept in large-scale distributed deep learning optimization.

  • ✗

    The storage throughput of the DBFS mount is too high.

    Why it's wrong here

    High storage throughput is a performance advantage, not a detriment. Storage issues typically manifest as I/O bottlenecks where workers wait for data, but they do not cause the model to converge slowly in terms of loss reduction per epoch. Slow convergence is an algorithmic issue, not a data retrieval issue.

  • ✗

    The number of epochs is too low for the current model size.

    Why it's wrong here

    While increasing epochs can help, it does not explain why the model converges *slowly*. If a model converges correctly on one node but slowly on many, the cause is distributed synchronization and batching, not the total number of epochs. Adjusting the learning rate is the standard solution for this behavior.

About these practice questions

One of 319 original Databricks-ML-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.