Courseiva
Model Development →mediumMultiple Choice

Databricks-ML-Assoc Model Development Practice Question

A data scientist is training a scikit-learn model on a Databricks cluster with autoscaling enabled. They observe that each epoch takes significantly longer than expected, and the Spark UI shows very low CPU utilization across the worker nodes. The dataset is small enough to fit in memory on a single node. What is the most likely cause of the slow training?

⚠ Common exam trap

The trap here is assuming that any Databricks cluster automatically distributes scikit-learn training across workers, when in fact standard scikit-learn runs only on the driver.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The model is being trained on the driver node only, and the worker nodes are idle because the training code does not distribute the workload.

The correct answer is that the model is being trained on the driver node only, leaving worker nodes idle. In Databricks, standard scikit-learn training runs on the driver unless you explicitly distribute it. The low CPU utilization on workers and slow epochs confirm that the workload is not parallelized. To speed up training, you could use Spark ML, Horovod, or parallelize hyperparameter tuning with Hyperopt.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    The cluster is using a runtime version that does not support scikit-learn, forcing a fallback to single-threaded execution.

    Why it's wrong here

    Databricks runtimes include scikit-learn and support it fully. A runtime version mismatch would typically cause import errors or version conflicts, not silent single-threaded execution. The symptom of low CPU utilization across workers points to the training not being distributed, not a missing library.

  • ✗

    The cluster's autoscaling policy is aggressively terminating worker nodes, causing the training to restart.

    Why it's wrong here

    Autoscaling does not restart running tasks; it only adjusts node count based on load. Since the dataset fits in memory and CPU utilization is low, autoscaling is not the bottleneck. The slow training is due to inefficient parallelism from the driver-centric training, not node termination.

  • ✓

    The model is being trained on the driver node only, and the worker nodes are idle because the training code does not distribute the workload.

    Why this is correct

    When you train a scikit-learn model using standard APIs like .fit(), the computation runs on the driver node. Worker nodes remain idle, resulting in low cluster CPU utilization and slow training. To leverage the cluster, you would need to use distributed training libraries such as Spark ML or Horovod, or parallelize hyperparameter tuning.

  • ✗

    The data is stored in DBFS, and reading it for each epoch introduces network latency that throttles training.

    Why it's wrong here

    If the dataset fits in memory, it is typically loaded once and cached. DBFS read latency would affect initial load time, not per-epoch training time. Low CPU utilization across workers suggests the workers are not participating in computation, not that I/O is slow.

About these practice questions

This Databricks-ML-Assoc question is part of Courseiva's 319-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.