Courseiva
ML Ops →easyMultiple Choice

Databricks-ML-Pro ML Ops Practice Question

A machine learning engineer is setting up a Databricks job to retrain a model every night. The job must use a specific Python library version that is not available in the default Databricks Runtime. The engineer wants to ensure that the retraining job has access to this library without affecting other jobs in the workspace. What is the recommended approach?

⚠ Common exam trap

The trap here is thinking that installing a library in a notebook with `%pip install` is sufficient for a scheduled job, but that installation is ephemeral and not isolated, so it fails to provide a persistent, job-specific environment.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Install the library on the job cluster by adding it as a cluster library, and configure the job to use that cluster.

For a Databricks job that requires a specific library, the best practice is to install the library on the job cluster. Job clusters are dedicated to the job and can be configured with libraries that are installed at cluster startup. This ensures the library is available for every run, isolates dependencies, and avoids impacting other workloads. It is simpler and more maintainable than custom Docker images or notebook-level installations.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Modify the global init script to install the library on all clusters in the workspace.

    Why it's wrong here

    Global init scripts run on all clusters and would install the library workspace-wide, affecting other jobs and potentially causing version conflicts. This violates the requirement to avoid affecting other jobs. Additionally, global init scripts require cluster restart to take effect and are not scoped to a specific job, making them unsuitable for this scenario.

  • ✓

    Install the library on the job cluster by adding it as a cluster library, and configure the job to use that cluster.

    Why this is correct

    In Databricks, you can install libraries at the cluster level, making them available to all notebooks and jobs that run on that cluster. For a job, you can define a job cluster with the required library installed. This ensures the library is present for every run without manual installation, and it isolates the library to that cluster, preventing impact on other jobs. This is the recommended and simplest approach for job-specific dependencies.

  • ✗

    Create a custom Docker image with the required library and configure the job cluster to use that image.

    Why it's wrong here

    While custom Docker images are supported in Databricks, they are typically used for specialized environments and require additional setup with a container registry. This approach is more complex than necessary for a single library and may not be the recommended first choice for a simple dependency. Databricks provides a more straightforward mechanism for library installation at the cluster level.

  • ✗

    Install the library using `%pip install` in the first cell of the notebook used by the job.

    Why it's wrong here

    Using `%pip install` in a notebook installs the library only for that notebook session, but it is not persistent across job runs unless the notebook is re-executed each time. For a scheduled job, this would require installing the library on every run, which adds overhead and may fail if the library is not available from the package repository. It also does not isolate the library from other jobs, as the installation is ephemeral.

About these practice questions

One of 300 original Databricks-ML-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.