Courseiva
Model Deployment →mediumMultiple Choice

Databricks-ML-Assoc Model Deployment Practice Question

A data scientist trains a model with a feature engineering pipeline and wants batch scoring to happen nightly on a Delta table using Databricks, producing predictions that downstream dashboards read. The scoring job must scale with data volume and be re-runnable if it fails. Which approach best fits these requirements?

⚠ Common exam trap

The trap here is assuming a real-time serving endpoint is the default way to use a model, when scheduled distributed batch scoring is the appropriate pattern for bulk table data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use MLflow's predict API to load the model in a Spark job and score the Delta table with a distributed DataFrame.

Nightly bulk scoring on a Delta table is a distributed batch workload, so the model should be loaded with the MLflow predict API inside a Spark job that reads the feature table and writes predictions back to Delta. That approach scales horizontally with data volume and integrates cleanly with Databricks job scheduling and retries, unlike endpoint-per-row calls, dashboard-embedded pickles or single-threaded SQL.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Deploy the model to a real-time Model Serving endpoint and call it row by row from a notebook loop.

    Why it's wrong here

    A real-time endpoint is designed for low-latency interactive scoring, and invoking it once per row incurs per-request overhead that scales terribly with table size. It also couples nightly batch throughput to endpoint capacity and cost, and a mid-job failure leaves partial results with no natural checkpoint. This is an anti-pattern when the workload is inherently bulk and scheduled.

  • ✗

    Export the model to a pickle file and have the dashboard application load and score records on demand.

    Why it's wrong here

    Embedding a pickle in a dashboard application pushes model serving into the presentation tier, where scaling, dependency management and versioning are all handled poorly. It also makes the nightly schedule meaningless because scoring happens whenever a user opens the dashboard. This design centralizes risk in the wrong layer and does not provide the distributed, re-runnable batch job the scenario requires.

  • ✓

    Use MLflow's predict API to load the model in a Spark job and score the Delta table with a distributed DataFrame.

    Why this is correct

    Loading the model with the MLflow predict API inside a Spark job lets the scoring run distributed across the cluster, so throughput scales with the data volume and executor count. The job can be scheduled with a Databricks job and retried on failure, and results are written back to Delta for dashboards. This aligns with a nightly, re-runnable bulk scoring pattern.

  • ✗

    Convert the model to a SQL UDF and run a single-threaded query against the feature table.

    Why it's wrong here

    A SQL UDF can score data efficiently, but a single-threaded query defeats the requirement to scale with data volume and would be far slower than a distributed job on large tables. It also complicates re-runnability because orchestration and retries would have to be built around ad hoc SQL. The scenario calls for a scheduled, distributed batch process rather than an interactive query.

About these practice questions

One of 319 original Databricks-ML-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.