Courseiva
Model Deployment →mediumMultiple Choice

Databricks-ML-Pro Model Deployment Practice Question

An ML engineer is deploying a scikit-learn model to a Databricks Model Serving endpoint. The model expects a single feature vector of 10 float values per request. The endpoint must return predictions in under 100 ms. Which approach should the engineer use to minimize per-request overhead?

⚠ Common exam trap

The trap here is assuming that autoscaling or request batching reduces the latency of a single small request, when they primarily improve throughput under load.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Log the model using the native scikit-learn flavor and deploy it with a workload size that provides sufficient CPU.

For low-latency single-vector inference, minimizing framework overhead is critical. The native scikit-learn flavor avoids the extra Python layer that pyfunc introduces, and a suitable workload size provides the CPU needed to execute the model quickly. Autoscaling and request batching address throughput rather than per-request latency, so they do not help meet the strict latency target.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Log the model with a signature that specifies a tensor input and enable the endpoint's request batching.

    Why it's wrong here

    Request batching can improve throughput by grouping multiple requests, but it adds latency because the server waits to accumulate a batch. For a strict sub-100 ms requirement with single-vector requests, batching is counterproductive. Also, specifying a tensor input does not inherently reduce per-request overhead; it changes the input schema but not the fundamental cost of invoking the model. This approach targets throughput, not minimal per-request latency.

  • ✗

    Enable autoscaling on the endpoint and send requests with a batch size of 1.

    Why it's wrong here

    Autoscaling adjusts the number of model server instances based on load, but it does not reduce the per-request overhead of processing a single input. With batch size 1, the endpoint still incurs the same preprocessing and framework invocation costs per request. Autoscaling helps with throughput under high concurrency, not with the latency of an individual small request. For low-latency single-vector inference, the overhead is dominated by framework call overhead, which autoscaling cannot eliminate.

  • ✗

    Use MLflow's pyfunc flavor with a custom predict method that processes one row at a time.

    Why it's wrong here

    The pyfunc flavor adds a Python-level wrapper around the model, which introduces additional overhead compared to the native scikit-learn flavor. Processing one row at a time further prevents any vectorized optimizations. While pyfunc is flexible, it is not the lowest-latency option for a standard scikit-learn model. For minimal per-request overhead, the native flavor that integrates directly with the model server is preferred.

  • ✓

    Log the model using the native scikit-learn flavor and deploy it with a workload size that provides sufficient CPU.

    Why this is correct

    The native scikit-learn flavor allows the model server to load and invoke the model directly without an extra Python wrapper, reducing per-request overhead. Choosing an appropriate workload size ensures enough CPU resources to handle the model's computation quickly. This combination minimizes framework overhead and provides the necessary compute, making it the best choice to meet the sub-100 ms latency requirement for single-vector requests.

About these practice questions

One of 300 original Databricks-ML-Pro practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.