Courseiva
Model Deployment →hardMultiple Choice

Databricks-ML-Pro Model Deployment Practice Question

An ML engineer is deploying a model to a Databricks Model Serving endpoint. The model's inference function logs predictions to a Delta table for monitoring. During testing, they notice that the logging adds significant latency. They need to reduce the impact on inference latency. Which approach should they take?

⚠ Common exam trap

The trap here is assuming that scaling resources or reducing logging frequency will eliminate the latency, when the real issue is synchronous blocking I/O.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Move the logging to an asynchronous background process.

Asynchronous logging moves the logging operation outside the critical path of inference, allowing predictions to be returned without waiting for the log write to complete. This is the most effective way to reduce latency caused by logging in a serving endpoint. Databricks Model Serving supports asynchronous logging patterns, such as using the inference table or custom async code.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable auto-scaling to handle the additional load from logging.

    Why it's wrong here

    Auto-scaling adjusts the number of workers based on traffic but does not change the synchronous nature of logging within each request. The latency added by logging per request remains. Auto-scaling might help with overall throughput but does not reduce the latency of individual requests. The bottleneck is the logging operation itself.

  • ✗

    Reduce the frequency of logging by sampling predictions.

    Why it's wrong here

    Sampling reduces the volume of logs but does not eliminate the latency caused by each logging call. If logging is synchronous, even occasional calls can add latency. Sampling might be acceptable for monitoring but does not solve the core issue of blocking I/O during inference. It is a trade-off between data completeness and performance.

  • ✓

    Move the logging to an asynchronous background process.

    Why this is correct

    Asynchronous logging decouples the logging operation from the inference path. The model returns predictions immediately, and logging occurs in the background. This significantly reduces latency because the response is not delayed by I/O operations. Databricks Model Serving supports asynchronous logging via custom code or by using the inference table feature, which logs asynchronously.

  • ✗

    Increase the workload size of the serving endpoint.

    Why it's wrong here

    Increasing workload size provides more compute resources but does not address the synchronous logging overhead. The logging operation still blocks the inference response. While larger instances may handle concurrent requests better, the latency per request due to logging remains. This is not a targeted solution for reducing logging-induced latency.

About these practice questions

Courseiva writes every Databricks-ML-Pro question from scratch — 300 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-ML-Pro practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Pro exam.