Databricks-ML-Assoc Model Development Practice Question
A data scientist is deploying a model to a Databricks Model Serving endpoint. They observe that the inference latency is high. What should they check first?
⚠ Common exam trap
Candidates immediately assume the ML model itself is unoptimized, failing to investigate expensive inline data transformations occurring before the model receives the payload.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The complexity of the input data and preprocessing transformations.
High inference latency is often caused by heavy preprocessing steps or unoptimized model complexity. Checking the 'inference profile' or analyzing the time taken for data transformation versus the actual model prediction is the most critical first step. By separating these concerns, the data scientist can identify whether the bottleneck lies in the feature engineering pipeline (which might need optimization) or the model architecture itself, enabling targeted performance improvements.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The number of users accessing the workspace.
Why it's wrong here
The number of concurrent users in a workspace generally does not impact the inference latency of a dedicated Model Serving endpoint. Inference endpoints typically run on isolated compute resources, so activity by other users is unlikely to be the primary cause of latency unless there is extreme resource contention on the cloud provider.
- ✓
The complexity of the input data and preprocessing transformations.
Why this is correct
Preprocessing steps executed during inference are a frequent cause of latency. If the data requires complex joins, heavy transformations, or slow library calls, this will add directly to the total latency. Ensuring that these transformations are optimized or pre-calculated is usually the most effective way to reduce overall request latency.
- ✗
The version of the Databricks Runtime installed on the cluster.
Why it's wrong here
While runtime versions can affect performance, it is rarely the first thing to check for high latency. Model Serving endpoints use specific, optimized runtimes designed for low-latency inference. Changing the runtime is unlikely to solve a bottleneck caused by inefficient code or model size, making this an ineffective initial diagnostic step.
- ✗
The storage location of the training dataset.
Why it's wrong here
The location of the training dataset has no bearing on inference latency. Once a model is trained and deployed, it is completely decoupled from the original training data. Investigating the training data path will not yield any insights into why the live inference requests are taking longer than expected.
About these practice questions
Courseiva writes every Databricks-ML-Assoc question from scratch — 319 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-ML-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-ML-Assoc exam.