Databricks-GenAI-Assoc Assembling and Deploying Apps Practice Question
A team has deployed a model that is experiencing high latency. How should they identify if the bottleneck is the model inference or the preprocessing code?
⚠ Common exam trap
Candidates often guess that general platform monitoring metrics are sufficient. They fail to realize that MLflow Tracing is specifically required to break down the request into discrete components like retrieval and generation.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use MLflow Tracing to inspect the execution pipeline.
Distinguishing between inference time and preprocessing time is vital for performance tuning. By using MLflow's tracing capabilities or custom timing logs, engineers can isolate specific segments of the request pipeline. This granular visibility allows for targeted optimization, such as optimizing the vector search or reducing prompt overhead, rather than blindly attempting to speed up the model itself, which is often significantly more resource-intensive.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Check the cluster cost in the Databricks billing console.
Why it's wrong here
Billing consoles show financial costs but provide no insight into request-level latency or performance bottlenecks within the application code. This data is entirely unrelated to the execution time of specific functions like inference or preprocessing, making it useless for debugging performance issues in the model serving pipeline.
- ✓
Use MLflow Tracing to inspect the execution pipeline.
Why this is correct
MLflow Tracing provides detailed visibility into the duration of every step in the request flow, from input preprocessing to model inference and output post-processing. This allows engineers to pinpoint exactly where the latency is occurring, enabling them to optimize the specific components that are causing the performance delays.
- ✗
Re-train the model on a larger dataset.
Why it's wrong here
Re-training on a larger dataset will likely increase the model size and inference latency, making the performance issue worse. It does not address the preprocessing overhead or provide any diagnostic information regarding where the current bottlenecks exist, therefore it is counterproductive for solving latency problems in production.
- ✗
Increase the memory limit for the inference cluster.
Why it's wrong here
Increasing memory limits may prevent out-of-memory errors but will not resolve latency bottlenecks caused by inefficient code paths or slow data retrieval. Without observability, this is a guess that likely wastes resources without improving the actual response time for the end-user, which is the core metric to address.
About these practice questions
One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.