Courseiva
Evaluation and Monitoring →mediumMultiple Choice

Databricks-GenAI-Assoc Evaluation and Monitoring Practice Question

An enterprise LLM application on Databricks is experiencing high latency and inconsistent responses. Which approach best enables observability to identify the root cause of these performance bottlenecks within the LLM pipeline?

⚠ Common exam trap

Candidates often choose standard cluster log analysis or simple end-to-end latency metrics instead of MLflow Tracing to inspect individual steps within an LLM chain.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use MLflow Tracing to record inputs, outputs, and execution duration of each step in the chain.

Integrating MLflow Tracing provides granular visibility into the execution flow, including model inputs, outputs, and latency per step. This is critical for diagnosing complex chains where intermediate calls may be stalling. By instrumenting the code, engineers can identify which specific retrieval step or prompt generation block is contributing to the high latency, allowing for targeted optimization of the RAG pipeline components.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Enable standard Spark UI logs to monitor cluster resource utilization during inference.

    Why it's wrong here

    Spark UI exposes cluster CPU, memory and shuffle metrics, not per-request LLM traces, prompt tokens or retrieval timings, so it cannot isolate pipeline latency. It is tempting because Spark UI is the default observability surface for Databricks jobs, and would suffice for diagnosing slow Spark ETL stages rather than LLM inference.

  • ✗

    Review the Databricks SQL query history to identify slow retrieval operations.

    Why it's wrong here

    SQL query history covers retrieval queries only, ignoring prompt construction, model inference and orchestration stages where latency also arises. It is tempting because retrieval is a common bottleneck, but it would be correct for diagnosing slow SQL; full pipeline observability needs MLflow tracing across all stages.

  • ✓

    Use MLflow Tracing to record inputs, outputs, and execution duration of each step in the chain.

    Why this is correct

    MLflow Tracing captures end-to-end execution details for LLM workflows. It allows developers to visualize the entire dependency graph and latency of each step, enabling precise identification of bottlenecks in chains, prompt templates, or vector search lookups.

  • ✗

    Deploy an external monitoring agent to scan the model server network traffic.

    Why it's wrong here

    Network traffic scanning captures transport metadata, not per-stage latency inside the LLM pipeline, so it cannot isolate the bottleneck. It is tempting as an infrastructure monitoring approach, but it would suit network fault diagnosis; MLflow tracing is required to observe prompt, retrieval and inference timings.

About these practice questions

One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official Databricks exam blueprint

This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.