Databricks-GenAI-Assoc Application Development Practice Question
A GenAI engineer has registered a RAG chain in Unity Catalog as a model and now needs to deploy it for real-time inference with per-request token usage and latency captured automatically. Which Databricks capability should they enable on the serving endpoint?
⚠ Common exam trap
The trap here is assuming that any monitoring-adjacent feature, such as autoscaling or lineage, will automatically capture request-level telemetry for a served model.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable inference tables on the Mosaic AI Model Serving endpoint.
Inference tables are the native Mosaic AI Model Serving feature that persists request and response payloads plus metadata to a Delta table, giving automatic capture of token usage and latency for real-time endpoints. Autoscaling, MLflow experiments, and Unity Catalog lineage address capacity, development tracking, and governance respectively, but none record per-request serving telemetry.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Attach an MLflow experiment to the endpoint and rely on the run history for telemetry.
Why it's wrong here
MLflow experiments track runs during development, not live serving traffic. Once a model is deployed to a serving endpoint, requests are not automatically logged as MLflow runs, so experiment history would not contain production token usage or latency. It is a development-time tracking tool, not a serving observability mechanism.
- ✗
Enable Unity Catalog lineage on the registered model version.
Why it's wrong here
Unity Catalog lineage records data and model dependencies such as source tables and notebooks. It does not capture per-request payloads, token counts, or latency from a serving endpoint. While valuable for governance and impact analysis, lineage operates at the asset relationship level, not the individual inference request level, so it cannot fulfill this requirement.
- ✗
Configure the endpoint with autoscaling enabled to log per-request metrics.
Why it's wrong here
Autoscaling adjusts the number of concurrent model replicas based on traffic. It has no role in capturing per-request payloads or token counts, and it does not write any request-level data to a Delta table. Enabling it changes capacity behavior, not observability, so it cannot satisfy the requirement for per-request token usage and latency capture.
- ✓
Enable inference tables on the Mosaic AI Model Serving endpoint.
Why this is correct
Inference tables on Mosaic AI Model Serving automatically capture the request payload, response, and metadata for each served request, enabling monitoring of token usage and latency without custom logging code. For a RAG chain registered in Unity Catalog, this gives immediate observability into production traffic and is the supported mechanism for capturing real-time inference telemetry.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.