Databricks-GenAI-Assoc Design Applications Practice Question
A team is designing a Databricks RAG application that must return answers with citations to source documents. They plan to use Databricks Vector Search and want the LLM to reference specific chunks. Which two design choices are required to produce reliable citations? (Choose two.)
⚠ Common exam trap
The trap here is focusing on retrieval tuning knobs such as similarity metrics while overlooking that citations depend on carrying source metadata from the index into the prompt.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Store a stable document identifier and chunk metadata alongside each embedding in the Vector Search index.
Reliable citations require two things: the retriever must return source identifiers, and those identifiers must reach the LLM in the prompt. Storing document IDs and chunk metadata in the Vector Search index enables the first, and including them in the prompt context enables the second. Similarity metrics, temperature, and inference tables do not provide the linkage needed for citations.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the LLM temperature to encourage the model to generate more detailed citations.
Why it's wrong here
Higher temperature increases randomness and makes citations less reliable, not more. Citations depend on the presence of source identifiers in the prompt, not on sampling diversity. Raising temperature can also cause the model to invent plausible-looking but incorrect references.
- ✓
Store a stable document identifier and chunk metadata alongside each embedding in the Vector Search index.
Why this is correct
Citations require mapping a retrieved chunk back to its source document. Including a stable document ID and chunk metadata in the index payload allows the retriever to return that identifier with each result. Without this, the LLM cannot reference a specific source, and citations become guesses or are omitted entirely.
- ✗
Enable inference tables on the Vector Search endpoint to capture citation metadata.
Why it's wrong here
Inference tables log requests and responses for monitoring, but they do not inject source identifiers into the prompt or the index. They are an observability feature, not a mechanism for producing citations. Relying on them would leave the LLM without the information needed to cite sources.
- ✗
Set the Vector Search index to use a cosine similarity metric instead of dot product.
Why it's wrong here
The similarity metric affects ranking quality but does not provide citation information. Both cosine and dot product can retrieve relevant chunks, yet neither carries source identifiers into the prompt. Changing the metric does not address the requirement to map answers back to specific documents.
- ✓
Include the retrieved chunk identifiers and source references in the prompt context so the LLM can cite them in its response.
Why this is correct
The LLM can only cite what it sees. Passing chunk identifiers and source references in the prompt gives the model the information needed to produce inline citations. If the prompt omits these references, the model may fabricate sources or decline to cite, undermining the reliability of the citations.
About these practice questions
One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.