Databricks-GenAI-Assoc Design Applications Practice Question
A GenAI engineer is designing a Databricks application that must ground answers in a large corpus of internal policy documents. The corpus is updated by a nightly Delta job, and the application must cite the source document for every answer. The engineer is deciding how to structure the retrieval and generation stages. Which TWO design choices best satisfy the grounding and citation requirements? (Choose two.)
⚠ Common exam trap
The trap here is believing a capable model can supply citations from memory, when only metadata carried through retrieval can produce citations that are verifiable against the source documents.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Instruct the model to answer only from the retrieved context and to include the provided document identifiers in its response
Grounding and citation depend on two things working together: retrieval that returns source identifiers with each chunk, and a generation prompt that restricts the model to that retrieved context while requiring it to surface the identifiers. Metadata columns supply verifiable provenance, and the constrained instruction keeps the answer tied to retrieved policy text instead of the model's own memory.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Instruct the model to answer only from the retrieved context and to include the provided document identifiers in its response
Why this is correct
Constraining the prompt to the retrieved context reduces unsupported generation, and requiring the model to echo the supplied identifiers ties each statement back to a retrieved chunk. Because the identifiers originate from the index metadata rather than the model's memory, the citations remain verifiable against the source documents.
- ✗
Rely on the model's pretrained knowledge of internal policies to fill gaps when retrieval returns no relevant chunks
Why it's wrong here
Internal policy documents are not part of any public pretrained corpus, so the model has no reliable knowledge of them. Allowing the model to fill gaps from its own weights produces confident but ungrounded statements that cannot be cited, directly undermining the citation requirement and increasing hallucination risk.
- ✓
Store chunk text and source metadata such as document ID and section as columns alongside the embeddings in the Vector Search index
Why this is correct
Keeping source metadata as columns next to each embedding means every similarity result returns the document ID and section needed to build a citation. Without this, the generation stage would receive text with no provenance, and the application could not reliably attribute an answer to a specific policy document or section.
- ✗
Cache the full corpus in the prompt on every request so the model always has complete policy context
Why it's wrong here
Loading the entire corpus into each prompt exceeds context limits for a large document set, inflates latency and cost dramatically, and still does not guarantee the model attends to the relevant passage. It also provides no structured provenance, so citing a specific document or section becomes guesswork rather than a lookup.
- ✗
Increase the model's temperature setting so the generation stage explores a wider range of possible answers
Why it's wrong here
Higher temperature increases randomness in token selection, which makes answers less deterministic and more likely to drift from the retrieved context. For a policy question-and-answer application requiring faithful citations, variability is a liability rather than a benefit, and it does nothing to improve grounding or provenance.
About these practice questions
This Databricks-GenAI-Assoc question is part of Courseiva's 330-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.