Databricks-GenAI-Assoc Design Applications Practice Question
An engineer is designing a Databricks RAG application where the same retrieved context must be reused across several prompt variants during evaluation. They want to reduce token cost and improve consistency between variants. Which design choice best supports this?
⚠ Common exam trap
It's easy for candidates to confuse generation-side caching with retrieval reuse, when the cost and consistency problem originates in repeated retriever calls.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Compute the retrieval step once, store the retrieved documents in a Delta table keyed by query, and have each prompt variant read from that table.
Separating retrieval from generation and persisting the retrieved context lets several prompt variants share one retrieval result. This lowers cost because embeddings and Vector Search are invoked once, and it improves consistency because all variants see the same evidence. The alternatives change generation or chunking behavior but do not eliminate redundant retrieval.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Use a higher temperature during evaluation so the model explores more responses and the best one can be selected.
Why it's wrong here
Temperature affects generation randomness, not retrieval reuse or token cost. Raising it would make variants less comparable and could increase variability in evaluation. It does nothing to prevent the retriever from being called once per variant, so it fails to meet the stated objective.
- ✓
Compute the retrieval step once, store the retrieved documents in a Delta table keyed by query, and have each prompt variant read from that table.
Why this is correct
Caching retrieval results in a Delta table decouples retrieval from generation, so multiple prompt variants consume the same context without repeating embedding lookups or Vector Search calls. This reduces token and compute cost and ensures every variant is evaluated against identical context, which makes comparisons fair and reproducible.
- ✗
Configure the model serving endpoint to cache identical requests so repeated prompts with the same context are served from cache.
Why it's wrong here
Endpoint request caching targets identical generation requests, but prompt variants differ in wording, so their requests are not identical and would not hit the same cache entry. It also does not avoid the retrieval calls that precede generation. The scenario needs shared retrieval results, not generation-level caching.
- ✗
Increase the chunk size in the Vector Search index so that each retrieval returns more context, reducing the need to call the retriever multiple times.
Why it's wrong here
Larger chunks do not prevent repeated retrieval calls for each prompt variant; they only change the amount of text returned per call. They can also dilute relevance and consume more tokens per prompt. The problem is repeated retrieval, not insufficient context, so this adjustment does not address the cost or consistency goal.
About these practice questions
Courseiva writes every Databricks-GenAI-Assoc question from scratch — 330 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.