Databricks-GenAI-Assoc Design Applications Practice Question
A team is deploying a RAG application using Databricks Model Serving with a foundation model endpoint and a Databricks Vector Search index. During load testing, they observe that p95 latency spikes when the retriever returns many chunks, and the LLM occasionally truncates context. They want to reduce latency while preserving answer quality. Which change is most appropriate?
⚠ Common exam trap
The trap here is treating temperature as a performance lever, when latency in RAG is dominated by prompt length and the volume of retrieved context, not by sampling randomness.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply a reranker to the retrieved chunks and pass only the top few highest-scoring chunks to the LLM.
Reranking retrieved chunks and passing only the top few to the LLM reduces prompt length, which lowers LLM latency and mitigates context truncation. Because the reranker prioritizes the most query-relevant chunks, answer quality is preserved even with fewer chunks. The other options either increase prompt size, change an unrelated parameter, or alter the index without addressing prompt length.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of retrieved chunks and rely on the LLM to ignore irrelevant content.
Why it's wrong here
Increasing retrieved chunks increases prompt length, which raises latency and cost and makes truncation more likely. Foundation models do not reliably ignore irrelevant context; extra chunks can distract the model and reduce answer quality. This change directly worsens the observed p95 latency spikes and context truncation rather than resolving them, so it is the opposite of the required optimization.
- ✗
Switch the Vector Search index to a smaller embedding dimension and keep the same number of chunks.
Why it's wrong here
Reducing embedding dimension can speed up vector similarity computation slightly, but the dominant latency contributor in this scenario is LLM prompt processing of many chunks. Keeping the same number of chunks means the prompt length and truncation risk remain unchanged. This change also requires re-embedding the corpus and may degrade retrieval quality, so it does not reliably solve the observed problem.
- ✓
Apply a reranker to the retrieved chunks and pass only the top few highest-scoring chunks to the LLM.
Why this is correct
A reranker scores retrieved chunks against the query with higher precision than the initial vector similarity, so the application can pass fewer but more relevant chunks to the LLM. This shortens the prompt, reducing latency and the chance of context truncation, while preserving answer quality because the most relevant evidence is retained. It directly addresses both symptoms observed during load testing.
- ✗
Lower the LLM temperature to zero to reduce response time.
Why it's wrong here
Temperature controls sampling randomness, not the number of tokens processed. Setting temperature to zero makes output more deterministic but does not shorten the prompt or reduce the computational cost of processing many retrieved chunks. The latency spikes are caused by prompt length and retrieval volume, so this change does not address the root cause and may not measurably improve p95 latency.
About these practice questions
One of 330 original Databricks-GenAI-Assoc practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official Databricks exam blueprint
This Databricks-GenAI-Assoc practice question is part of Courseiva's free Databricks certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the Databricks-GenAI-Assoc exam.