A developer is tuning a retrieval-augmented generation pipeline that uses NVIDIA NIM embeddings and a NIM LLM. Latency is dominated by embedding thousands of document chunks at query time because the team re-embeds the whole corpus on every request. Which change most directly fixes the architecture?
The latency comes from re-embedding the full corpus per request, which is wasteful because document chunks are static. Generating embeddings once and storing them in a vector index means each query only needs a single embedding call plus a nearest-neighbor lookup. This is the standard RAG pattern and removes the dominant cost while preserving retrieval quality.
Why this answer
When latency is driven by re-embedding static documents per request, the correct remedy is to embed the corpus once, store the vectors in an index, and embed only the query at serving time. Larger embedding models, longer generation, and answer caches do not remove the redundant full-corpus embedding work and can even increase cost or staleness.
Exam trap
The trap here is treating an answer cache or a bigger embedding model as the fix, when the real problem is embedding static documents repeatedly instead of once.