Courseiva
Software Development →mediumMultiple Choice

NCA-GENL Software Development Practice Question

A developer is tuning a retrieval-augmented generation pipeline that uses NVIDIA NIM embeddings and a NIM LLM. Latency is dominated by embedding thousands of document chunks at query time because the team re-embeds the whole corpus on every request. Which change most directly fixes the architecture?

⚠ Common exam trap

The trap here is treating an answer cache or a bigger embedding model as the fix, when the real problem is embedding static documents repeatedly instead of once.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Precompute and persist chunk embeddings in a vector index once, then embed only the incoming query and retrieve nearest neighbors at request time.

When latency is driven by re-embedding static documents per request, the correct remedy is to embed the corpus once, store the vectors in an index, and embed only the query at serving time. Larger embedding models, longer generation, and answer caches do not remove the redundant full-corpus embedding work and can even increase cost or staleness.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Cache the final LLM answers in memory and serve repeats, leaving the per-request embedding of the corpus unchanged.

    Why it's wrong here

    Answer caching helps only for repeated identical queries and does nothing for first-time or varied questions, which still trigger full-corpus embedding. It also risks serving stale answers when documents change. The dominant cost identified in the scenario remains untouched, so this is a partial workaround rather than a fix.

  • ✗

    Switch the embedding NIM to a larger model with higher dimensionality so fewer chunks are needed to cover the corpus.

    Why it's wrong here

    A larger embedding model increases per-call cost and storage, and it does not change the fact that the whole corpus is re-embedded on every request. Higher dimensionality can slightly improve recall but makes the lookup and storage heavier. The architectural flaw of redundant full-corpus embedding remains, so latency would likely worsen rather than improve.

  • ✗

    Increase the LLM's max_tokens so it can summarize the raw documents directly instead of retrieving chunks.

    Why it's wrong here

    Raising max_tokens lengthens generation and increases latency, and the model still needs the document text in context, which requires retrieval or embedding anyway. Summarizing raw documents per request does not eliminate the embedding bottleneck and risks exceeding the context window. This choice trades one cost for a larger one without fixing the architecture.

  • ✓

    Precompute and persist chunk embeddings in a vector index once, then embed only the incoming query and retrieve nearest neighbors at request time.

    Why this is correct

    The latency comes from re-embedding the full corpus per request, which is wasteful because document chunks are static. Generating embeddings once and storing them in a vector index means each query only needs a single embedding call plus a nearest-neighbor lookup. This is the standard RAG pattern and removes the dominant cost while preserving retrieval quality.

About these practice questions

Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.