NCA-GENL Core Machine Learning and AI Knowledge Practice Question
A developer is building a retrieval-augmented generation (RAG) system for an internal knowledge base. The system must answer questions using company documents that are updated frequently. Which component is primarily responsible for retrieving the most relevant document chunks to include in the LLM's context?
⚠ Common exam trap
The trap here is attributing retrieval to the LLM's attention or fine-tuning, when in a RAG architecture the vector database and similarity search perform the actual document selection.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The vector database and its similarity search.
Retrieval-augmented generation separates knowledge retrieval from generation. The vector database stores embeddings of document chunks, and a similarity search matches the query embedding to the most relevant chunks. Those chunks are then inserted into the LLM's context. This design allows the knowledge base to be updated independently of the model, which is essential for frequently changing internal documents.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The fine-tuning dataset used to adapt the LLM.
Why it's wrong here
Fine-tuning adapts model weights to a task or domain but does not provide dynamic retrieval of current documents. In a RAG setup, the knowledge base changes frequently, and fine-tuning would be too slow and costly to update. The retrieval component must query an external index at inference time, not rely on static fine-tuning data.
- ✓
The vector database and its similarity search.
Why this is correct
In a RAG system, documents are chunked and embedded into vectors stored in a vector database. When a query arrives, its embedding is compared to stored vectors using similarity search to retrieve the most relevant chunks. This retrieval step is what supplies the LLM with grounded context, making the vector database and its search the primary responsible component.
- ✗
The LLM's attention mechanism.
Why it's wrong here
The attention mechanism operates within the LLM's context window to weigh relationships between tokens already provided. It does not search an external document store or select chunks from a knowledge base. Retrieval happens before the LLM receives the context, so attention is not the component responsible for fetching relevant documents.
- ✗
The tokenizer used to preprocess the query.
Why it's wrong here
The tokenizer converts text into tokens for the model, but it does not perform retrieval or rank documents. It is a preprocessing step that prepares input for embedding or generation. While tokenization affects how text is represented, it does not select relevant chunks from a knowledge base, so it is not the retrieval component.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.