NCP-GENL Data Preparation Practice Question
You are preparing a dataset of support tickets for a RAG system using NVIDIA NeMo. Many tickets are short and contain little context, which hurts retrieval quality. Which data preparation technique best improves retrieval by enriching each ticket with related information before embedding?
⚠ Common exam trap
The trap here is assuming that retrieval-time tuning, such as lowering a similarity threshold, can compensate for source documents that simply lack contextual content.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use an LLM to generate contextual summaries or metadata for each ticket and prepend them to the text before embedding.
Short tickets embed poorly because they carry little semantic signal. Using an LLM to generate contextual summaries or metadata and prepending that text before embedding enriches each record, so the vector captures more of the ticket's intent and domain. This improves matching with user queries. Deduplication, larger embeddings, and lower thresholds do not add missing context, so they cannot resolve the sparsity problem in this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the embedding dimension of the retrieval model.
Why it's wrong here
A larger embedding dimension can capture more detail, but it cannot invent information that is absent from a short ticket. The fundamental problem is insufficient content, not insufficient representational capacity. Changing the embedding size does not enrich the source text and therefore does not resolve the retrieval quality issue described.
- ✗
Lower the similarity threshold used during retrieval.
Why it's wrong here
Lowering the similarity threshold returns more matches, including irrelevant ones, which can reduce precision without fixing the underlying sparse representations. The tickets still lack context, so their embeddings remain weak. This retrieval-time adjustment does not perform data preparation and fails to address the root cause of poor retrieval quality.
- ✗
Apply deduplication to remove similar tickets from the corpus.
Why it's wrong here
Deduplication reduces redundancy but does not add context to short tickets. After deduplication, the remaining tickets are still sparse and may embed poorly, so retrieval quality does not improve in the way the scenario requires. This technique addresses corpus size, not the lack of contextual information in individual records.
- ✓
Use an LLM to generate contextual summaries or metadata for each ticket and prepend them to the text before embedding.
Why this is correct
Generating a summary or metadata adds missing context to each short ticket, so the embedded representation captures more relevant semantics and matches user queries better. Prepending this enriched text before embedding directly addresses the sparsity problem and improves retrieval, making it the appropriate technique for the scenario.
Visual reference
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.