Courseiva
Data Preparation →mediumMultiple Choice

NCP-GENL Data Preparation Practice Question

You are preparing a dataset of support tickets for a RAG system using NVIDIA NeMo. Many tickets are short and contain little context, which hurts retrieval quality. Which data preparation technique best improves retrieval by enriching each ticket with related information before embedding?

⚠ Common exam trap

The trap here is assuming that retrieval-time tuning, such as lowering a similarity threshold, can compensate for source documents that simply lack contextual content.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use an LLM to generate contextual summaries or metadata for each ticket and prepend them to the text before embedding.

Short tickets embed poorly because they carry little semantic signal. Using an LLM to generate contextual summaries or metadata and prepending that text before embedding enriches each record, so the vector captures more of the ticket's intent and domain. This improves matching with user queries. Deduplication, larger embeddings, and lower thresholds do not add missing context, so they cannot resolve the sparsity problem in this scenario.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the embedding dimension of the retrieval model.

    Why it's wrong here

    A larger embedding dimension can capture more detail, but it cannot invent information that is absent from a short ticket. The fundamental problem is insufficient content, not insufficient representational capacity. Changing the embedding size does not enrich the source text and therefore does not resolve the retrieval quality issue described.

  • ✗

    Lower the similarity threshold used during retrieval.

    Why it's wrong here

    Lowering the similarity threshold returns more matches, including irrelevant ones, which can reduce precision without fixing the underlying sparse representations. The tickets still lack context, so their embeddings remain weak. This retrieval-time adjustment does not perform data preparation and fails to address the root cause of poor retrieval quality.

  • ✗

    Apply deduplication to remove similar tickets from the corpus.

    Why it's wrong here

    Deduplication reduces redundancy but does not add context to short tickets. After deduplication, the remaining tickets are still sparse and may embed poorly, so retrieval quality does not improve in the way the scenario requires. This technique addresses corpus size, not the lack of contextual information in individual records.

  • ✓

    Use an LLM to generate contextual summaries or metadata for each ticket and prepend them to the text before embedding.

    Why this is correct

    Generating a summary or metadata adds missing context to each short ticket, so the embedded representation captures more relevant semantics and matches user queries better. Prepending this enriched text before embedding directly addresses the sparsity problem and improves retrieval, making it the appropriate technique for the scenario.

Visual reference

Client Recursive Resolver Root DNS (13 root servers) TLD DNS (.com, .org, …) Authoritative example.com query IP addr answer

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.