NCP-GENL Data Preparation Practice Question
You are curating instruction-tuning data for an NVIDIA NIM-deployed LLM. The raw dataset contains many near-duplicate instruction-response pairs that differ only by punctuation and whitespace. Which data preparation step is most appropriate to remove these before fine-tuning?
⚠ Common exam trap
The trap here is thinking that deduplication means exact string matching or global text normalization, when near-duplicates require similarity-based methods that preserve the original text.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Apply fuzzy deduplication using MinHash LSH over normalized instruction-response text.
Fuzzy deduplication with MinHash LSH is the standard approach for near-duplicate removal in instruction-tuning datasets. It normalizes text for comparison, generates MinHash signatures, and uses locality-sensitive hashing to find similar pairs efficiently at scale. This directly removes the redundant examples described, improving training efficiency and reducing overfitting without altering the semantic content of the retained data.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the batch size during fine-tuning to average out duplicate examples.
Why it's wrong here
Larger batches do not remove duplicates; they merely change the gradient averaging. Near-duplicate examples still bias the loss toward those patterns and can cause the model to memorize them. Data cleaning must happen before training, not be compensated for by hyperparameters. Increasing batch size also raises memory requirements and may not be feasible on the available GPU, without solving the underlying data quality issue.
- ✗
Convert all instructions to lowercase and remove all punctuation globally.
Why it's wrong here
Global lowercasing and punctuation removal would destroy meaningful distinctions in the data, such as code snippets, proper nouns, and formatting that the model needs to learn. It also does not guarantee duplicate removal, because different instructions could become identical after normalization. Proper deduplication should normalize only for comparison purposes while preserving the original text for training.
- ✗
Use NVIDIA Triton Inference Server dynamic batching to filter duplicates at serving time.
Why it's wrong here
Triton dynamic batching groups inference requests to improve throughput; it has no mechanism to detect or remove duplicate training examples. Serving-time batching operates on live requests, not on a static dataset. Using it here would not affect the fine-tuning corpus at all, and the duplicates would still be present during training, leading to the same overfitting and wasted compute.
- ✓
Apply fuzzy deduplication using MinHash LSH over normalized instruction-response text.
Why this is correct
Fuzzy deduplication with MinHash LSH is designed to catch near-duplicates that differ by punctuation, whitespace, or minor edits. Normalizing text before hashing ensures that trivial variations map to the same signature. Removing these duplicates prevents the model from overfitting to repeated examples and reduces wasted training compute, which is especially important when fine-tuning an instruction-following model on a curated dataset.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.