Courseiva
Data Preparation →mediumMultiple Choice

NCP-GENL Data Preparation Practice Question

You are curating a 200 GB instruction-tuning corpus for an NVIDIA NeMo fine-tuning job on a Llama-based model. Post-training evaluation reveals the model regurgitates exact validation-set passages verbatim. An audit shows that near-duplicate instruction/response pairs were split randomly at the record level across train and validation partitions. Which data preparation change most directly eliminates this leakage while preserving the maximum amount of usable training data?

⚠ Common exam trap

The trap here is assuming that deduplication must be exact or that adjusting the split ratio can fix leakage caused by near-duplicate records spanning partitions.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Perform MinHash-based near-duplicate detection across the full corpus, cluster records by similarity, and assign whole clusters to either the train or validation split using a deterministic hash of the cluster ID.

The leakage arises because near-duplicate instruction/response pairs were split at the record level, allowing paraphrased twins to appear in both partitions. Clustering near-duplicates and assigning entire clusters to one split closes that pathway while retaining all unique content for training. Deterministic cluster-to-split hashing also makes the partition reproducible, which matters for auditing and for comparing fine-tuning runs fairly.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Increase the validation split ratio from 5% to 20% so that fewer training records overlap with the validation partition.

    Why it's wrong here

    Changing the split ratio does not stop near-duplicate pairs from straddling the boundary; with random record-level splitting, a paraphrased twin of a validation example can still land in training regardless of the ratio. This approach also removes 15% more data from training, reducing usable corpus size without addressing the actual duplication mechanism that produced the leakage.

  • ✓

    Perform MinHash-based near-duplicate detection across the full corpus, cluster records by similarity, and assign whole clusters to either the train or validation split using a deterministic hash of the cluster ID.

    Why this is correct

    Grouping near-duplicates into clusters and assigning entire clusters to one partition prevents the same or paraphrased content from appearing in both train and validation, which is exactly what caused the verbatim regurgitation. Because only duplicate clusters are collapsed rather than all similar-looking records, the maximum amount of unique training data is retained, and the deterministic hash keeps splits reproducible across reruns.

  • ✗

    Apply aggressive token-level cleaning to strip boilerplate phrases and punctuation from every record before splitting the dataset randomly.

    Why it's wrong here

    Token-level normalization may reduce superficial overlap, but semantically equivalent instructions and responses survive cleaning and still straddle the random split. The root cause is record-level partitioning of near-duplicate content, not noisy tokens. This option therefore leaves the leakage pathway intact while discarding potentially informative surface features and adding preprocessing cost without measurable benefit.

  • ✗

    Deduplicate only the validation partition by removing any record whose exact text appears in the training partition, then keep the original random split.

    Why it's wrong here

    Exact-match deduplication misses paraphrased or partially overlapping duplicates, which are precisely the pairs that caused verbatim regurgitation. It also only cleans the validation side, so a near-duplicate in training with no exact validation twin still leaks. The leakage mechanism is near-duplication, and exact matching is too weak a criterion to catch it reliably.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.