You are curating a 200 GB instruction-tuning corpus for an NVIDIA NeMo fine-tuning job on a Llama-based model. Post-training evaluation reveals the model regurgitates exact validation-set passages verbatim. An audit shows that near-duplicate instruction/response pairs were split randomly at the record level across train and validation partitions. Which data preparation change most directly eliminates this leakage while preserving the maximum amount of usable training data?
Grouping near-duplicates into clusters and assigning entire clusters to one partition prevents the same or paraphrased content from appearing in both train and validation, which is exactly what caused the verbatim regurgitation. Because only duplicate clusters are collapsed rather than all similar-looking records, the maximum amount of unique training data is retained, and the deterministic hash keeps splits reproducible across reruns.
Why this answer
The leakage arises because near-duplicate instruction/response pairs were split at the record level, allowing paraphrased twins to appear in both partitions. Clustering near-duplicates and assigning entire clusters to one split closes that pathway while retaining all unique content for training. Deterministic cluster-to-split hashing also makes the partition reproducible, which matters for auditing and for comparing fine-tuning runs fairly.
Exam trap
The trap here is assuming that deduplication must be exact or that adjusting the split ratio can fix leakage caused by near-duplicate records spanning partitions.