Courseiva
Data Preparation →hardMultiple Select

NCP-GENL Data Preparation Practice Question

You are preparing a massive dataset for training a NeMo-based LLM. Which TWO data preprocessing steps are critical to prevent data leakage and ensure model quality?

⚠ Common exam trap

Test-takers often confuse basic random splitting with temporal or content-based splitting, missing the fact that standard random splits fail to prevent data leakage in LLM datasets.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Performing temporal or content-based splitting to isolate the validation set.

Data leakage occurs when test data is inadvertently included in the training set, leading to inflated performance metrics. Deduplication is equally vital, as repetitive data causes the model to memorize samples rather than generalize. In NVIDIA workflows, these steps are typically performed via distributed scripts on the cluster before tokenization, ensuring that the model learns unique, non-overlapping information across all shards.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Performing temporal or content-based splitting to isolate the validation set.

    Why this is correct

    Isolating data based on timestamps or content clusters prevents the model from seeing future data or overlapping information during training. This ensures the evaluation set remains truly unseen, providing a realistic assessment of how the model will perform on new, unseen data in production environments.

  • ✓

    Implementing fuzzy deduplication to remove near-duplicate documents.

    Why this is correct

    Fuzzy deduplication identifies and removes documents that are structurally similar, preventing the model from over-indexing on repetitive patterns. This improves training efficiency and prevents the model from developing a bias toward specific, over-represented phrases or document structures, leading to better generalization on diverse, novel inputs.

  • ✗

    Increasing the frequency of data augmentation in the training pipeline.

    Why it's wrong here

    While augmentation can diversify data, it does not address data leakage or inherent redundancies. Excessive augmentation might actually introduce synthetic artifacts that confuse the model, and it is not a substitute for rigorous dataset sanitization and splitting protocols required to ensure scientific validity in model training.

  • ✗

    Adding metadata tags to all training sequences for classification.

    Why it's wrong here

    Metadata tagging is useful for instruct-tuning or multi-task learning but does not prevent leakage or deduplication errors. Including metadata without proper filtering might actually lead to unintentional leakage if the tags provide hints about the content that is meant to be held out for testing.

  • ✗

    Converting all text to lowercase to increase vocabulary efficiency.

    Why it's wrong here

    Lowercasing is a normalization technique, not a strategy for preventing leakage. While it reduces the vocabulary size, it can result in the loss of important semantic information, such as distinguishing between proper nouns and common nouns, which may be detrimental for specific downstream LLM applications.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.