Courseiva
Data Preparation →mediumMultiple Choice

NCP-GENL Data Preparation Practice Question

You are curating a 2 TB corpus of NVIDIA technical documentation and Python code for continued pretraining of a NeMo-based LLM. A colleague proposes filtering out any document containing the token sequence 'CUDA' to reduce hardware-specific bias. What is the most appropriate response?

⚠ Common exam trap

The trap here is assuming that removing a vendor-specific keyword reduces bias, when it actually strips high-value domain signal that the continued pretraining corpus was built to provide.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Reject the filter, because removing documents based on a single high-signal domain token destroys the domain-specific signal the model needs and is better handled by deduplication and quality scoring.

Keyword-based deletion of a domain-defining token is a destructive filter, not a quality filter. Effective NeMo data curation relies on deduplication, quality heuristics, and classifier scoring to remove low-value or redundant records while preserving coherent, in-domain technical content. The corpus exists to teach NVIDIA-specific concepts, so removing 'CUDA' would directly undermine the training objective.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Accept the filter, because removing every occurrence of a hardware-specific token guarantees the model will not overfit to NVIDIA-specific APIs and will generalize better to other vendors.

    Why it's wrong here

    Generalization is not achieved by deleting domain-defining tokens during continued pretraining. The goal of this corpus is to strengthen in-domain capability, and stripping 'CUDA' removes coherent technical passages rather than reducing bias. Overfitting is mitigated through deduplication, data mixing ratios, and regularization, not through keyword excision that fragments otherwise high-quality documents.

  • ✗

    Accept the filter but apply it only to documents where 'CUDA' appears more than ten times, since low-frequency occurrences are harmless and high-frequency ones indicate redundant marketing material.

    Why it's wrong here

    Frequency thresholds on a single token are arbitrary and do not measure redundancy. A document mentioning 'CUDA' many times may be a legitimate programming guide with high training value, while one mentioning it once may be boilerplate. Redundancy is properly detected with MinHash-based near-duplicate detection or semantic clustering, not with per-token count cutoffs.

  • ✓

    Reject the filter, because removing documents based on a single high-signal domain token destroys the domain-specific signal the model needs and is better handled by deduplication and quality scoring.

    Why this is correct

    The proposed filter is a crude keyword removal that discards entire documents whose core subject is the target domain. In NeMo Curator pipelines, quality is improved through exact and fuzzy deduplication, heuristic quality filters, and classifier-based scoring, not by excising a single token that defines the domain. Removing these documents would leave the model undertrained on the very terminology it must learn.

  • ✗

    Accept the filter, because NVIDIA NeMo requires that proprietary product names be excluded from pretraining corpora to comply with the model card's data provenance requirements.

    Why it's wrong here

    NeMo does not impose any such requirement to exclude product names from pretraining data. Data provenance documentation describes dataset origin and licensing, but it does not mandate removal of vendor terminology. In fact, maintaining accurate domain terminology is essential when the downstream objective is to answer questions about NVIDIA platforms and their APIs.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.