NCP-GENL Data Preparation Practice Question
You are using NVIDIA NeMo Curator to filter a 600 GB web-crawl corpus before pre-training. Your team wants to remove exact duplicates and near-duplicates to reduce memorization and speed up training. Which NeMo Curator stage should you apply?
⚠ Common exam trap
The trap here is assuming that any quality filter removes duplicates, when quality filters score documents individually rather than comparing them to each other.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
ExactDuplicates and FuzzyDuplicates stages
NeMo Curator's ExactDuplicates and FuzzyDuplicates stages are purpose-built for removing redundant documents. ExactDuplicates handles byte-level matches, while FuzzyDuplicates uses MinHash and LSH to catch near-identical texts. Together they reduce corpus size without discarding unique content, which directly lowers memorization risk and training cost on a large web-crawl dataset.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
HeuristicFilter stage with a quality threshold
Why it's wrong here
HeuristicFilter removes low-quality documents based on rules such as word count, perplexity, or symbol ratios, but it does not compare documents against each other. It cannot detect exact or near-duplicate content, so it leaves redundancy in the corpus and does not address the memorization and compute concerns described.
- ✗
DocumentDownloader stage followed by DocumentResolver
Why it's wrong here
DocumentDownloader and DocumentResolver are ingestion stages that fetch and resolve source documents; they do not perform deduplication. Running them would at best re-fetch the same content, and they provide no exact or fuzzy matching capability, so the near-duplicate and exact-duplicate removal goal would not be met.
- ✓
ExactDuplicates and FuzzyDuplicates stages
Why this is correct
NeMo Curator provides ExactDuplicates and FuzzyDuplicates stages that identify redundant documents; ExactDuplicates catches byte-identical texts, while FuzzyDuplicates uses MinHash/LSH to find near-duplicates. Applying both in sequence on the 600 GB corpus removes redundant content, lowering memorization risk and reducing the effective training token count.
- ✗
ClassifierFilter stage using a domain classifier
Why it's wrong here
ClassifierFilter scores documents with a trained model to keep or drop them by domain relevance or quality. It evaluates each document independently and has no pairwise comparison mechanism, so it will not identify duplicate or near-duplicate documents, leaving the redundancy problem unsolved.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.