NCP-GENL Data Preparation Practice Question
You are preparing a dataset for continued pre-training of an LLM on internal engineering documents. The corpus contains many documents with boilerplate headers, footers, and legal disclaimers repeated across files. Which NeMo Curator approach best reduces this boilerplate while preserving unique technical content?
⚠ Common exam trap
The trap here is assuming that document-level deduplication or filtering will remove boilerplate, when the repeated text is only a small portion of otherwise unique documents.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the n-gram or sentence-level deduplication stage to remove repeated spans
Span-level deduplication in NeMo Curator detects repeated n-grams or sentences across documents and removes them, which precisely targets boilerplate headers, footers, and legal disclaimers. Unlike document-level filtering, it preserves the unique technical content within each file, making it the best choice for a corpus where redundancy is localized rather than whole-document.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Use the n-gram or sentence-level deduplication stage to remove repeated spans
Why this is correct
NeMo Curator supports span-level deduplication, which can identify and remove repeated n-grams or sentences that appear across many documents. This directly targets boilerplate headers, footers, and disclaimers while leaving unique technical passages intact. It is the most precise way to reduce redundancy without discarding entire documents.
- ✗
Apply a heuristic filter that drops documents below a word-count threshold
Why it's wrong here
A word-count threshold removes short documents but does not target repeated boilerplate within longer documents. Engineering documents with substantial unique content but also boilerplate would survive, so the repetitive text remains. This approach also risks discarding short but valuable technical notes, making it a poor fit for the scenario.
- ✗
Lowercase all text and remove punctuation
Why it's wrong here
Lowercasing and punctuation removal are normalization steps that change the text's form but do not eliminate repeated boilerplate. The same disclaimers would still appear, just in a different case or without punctuation. This approach does not address the core problem of redundancy and may harm technical content by removing meaningful punctuation.
- ✗
Split documents into fixed-size chunks and keep only the first chunk
Why it's wrong here
Keeping only the first chunk assumes boilerplate is always at the beginning, which is not true for footers and disclaimers. It also discards potentially unique technical content from later chunks. This crude truncation would reduce data volume but not selectively remove boilerplate, and it risks losing valuable information.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.