Courseiva
Data Preparation →hardMultiple Choice

NCP-GENL Data Preparation Practice Question

You are preparing a dataset for continued pre-training of an LLM on internal engineering documents. The corpus contains many documents with boilerplate headers, footers, and legal disclaimers repeated across files. Which NeMo Curator approach best reduces this boilerplate while preserving unique technical content?

⚠ Common exam trap

The trap here is assuming that document-level deduplication or filtering will remove boilerplate, when the repeated text is only a small portion of otherwise unique documents.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use the n-gram or sentence-level deduplication stage to remove repeated spans

Span-level deduplication in NeMo Curator detects repeated n-grams or sentences across documents and removes them, which precisely targets boilerplate headers, footers, and legal disclaimers. Unlike document-level filtering, it preserves the unique technical content within each file, making it the best choice for a corpus where redundancy is localized rather than whole-document.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Use the n-gram or sentence-level deduplication stage to remove repeated spans

    Why this is correct

    NeMo Curator supports span-level deduplication, which can identify and remove repeated n-grams or sentences that appear across many documents. This directly targets boilerplate headers, footers, and disclaimers while leaving unique technical passages intact. It is the most precise way to reduce redundancy without discarding entire documents.

  • ✗

    Apply a heuristic filter that drops documents below a word-count threshold

    Why it's wrong here

    A word-count threshold removes short documents but does not target repeated boilerplate within longer documents. Engineering documents with substantial unique content but also boilerplate would survive, so the repetitive text remains. This approach also risks discarding short but valuable technical notes, making it a poor fit for the scenario.

  • ✗

    Lowercase all text and remove punctuation

    Why it's wrong here

    Lowercasing and punctuation removal are normalization steps that change the text's form but do not eliminate repeated boilerplate. The same disclaimers would still appear, just in a different case or without punctuation. This approach does not address the core problem of redundancy and may harm technical content by removing meaningful punctuation.

  • ✗

    Split documents into fixed-size chunks and keep only the first chunk

    Why it's wrong here

    Keeping only the first chunk assumes boilerplate is always at the beginning, which is not true for footers and disclaimers. It also discards potentially unique technical content from later chunks. This crude truncation would reduce data volume but not selectively remove boilerplate, and it risks losing valuable information.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.