Courseiva
Data Preparation →hardMultiple Choice

NCP-GENL Data Preparation Practice Question

You are building a pretraining dataset from a large collection of source-code repositories for an NVIDIA NeMo LLM. The data includes many files with licenses, generated code, and minified JavaScript. Which NeMo Curator-based approach best improves code data quality before tokenization?

⚠ Common exam trap

The trap here is believing that a more powerful tokenizer or a larger context window can compensate for low-quality or repetitive code data.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Use language-specific filters that detect and remove minified files, license headers, and auto-generated code, then apply deduplication.

Language-specific filters combined with deduplication directly target the described quality issues. Minified files, license headers, and generated code are common in code corpora and can be detected with heuristics such as average line length, ratio of whitespace, and presence of standard license text. Deduplication removes repeated generated files. This preprocessing ensures the pretraining corpus contains meaningful code, which improves the model's code capabilities.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Convert all code to a single programming language using an automated transpiler before tokenization.

    Why it's wrong here

    Transpiling all code to one language is infeasible and would destroy language-specific syntax and idioms that the model needs to learn. It also introduces transpilation errors and loses information. The goal is to pretrain on diverse code, so preserving the original languages is essential. This approach does not remove minified or generated files and adds unnecessary complexity.

  • ✗

    Increase the model's context window to 32k tokens so that entire repositories fit in one sequence.

    Why it's wrong here

    A larger context window does not clean the data; it only allows longer sequences. Minified files and license headers would still be present and would consume context with low-value tokens. The scenario is about data quality, not sequence length. Increasing context also raises memory and compute costs significantly, without addressing the root issue of noisy or repetitive code.

  • ✗

    Tokenize all files with a byte-level BPE tokenizer and skip any further filtering.

    Why it's wrong here

    Byte-level BPE handles any character, but it does not remove low-quality or repetitive content. Minified files would still be tokenized into long sequences of short tokens, and license headers would remain. Skipping filtering means the model trains on noise, which can degrade performance. Tokenization is necessary but not sufficient; filtering and deduplication are also required.

  • ✓

    Use language-specific filters that detect and remove minified files, license headers, and auto-generated code, then apply deduplication.

    Why this is correct

    Minified JavaScript, license headers, and generated code are low-value or repetitive patterns that harm code pretraining. NeMo Curator supports custom filters for line length, ratio of alphanumeric characters, and detection of boilerplate. Deduplication removes repeated generated files. Applying these before tokenization ensures the model learns from meaningful code rather than noise, improving downstream code generation and understanding.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.