NCP-GENL Data Preparation Practice Question
You are building a pretraining dataset from a large collection of source-code repositories for an NVIDIA NeMo LLM. The data includes many files with licenses, generated code, and minified JavaScript. Which NeMo Curator-based approach best improves code data quality before tokenization?
⚠ Common exam trap
The trap here is believing that a more powerful tokenizer or a larger context window can compensate for low-quality or repetitive code data.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use language-specific filters that detect and remove minified files, license headers, and auto-generated code, then apply deduplication.
Language-specific filters combined with deduplication directly target the described quality issues. Minified files, license headers, and generated code are common in code corpora and can be detected with heuristics such as average line length, ratio of whitespace, and presence of standard license text. Deduplication removes repeated generated files. This preprocessing ensures the pretraining corpus contains meaningful code, which improves the model's code capabilities.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Convert all code to a single programming language using an automated transpiler before tokenization.
Why it's wrong here
Transpiling all code to one language is infeasible and would destroy language-specific syntax and idioms that the model needs to learn. It also introduces transpilation errors and loses information. The goal is to pretrain on diverse code, so preserving the original languages is essential. This approach does not remove minified or generated files and adds unnecessary complexity.
- ✗
Increase the model's context window to 32k tokens so that entire repositories fit in one sequence.
Why it's wrong here
A larger context window does not clean the data; it only allows longer sequences. Minified files and license headers would still be present and would consume context with low-value tokens. The scenario is about data quality, not sequence length. Increasing context also raises memory and compute costs significantly, without addressing the root issue of noisy or repetitive code.
- ✗
Tokenize all files with a byte-level BPE tokenizer and skip any further filtering.
Why it's wrong here
Byte-level BPE handles any character, but it does not remove low-quality or repetitive content. Minified files would still be tokenized into long sequences of short tokens, and license headers would remain. Skipping filtering means the model trains on noise, which can degrade performance. Tokenization is necessary but not sufficient; filtering and deduplication are also required.
- ✓
Use language-specific filters that detect and remove minified files, license headers, and auto-generated code, then apply deduplication.
Why this is correct
Minified JavaScript, license headers, and generated code are low-value or repetitive patterns that harm code pretraining. NeMo Curator supports custom filters for line length, ratio of alphanumeric characters, and detection of boilerplate. Deduplication removes repeated generated files. Applying these before tokenization ensures the model learns from meaningful code rather than noise, improving downstream code generation and understanding.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.