NCP-GENL Data Preparation Practice Question
You are preparing a 2 TB corpus of English and German web text for continued pretraining of a NeMo-based LLM. The German portion includes many pages with unescaped HTML entities and mixed-language sentences. Which NeMo Curator stage should you apply to remove boilerplate, fix HTML artifacts, and filter low-quality documents before tokenization?
⚠ Common exam trap
The trap here is assuming that any NVIDIA component that processes text, such as Guardrails or Triton, can substitute for a dedicated data curation pipeline when preparing a training corpus.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
NeMo Curator's text cleaning and heuristic filtering pipeline, including HTML unescaping, language identification, and quality classifier stages.
NeMo Curator is the correct tool because it supplies production-grade stages for HTML unescaping, language identification, and quality filtering that operate on large-scale text before tokenization. These stages directly address the German/English mixed-language and malformed HTML issues. The other options are training, runtime safety, or serving components, none of which clean a pretraining corpus.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
NeMo Curator's text cleaning and heuristic filtering pipeline, including HTML unescaping, language identification, and quality classifier stages.
Why this is correct
NeMo Curator provides modular stages for exactly this: HTML unescaping, boilerplate removal, language identification using fastText, and quality filtering with a classifier or heuristic scores. Running these before tokenization ensures the German and English subsets are clean and consistently language-tagged, which reduces noise during continued pretraining and prevents mixed-language documents from degrading the model's language modeling.
- ✗
NeMo's Megatron-LM pretraining script with an increased dropout rate on the embedding layer.
Why it's wrong here
Dropout is a regularization technique applied during training, not a data preparation stage. It cannot remove HTML entities, detect language, or filter low-quality documents. Applying higher dropout to embeddings would not fix malformed text and may even harm learning, since the model still consumes noisy tokens. Data cleaning must occur before the tokenizer and training pipeline, not be addressed through model hyperparameters.
- ✗
NeMo Guardrails with a custom Colang flow that blocks documents containing HTML tags.
Why it's wrong here
NeMo Guardrails is designed for runtime safety and conversational control of LLM applications, not for offline corpus cleaning. It operates on user interactions and model outputs, not on raw web documents. While a Colang flow could theoretically match patterns, it is not a scalable data preparation tool and lacks the language identification, quality scoring, and deduplication capabilities needed for a multi-terabyte corpus.
- ✗
NVIDIA Triton Inference Server with a Python backend that preprocesses each document at inference time.
Why it's wrong here
Triton Inference Server serves models for inference; it is not a data preparation framework. Using it to preprocess documents at inference time would add latency and complexity without cleaning the training corpus. The scenario requires cleaning before tokenization and training, so a serving solution is architecturally misplaced and does not provide the required filtering, language detection, or deduplication stages.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.