NCP-GENL Data Preparation Practice Question
You are preparing a multilingual corpus for pre-training an LLM. The corpus contains documents in English, Spanish, and German, but the English portion is 80% of the data. You want to avoid the model becoming biased toward English. Which data preparation technique should you apply?
⚠ Common exam trap
The trap here is thinking that tokenizer changes or translation solve imbalance, when the core issue is the proportion of training examples per language.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Oversample the non-English documents or undersample the English documents
Balancing the language distribution by oversampling minority languages or undersampling the dominant one directly addresses the 80% English skew. This ensures the model sees a more representative mix during training, reducing bias toward English and improving performance on Spanish and German. It is a standard data preparation technique for multilingual corpora.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Remove all English documents to force multilingual learning
Why it's wrong here
Removing English entirely would discard a large amount of valuable data and likely degrade overall performance, including English proficiency. The goal is balance, not elimination. This extreme measure would harm the model's ability to generalize and is not a recommended data preparation technique for multilingual training.
- ✗
Translate all documents into English
Why it's wrong here
Translating everything into English would eliminate multilingual capability entirely and defeat the purpose of a multilingual corpus. It also introduces translation artifacts and loses native language nuances. This approach does not address the imbalance; it merely removes the other languages, which is counterproductive for a multilingual model.
- ✗
Apply a language-specific tokenizer for each language
Why it's wrong here
Using separate tokenizers per language would complicate training and is not typical for a single multilingual model. While tokenizers can impact efficiency, they do not directly address data imbalance. The scenario asks for a technique to avoid English bias, and tokenizer choice alone does not change the proportion of training data per language.
- ✓
Oversample the non-English documents or undersample the English documents
Why this is correct
Adjusting the sampling ratios by oversampling minority languages or undersampling the dominant language balances the effective training distribution. This reduces English bias and encourages the model to allocate capacity to Spanish and German. It is a standard technique in multilingual data preparation and can be implemented via NeMo Curator or custom sampling logic.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.