Courseiva
Data Preparation →mediumMultiple Choice

NCP-GENL Data Preparation Practice Question

You are preparing a multilingual corpus for pre-training an LLM. The corpus contains documents in English, Spanish, and German, but the English portion is 80% of the data. You want to avoid the model becoming biased toward English. Which data preparation technique should you apply?

⚠ Common exam trap

The trap here is thinking that tokenizer changes or translation solve imbalance, when the core issue is the proportion of training examples per language.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Oversample the non-English documents or undersample the English documents

Balancing the language distribution by oversampling minority languages or undersampling the dominant one directly addresses the 80% English skew. This ensures the model sees a more representative mix during training, reducing bias toward English and improving performance on Spanish and German. It is a standard data preparation technique for multilingual corpora.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Remove all English documents to force multilingual learning

    Why it's wrong here

    Removing English entirely would discard a large amount of valuable data and likely degrade overall performance, including English proficiency. The goal is balance, not elimination. This extreme measure would harm the model's ability to generalize and is not a recommended data preparation technique for multilingual training.

  • ✗

    Translate all documents into English

    Why it's wrong here

    Translating everything into English would eliminate multilingual capability entirely and defeat the purpose of a multilingual corpus. It also introduces translation artifacts and loses native language nuances. This approach does not address the imbalance; it merely removes the other languages, which is counterproductive for a multilingual model.

  • ✗

    Apply a language-specific tokenizer for each language

    Why it's wrong here

    Using separate tokenizers per language would complicate training and is not typical for a single multilingual model. While tokenizers can impact efficiency, they do not directly address data imbalance. The scenario asks for a technique to avoid English bias, and tokenizer choice alone does not change the proportion of training data per language.

  • ✓

    Oversample the non-English documents or undersample the English documents

    Why this is correct

    Adjusting the sampling ratios by oversampling minority languages or undersampling the dominant language balances the effective training distribution. This reduces English bias and encourages the model to allocate capacity to Spanish and German. It is a standard technique in multilingual data preparation and can be implemented via NeMo Curator or custom sampling logic.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.