NCP-GENL Data Preparation Practice Question
A team is preparing a mixed-language corpus for continued pretraining of an NVIDIA NeMo Megatron model. The corpus contains English, Japanese, and Arabic documents. Tokenizer analysis shows the current English-centric BPE vocabulary produces very long token sequences for Japanese and Arabic, inflating sequence length and compute cost. The team wants to reduce sequence length for non-English text without retraining the tokenizer from scratch and without degrading English performance. Which data preparation action best achieves this?
⚠ Common exam trap
The trap here is treating long token sequences as a context-length problem to be solved by enlarging the window rather than as a vocabulary coverage problem rooted in the tokenizer.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Extend the existing BPE vocabulary with additional merges learned from a balanced multilingual sample, then retrain only the embedding and output layers while freezing the rest of the model.
The high token counts for Japanese and Arabic stem from a vocabulary that lacks subword units for those scripts. Extending the BPE vocabulary with multilingual merges shortens their token sequences while preserving English merges, and retraining only the embedding and output layers adapts the new entries without perturbing the pretrained transformer. This lowers compute cost and keeps English behavior stable.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Lowercase and strip diacritics from all Japanese and Arabic text so that the English-centric vocabulary can represent it more efficiently.
Why it's wrong here
Arabic diacritics and Japanese kana are not equivalent to Latin case or accent marks; removing them destroys orthographic and semantic information and does not make the English BPE merges cover these scripts any better. The token sequences remain long because the vocabulary simply lacks the relevant subword units, so compute cost stays high while data quality drops.
- ✗
Transliterate all Japanese and Arabic documents into Latin script using a standard romanization scheme before tokenization.
Why it's wrong here
Romanization replaces native script with Latin approximations, which the English-centric vocabulary tokenizes more compactly, but it introduces ambiguity and loses information that the model would otherwise learn from authentic text. Downstream generation and evaluation would also require romanized inputs, creating a mismatch with real-world usage. The sequence-length benefit is achieved at the cost of linguistic fidelity and practical deployability.
- ✗
Increase the model's maximum sequence length and positional embedding size so that long token sequences for Japanese and Arabic fit without truncation.
Why it's wrong here
Extending the context window accommodates longer sequences but does not reduce their length, so per-sample compute cost stays high and grows with the larger attention matrix. It also requires modifying positional embeddings, which can destabilize the pretrained model. The underlying inefficiency, an English-centric vocabulary that over-segments non-English scripts, remains unaddressed.
- ✓
Extend the existing BPE vocabulary with additional merges learned from a balanced multilingual sample, then retrain only the embedding and output layers while freezing the rest of the model.
Why this is correct
Adding multilingual merges to the existing vocabulary reduces the number of tokens needed to represent Japanese and Arabic text, directly cutting sequence length and compute. Keeping the original English merges preserves English tokenization behavior, and retraining only the embedding and output layers adapts the new vocabulary entries without disturbing the pretrained transformer weights, which limits the risk of degrading English performance.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.