NCP-GENL • Practice Test 5
Free NCP-GENL practice test — 10 questions with explanations. Set 5. No signup required.
A team is preparing a mixed-language corpus for continued pretraining of an NVIDIA NeMo Megatron model. The corpus contains English, Japanese, and Arabic documents. Tokenizer analysis shows the current English-centric BPE vocabulary produces very long token sequences for Japanese and Arabic, inflating sequence length and compute cost. The team wants to reduce sequence length for non-English text without retraining the tokenizer from scratch and without degrading English performance. Which data preparation action best achieves this?
Choose an answer to begin — your selection is scored in the full session.
10 questions · instant feedback and full explanations after every question.