NCP-GENL • Practice Exam 3 — 20 Questions
Free NCP-GENL practice exam 3 — 20 questions with explanations. No signup required.
A team is preparing a mixed-language corpus for continued pretraining of an NVIDIA NeMo Megatron model. The corpus contains English, Japanese, and Arabic documents. Tokenizer analysis shows the current English-centric BPE vocabulary produces very long token sequences for Japanese and Arabic, inflating sequence length and compute cost. The team wants to reduce sequence length for non-English text without retraining the tokenizer from scratch and without degrading English performance. Which data preparation action best achieves this?
Choose an answer to begin — your selection is scored in the full session.
20 questions · instant feedback and full explanations after every question.