Courseiva

NCA-GENL Core Machine Learning and AI Knowledge Practice Question

A data scientist is pretraining a 12-layer transformer encoder on a corpus of legal contracts. To prevent the model from simply copying each token to its output during masked language modeling, the team needs a strategy that forces the model to learn bidirectional context. Which masking approach should they apply?

⚠ Common exam trap

The trap here is assuming that any token replacement strategy will work, when the specific 80/10/10 split is designed to prevent the model from ignoring context and to align pretraining with fine-tuning.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Mask 15% of input tokens at random, replacing 80% with [MASK], 10% with a random token, and 10% with the original token.

The standard masked language modeling objective masks a small percentage of tokens and requires the model to reconstruct them from bidirectional context. Using mostly [MASK] with some random and unchanged tokens balances learning signal and reduces pretraining-finetuning mismatch. This is the correct way to force bidirectional understanding in a transformer encoder.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Mask the final token in every sequence and train the model to predict only that token.

    Why it's wrong here

    Masking only the final token turns the task into autoregressive next-token prediction, which is unidirectional and would not force the encoder to use bidirectional context. Legal contracts benefit from both left and right context, so this approach underutilizes the encoder's capacity and fails the stated goal.

  • ✗

    Replace 50% of tokens with [MASK] and leave the rest unchanged.

    Why it's wrong here

    Masking half the tokens removes too much signal and makes the task unnecessarily difficult, often slowing convergence and degrading representation quality. The standard rate is 15%, not 50%. This aggressive masking also increases the mismatch between pretraining and downstream fine-tuning where no masks appear.

  • ✓

    Mask 15% of input tokens at random, replacing 80% with [MASK], 10% with a random token, and 10% with the original token.

    Why this is correct

    This is the standard BERT-style masking recipe. Replacing 80% with [MASK] forces the model to predict from context, while the 10% random and 10% unchanged tokens reduce pretrain-finetune mismatch and discourage the model from ignoring non-masked tokens. It directly supports bidirectional learning of legal contract language.

  • ✗

    Randomly shuffle the order of tokens within each sequence before feeding them to the model.

    Why it's wrong here

    Shuffling tokens destroys the sequential structure that the model must learn. It would prevent the encoder from learning meaningful syntactic and semantic relationships in legal contracts. This is not a masking strategy and does not create a prediction task that encourages bidirectional understanding.

About these practice questions

This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.