NCA-GENL Core Machine Learning and AI Knowledge Practice Question
In the context of transformer models, what is the purpose of the 'Attention Mask' during the training process?
⚠ Common exam trap
Candidates frequently mistake the attention mask for padding management used to handle variable sequence lengths, overlooking its critical causal role in preventing autoregressive models from seeing future tokens.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
It prevents the model from attending to future tokens in autoregressive models.
The attention mask is essential for transformer architectures to handle sequences of varying lengths and to enforce causal constraints. In tasks like sequence generation, it prevents the model from 'peeking' at future tokens, ensuring that the prediction at each position depends only on preceding tokens. This mechanism is critical for maintaining the autoregressive nature of models like GPT and ensuring correct model training.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
It serves to reduce the number of parameters in the self-attention layer.
Why it's wrong here
The attention mask does not change the number of weights or parameters in the model. It is a binary or additive mask applied to the attention scores before the softmax operation, acting as a filter for specific positions in the input sequence, not a structural parameter optimization.
- ✓
It prevents the model from attending to future tokens in autoregressive models.
Why this is correct
In autoregressive LLMs, the model must predict the next token based only on previous ones. The attention mask sets the attention scores for future positions to negative infinity before softmax, effectively nullifying their influence. This ensures the model learns causal dependencies during its training phase.
- ✗
It optimizes the data movement between the GPU's L1 and L2 cache.
Why it's wrong here
The attention mask is a functional component of the model architecture, not a low-level hardware optimization. While efficient implementation of masking can affect performance, its primary role is to logically constrain the attention mechanism rather than to manage hardware-level cache resources directly.
- ✗
It performs weight pruning to compress the model size after training.
Why it's wrong here
Weight pruning is a separate post-training technique used to reduce model size by removing unimportant weights. The attention mask is used during the training loop to control the information flow within the self-attention mechanism, and it is unrelated to the structural pruning of model weights.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.