Courseiva
LLM Architecture →easyMultiple Choice

NCP-GENL LLM Architecture Practice Question

What is the function of the 'Masked' component in a Decoder-only Transformer's self-attention during training?

⚠ Common exam trap

Candidates often think the mask is for ignoring padding tokens. While padding masks exist, the primary 'Masked' component in decoder-only training is strictly for preventing look-ahead at future tokens.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

To prevent the model from attending to future tokens in the sequence.

The mask ensures that each token can only attend to itself and the tokens that precede it in the sequence. This autoregressive property is essential for training decoder models, as it prevents the model from 'cheating' by looking at future tokens. This ensures that the model learns to predict the next word based solely on the context that would be available during actual inference.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    To filter out stop words from the input text.

    Why it's wrong here

    Masking in Transformers is about temporal causality, not linguistic content filtering. Stop words are handled during the preprocessing or tokenization stages. The attention mask is a structural mechanism used to enforce the autoregressive nature of the generation process, which is unrelated to the semantics of the words being processed.

  • ✓

    To prevent the model from attending to future tokens in the sequence.

    Why this is correct

    The causal mask is a triangular matrix applied to the attention scores that sets values for future tokens to negative infinity. This ensures that when the softmax is applied, the weights for these positions become zero, effectively hiding the future context from the model during training and ensuring autoregressive integrity.

  • ✗

    To increase the randomness of the model's predictions.

    Why it's wrong here

    Randomness in model predictions is controlled by parameters like temperature, top-k, or top-p sampling during inference, not by the attention mask. The attention mask is a deterministic architectural constraint designed to enforce sequence ordering; it does not introduce or manage stochasticity in the model's output generation process.

  • ✗

    To reduce the computation of the attention mechanism by 50%.

    Why it's wrong here

    While masking does effectively compute fewer values, it is not an efficiency-focused optimization; it is a correctness-focused architectural requirement. The computation reduction is a side effect of ignoring future positions, but the primary goal is ensuring that the model does not violate the causal dependencies of language generation.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.