NCP-GENL LLM Architecture Practice Question
An engineer is reviewing the attention implementation of a decoder-only LLM used for chat. During inference with a KV cache, generated tokens must not attend to future positions. Which mechanism enforces this constraint inside scaled dot-product attention?
⚠ Common exam trap
Watch out — candidates often confuse numerical stabilization techniques like score scaling or QK-norm with the masking mechanism that actually enforces causal visibility.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A causal mask that sets attention scores for future positions to negative infinity before the softmax.
Causal masking is the standard mechanism that makes decoder-only attention autoregressive. By adding negative infinity to scores of future positions before softmax, those positions receive zero weight, so each token can only attend to itself and prior tokens. Other listed techniques affect numerical stability or positional encoding but do not enforce the temporal constraint.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Dropping the scaling factor 1/sqrt(d_k) so that scores decay for distant tokens.
Why it's wrong here
Removing the scaling factor changes the magnitude of dot products and can destabilize softmax, but it does not prevent attention to future tokens. Distant future positions could still receive high scores if their key vectors align with the query. Masking is the only reliable mechanism for enforcing causal attention; scaling is about numerical conditioning, not causality.
- ✓
A causal mask that sets attention scores for future positions to negative infinity before the softmax.
Why this is correct
A causal (lower-triangular) mask adds negative infinity to scores at positions beyond the current token, so after softmax those weights become zero. This guarantees each position only attends to itself and earlier tokens, which is exactly the autoregressive constraint required for decoder-only generation, and it works identically whether or not a KV cache is used.
- ✗
Applying LayerNorm to the query and key projections before computing attention scores.
Why it's wrong here
Normalizing queries and keys (as in QK-norm) improves training stability and can prevent attention logit explosion, but it has no directional effect. Every position still computes scores against every other position, including future ones. LayerNorm cannot impose the lower-triangular structure needed for autoregressive decoding, so it does not enforce causality.
- ✗
Shifting the position IDs of cached keys so they appear before the current token.
Why it's wrong here
Position IDs affect how positional information is encoded, not which tokens are visible. Reassigning IDs to cached keys would corrupt positional semantics without blocking attention to genuinely future tokens during training-time parallel computation. Causality in attention comes from masking the score matrix, not from manipulating position indices.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.