NCP-GENL LLM Architecture Practice Question
A developer is inspecting a decoder-only Transformer and notices that during training, the model attends to future tokens in the sequence, causing the loss to drop unrealistically fast but generation to be incoherent. Which architectural mechanism is missing or misconfigured?
⚠ Common exam trap
Watch out — candidates often confuse positional encoding with causal masking, assuming that relative position information alone prevents a decoder from attending to future tokens.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Causal attention mask in the self-attention layers
Autoregressive language models require a causal mask that sets attention scores to negative infinity for positions after the current token. When that mask is absent, the training objective becomes trivial because the target token is visible in the input context, producing the fast loss drop and incoherent generation. Positional encodings, dropout, and normalization placement do not enforce temporal directionality.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Pre-layer normalization placement before attention
Why it's wrong here
Pre-LN versus post-LN changes gradient flow and training stability, but neither placement restricts which tokens a position can attend to. A model with pre-LN and no causal mask still leaks future information, so normalization placement is unrelated to the observed training-versus-generation discrepancy.
- ✓
Causal attention mask in the self-attention layers
Why this is correct
In a decoder-only model, self-attention must be masked so each position can only attend to itself and earlier positions. Without this causal mask, the model sees future tokens during training, trivially minimizing next-token loss while learning nothing useful for autoregressive generation, which exactly matches the described symptom.
- ✗
Dropout regularization on attention weights
Why it's wrong here
Dropout randomly zeros attention probabilities during training to reduce overfitting. It has no role in enforcing directionality of attention. Removing or misconfiguring dropout would change generalization behavior, not allow the model to see future tokens, so it cannot explain the unrealistic loss curve described.
- ✗
Rotary positional embedding frequency base
Why it's wrong here
RoPE encodes relative position through rotation of query and key vectors and affects how distance is represented, but it does not by itself prevent attending to future tokens. Even with a perfect RoPE base, the model would still leak future information if no causal mask is applied in the attention score matrix.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.