NCP-GENL LLM Architecture Practice Question
An engineer is analyzing why a decoder-only LLM with 32,000-token context length fails to answer questions that require information from the beginning of a long document when the answer is near the end. The model was trained with standard causal attention. Which two architectural or training factors are most likely contributing to this failure? (Choose two.)
⚠ Common exam trap
The trap here is blaming the causal attention mask for limiting long-context reasoning, when causal masking is a defining feature of decoder-only models and does not prevent attending to earlier tokens.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The positional encoding scheme may not generalize well beyond the sequence lengths seen during pretraining.
Long-context retrieval failures in decoder-only LLMs typically stem from two sources: pretraining on sequences much shorter than the target context, which prevents the model from learning to attend across the full window, and positional encoding schemes that do not generalize beyond training lengths. Together these cause the model to underuse information at the start of a long document. The causal mask, feed-forward network, and vocabulary size are not primary causes of this behavior.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The feed-forward network in each layer compresses the hidden state, discarding information from early tokens.
Why it's wrong here
The feed-forward network operates on each position independently and does not selectively discard early-token information. While it transforms representations, the residual stream preserves information across layers. The bottleneck for long-context retrieval is attention and positional encoding, not feed-forward compression, which is not a recognized cause of this failure mode.
- ✗
The vocabulary size is too small to represent the document's domain-specific terms, causing tokenization errors.
Why it's wrong here
Vocabulary size affects tokenization granularity but does not explain why information at the beginning of a long context is ignored. Even with a large vocabulary, the model could still fail to attend across 32,000 tokens if it was not trained on such lengths or if positional encoding does not generalize. This factor is unrelated to the described long-context retrieval failure.
- ✓
The positional encoding scheme may not generalize well beyond the sequence lengths seen during pretraining.
Why this is correct
Many positional encoding schemes, especially learned absolute embeddings or RoPE without scaling, degrade when sequences exceed the lengths seen during training. The model cannot reliably distinguish or weight positions far beyond its training range, so attention to early tokens becomes imprecise. This directly contributes to the failure to retrieve information from the beginning of a long document.
- ✗
Causal attention masks prevent each token from attending to future tokens, which limits bidirectional reasoning over the document.
Why it's wrong here
Causal masking is inherent to decoder-only models and is not a defect; it ensures autoregressive generation. The question is about attending to earlier tokens, which causal attention does allow. The failure to use early information stems from training distribution and positional encoding behavior, not from the causal mask itself, which is functioning as intended.
- ✓
The model was pretrained primarily on sequences much shorter than 32,000 tokens, so it did not learn to attend across the full context.
Why this is correct
If pretraining used sequences far shorter than the target context length, the attention patterns never learned to connect distant tokens. Extending context at inference without training on long sequences leads to poor utilization of early tokens. This is a well-documented cause of the lost-in-the-middle effect, where information at the extremes of a long context is underused.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.