NCP-GENL LLM Architecture Practice Question
A team is pretraining a decoder-only Transformer LLM on a large corpus of code and natural language. They observe that the model's training loss decreases smoothly, but during generation it sometimes produces degenerate repetition, and attention entropy on long sequences collapses. They suspect the issue is related to the positional encoding scheme. Which architectural change is most likely to mitigate the attention entropy collapse while preserving the model's ability to generalize to sequences longer than those seen during pretraining?
⚠ Common exam trap
The trap here is assuming that any positional encoding change, such as switching to sinusoidal or adding normalization, will fix length generalization, when only a relative-position scheme applied inside attention addresses the described entropy collapse.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Replace learned absolute positional embeddings with Rotary Positional Embeddings (RoPE) applied to queries and keys in every self-attention layer.
The symptoms point to a positional encoding that does not generalize beyond the pretraining sequence length. Rotary Positional Embeddings inject relative position directly into the attention computation by rotating queries and keys, so attention scores depend on relative offsets rather than absolute indices. This preserves attention entropy on longer sequences and reduces degenerate repetition, making it the appropriate architectural change for this scenario.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the number of attention heads while keeping the total hidden dimension constant, without changing the positional encoding.
Why it's wrong here
Adding more heads with a fixed hidden dimension reduces the per-head dimension, which can change representational capacity but does not alter how position is encoded. The root cause described is positional extrapolation and entropy collapse, not head count. This change is unlikely to fix repetition on long sequences and may even degrade per-head expressiveness.
- ✗
Apply Layer Normalization before the self-attention sublayer instead of after it, leaving the positional encoding unchanged.
Why it's wrong here
Pre-LN versus post-LN affects gradient flow and training stability, but it does not change the positional encoding mechanism. The scenario attributes the repetition and attention entropy collapse to positional handling during long-sequence generation. Reordering normalization alone will not provide relative-position generalization or prevent entropy collapse when extrapolating beyond the pretraining length.
- ✓
Replace learned absolute positional embeddings with Rotary Positional Embeddings (RoPE) applied to queries and keys in every self-attention layer.
Why this is correct
RoPE encodes relative position by rotating query and key vectors by an angle proportional to their absolute position. Because attention scores depend only on the relative rotation between queries and keys, the model naturally generalizes to longer contexts and avoids the entropy collapse seen when learned absolute embeddings are extrapolated beyond their trained range. This matches the observed repetition and entropy issues.
- ✗
Add a sinusoidal positional encoding to the input embeddings and freeze those embeddings during training.
Why it's wrong here
Sinusoidal encodings are fixed and added only at the input, so they do not provide the relative-position inductive bias inside each attention layer. Freezing input embeddings further limits the model's ability to adapt position representations to the code and language mixture. This does not address entropy collapse in deep attention layers or improve length generalization in the way the scenario requires.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.