NCP-GENL LLM Architecture Practice Question
A team is pretraining a 13B-parameter decoder-only LLM on English text using byte-pair encoding with a 50,000-token vocabulary. They observe that the model produces fluent but repetitive continuations and that the average log-probability assigned to ground-truth tokens plateaus early. The training loss curve shows the model is underfitting rather than overfitting. Which architectural change is most likely to improve the model's capacity to capture long-range dependencies?
⚠ Common exam trap
The trap here is assuming that any architectural tweak that increases parameter count (such as adding attention heads) will resolve underfitting, when in fact head rebalancing at a fixed hidden dimension does not add representational capacity.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Increase the number of transformer layers and the hidden dimension proportionally.
Underfitting in a large decoder-only LLM indicates the model lacks sufficient capacity for the data distribution. Scaling both depth and width increases the parameter count and the number of nonlinear transformations applied to the residual stream, which improves the model's ability to capture long-range dependencies. The other options either rebalance existing capacity, change positional encoding without adding parameters, or reduce computation without improving fit.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Reduce the vocabulary size to 32,000 tokens to lower the softmax computation cost.
Why it's wrong here
Shrinking the vocabulary reduces the output projection and softmax cost but increases sequence length for the same text, and it does not address underfitting. A smaller vocabulary can actually make the modeling task harder because common words get split into more subword tokens. This change targets efficiency, not capacity, and would not lift the log-probability plateau.
- ✗
Increase the number of attention heads while keeping the hidden dimension fixed.
Why it's wrong here
Increasing head count while keeping the hidden dimension fixed reduces the per-head dimension, which can fragment representational capacity rather than add it. Underfitting in a 13B model is rarely resolved by rebalancing heads; the bottleneck is more likely depth or width in the residual stream. This change does not increase the total parameter count or the model's ability to model longer-range dependencies meaningfully.
- ✓
Increase the number of transformer layers and the hidden dimension proportionally.
Why this is correct
Underfitting at 13B parameters on English text signals insufficient model capacity. Scaling depth (more layers) and width (larger hidden dimension) together increases both the number of sequential transformation steps and the representational bandwidth, allowing the network to model longer-range dependencies and richer token distributions. This directly addresses the plateau in log-probability on ground-truth tokens.
- ✗
Replace the learned positional embeddings with sinusoidal absolute positional encodings.
Why it's wrong here
Swapping learned positional embeddings for sinusoidal ones changes how position is encoded but does not add model capacity. Sinusoidal encodings are parameter-free, so this would actually reduce the parameter count slightly. The underfitting symptom would persist because the underlying issue is insufficient depth and width, not the choice of positional encoding scheme.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.