NCP-GENL LLM Architecture Practice Question
What is the architectural role of Layer Normalization in a Transformer, and where is it typically placed to ensure stable training?
⚠ Common exam trap
Candidates often confuse Pre-LN with Post-LN. They may remember Layer Normalization exists but fail to recognize that Pre-LN is the modern standard for training stability in deep Transformers.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
It is placed before each sub-layer (Pre-LN) to improve training stability.
Layer normalization stabilizes the hidden states by ensuring their mean and variance remain within a controlled range throughout the network layers. In modern LLMs, placing it before the attention and FFN blocks (Pre-LN) is preferred over the original Post-LN. This prevents gradient explosion early in training, allowing for higher learning rates and more reliable convergence for deep models during their pre-training phase.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
It normalizes the entire batch to reduce training time.
Why it's wrong here
Batch normalization is rarely used in NLP models because it depends on the statistics of the batch, which can vary significantly across sequence lengths. Layer normalization is preferred because it normalizes features within each individual token, maintaining consistency across sequences and avoiding dependencies on the broader training batch statistics.
- ✓
It is placed before each sub-layer (Pre-LN) to improve training stability.
Why this is correct
Pre-LN configurations move the normalization layer inside the residual branch, which creates a more stable gradient flow. This architectural choice is standard for modern LLMs as it significantly reduces the risk of loss divergence during the early stages of training compared to the older Post-LN approach.
- ✗
It is used to perform dimensionality reduction on input embeddings.
Why it's wrong here
Layer normalization performs statistical standardization, not dimensionality reduction. Reducing the dimension would imply discarding information or projecting it into a smaller space, whereas normalization preserves the original feature count while scaling the magnitude of the values to ensure the network remains stable during backpropagation.
- ✗
It replaces the need for residual connections in the model.
Why it's wrong here
Residual connections and normalization are complementary; they work together to enable the training of deep models. Residual connections allow gradients to flow through the network without disappearing, while normalization ensures these signals stay within a range where the model's non-linearities can effectively function and learn complex patterns.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.