Courseiva
LLM Architecture →mediumMultiple Choice

NCP-GENL LLM Architecture Practice Question

What is the architectural role of Layer Normalization in a Transformer, and where is it typically placed to ensure stable training?

⚠ Common exam trap

Candidates often confuse Pre-LN with Post-LN. They may remember Layer Normalization exists but fail to recognize that Pre-LN is the modern standard for training stability in deep Transformers.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

It is placed before each sub-layer (Pre-LN) to improve training stability.

Layer normalization stabilizes the hidden states by ensuring their mean and variance remain within a controlled range throughout the network layers. In modern LLMs, placing it before the attention and FFN blocks (Pre-LN) is preferred over the original Post-LN. This prevents gradient explosion early in training, allowing for higher learning rates and more reliable convergence for deep models during their pre-training phase.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    It normalizes the entire batch to reduce training time.

    Why it's wrong here

    Batch normalization is rarely used in NLP models because it depends on the statistics of the batch, which can vary significantly across sequence lengths. Layer normalization is preferred because it normalizes features within each individual token, maintaining consistency across sequences and avoiding dependencies on the broader training batch statistics.

  • ✓

    It is placed before each sub-layer (Pre-LN) to improve training stability.

    Why this is correct

    Pre-LN configurations move the normalization layer inside the residual branch, which creates a more stable gradient flow. This architectural choice is standard for modern LLMs as it significantly reduces the risk of loss divergence during the early stages of training compared to the older Post-LN approach.

  • ✗

    It is used to perform dimensionality reduction on input embeddings.

    Why it's wrong here

    Layer normalization performs statistical standardization, not dimensionality reduction. Reducing the dimension would imply discarding information or projecting it into a smaller space, whereas normalization preserves the original feature count while scaling the magnitude of the values to ensure the network remains stable during backpropagation.

  • ✗

    It replaces the need for residual connections in the model.

    Why it's wrong here

    Residual connections and normalization are complementary; they work together to enable the training of deep models. Residual connections allow gradients to flow through the network without disappearing, while normalization ensures these signals stay within a range where the model's non-linearities can effectively function and learn complex patterns.

About these practice questions

One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.