Courseiva
LLM Architecture →mediumMultiple Choice

NCP-GENL LLM Architecture Practice Question

A research team is pretraining a decoder-only LLM and observes that gradient magnitudes in the earliest layers are extremely small while later layers train normally, causing slow convergence. They are using post-layer normalization. Which architectural change is most likely to improve gradient flow to the early layers?

⚠ Common exam trap

The trap here is treating vanishing early-layer gradients as a capacity or width problem and adjusting head counts, when the root cause is where normalization sits relative to the residual branch.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Move layer normalization to before each sublayer (pre-LN) and add a final normalization before the output projection

Post-LN places normalization inside the residual branch, so gradients must pass through normalization at every layer, attenuating them in deep stacks. Pre-LN normalizes the sublayer input and leaves the residual stream unnormalized, giving a direct gradient highway from the loss to early parameters. That structural change, plus a final normalization before the output head, is the recognized fix for the described symptom.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Move layer normalization to before each sublayer (pre-LN) and add a final normalization before the output projection

    Why this is correct

    Pre-LN applies normalization on the sublayer input, creating a clean residual path that carries gradients directly from the loss to early layers. This is the standard remedy for vanishing gradients in deep Transformers and typically removes the need for learning-rate warmup. Adding a final normalization stabilizes the output scale before the vocabulary projection.

  • ✗

    Replace layer normalization with batch normalization across the sequence dimension

    Why it's wrong here

    Batch normalization statistics depend on other examples in the batch and behave poorly with variable-length sequences and autoregressive decoding, where batch composition changes per step. It also does not specifically fix residual gradient attenuation the way pre-LN does, and it introduces train/inference mismatch that complicates generation.

  • ✗

    Remove residual connections to shorten the gradient path

    Why it's wrong here

    Residual connections are precisely what allow gradients to bypass sublayers and reach early layers. Removing them lengthens the effective gradient path and worsens vanishing gradients, making the described problem more severe rather than improving early-layer learning in a deep network.

  • ✗

    Increase the number of attention heads while keeping head dimension constant

    Why it's wrong here

    Adding attention heads changes the model's representational capacity and parameter count but does not address how gradients propagate through the residual stream and normalization layers. The early-layer gradient attenuation described is a normalization-placement issue, not a head-count issue.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.