NCP-GENL LLM Architecture Practice Question
A team is pre-training a 13B-parameter decoder-only LLM on a cluster of NVIDIA GPUs. They observe that gradient norms spike sharply during the first few hundred steps, destabilizing training. They want to keep the standard post-layer-normalization placement but stabilize early optimization. Which architectural technique should they apply?
⚠ Common exam trap
The trap here is assuming that learning rate warmup alone fixes post-LN instability when the underlying issue is depth-dependent residual variance growth.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Initialize all residual output projections with a small scaling factor such as 1/sqrt(2N).
Post-layer-normalization Transformers suffer from residual stream variance that grows with depth, producing large gradients early in training. Scaling residual branch outputs at initialization by a depth-dependent factor keeps variance bounded, directly stabilizing optimization without changing normalization placement or the optimizer schedule. Optimizer tweaks and activation swaps do not address the architectural root cause.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Initialize all residual output projections with a small scaling factor such as 1/sqrt(2N).
Why this is correct
Scaling residual branch outputs at initialization by a depth-dependent factor keeps the variance of the residual stream roughly constant across layers. This directly tames the growth that causes exploding gradients in post-LN Transformers, allowing stable early training without moving normalization. It is a recognized architectural fix used in deep Transformer initialization schemes such as those in Megatron-style training.
- ✗
Apply learning rate warmup over the first several thousand steps.
Why it's wrong here
Warmup mitigates large early updates and is standard practice, but it does not address the root cause of post-LN instability: the residual stream magnitude grows with depth, so gradients through deep blocks remain ill-conditioned. Spikes can still occur after warmup completes, and the question asks for an architectural technique, not an optimizer schedule change.
- ✗
Insert RMSNorm immediately after the token embedding layer only.
Why it's wrong here
Adding RMSNorm only after embeddings normalizes the input scale but does nothing to constrain the growing activations inside each Transformer block, which is where the gradient spikes originate. Early training instability with post-LN placement comes from unnormalized residual stream growth across many layers, so a single normalization at the input cannot prevent the exploding gradients observed during the first few hundred steps.
- ✗
Replace GELU activations in the feed-forward network with ReLU.
Why it's wrong here
Changing the activation function alters nonlinearity shape but does not control residual stream variance, which is the mechanism behind the observed gradient spikes. ReLU can even worsen dead-neuron issues at high learning rates. The instability is a depth-scaling problem, so swapping activations leaves the fundamental cause untouched and may reduce model quality.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.