NCA-GENL Core Machine Learning and AI Knowledge Practice Question
What is the primary function of Layer Normalization in a transformer architecture?
⚠ Common exam trap
Candidates often confuse Layer Normalization with Batch Normalization, incorrectly assuming that normalization across the batch dimension is the standard approach used in transformer architectures for stabilizing hidden states.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
To stabilize training by normalizing the inputs to each layer.
Layer normalization stabilizes the hidden state distributions across layers by normalizing inputs to have zero mean and unit variance. This prevents internal covariate shift and keeps activations within a stable range, allowing for faster convergence and deeper networks. In transformers, it is typically applied to the output of sub-layers, ensuring that subsequent computations remain numerically stable and well-conditioned for backpropagation.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
To increase the total parameter count of the model.
Why it's wrong here
Layer normalization adds a minimal number of learnable parameters (scale and shift). Its primary purpose is not to increase model capacity, but to normalize the inputs to layers. Adding parameters for the sake of complexity would likely hinder generalization rather than provide any structural benefit.
- ✓
To stabilize training by normalizing the inputs to each layer.
Why this is correct
By normalizing activations to have zero mean and unit variance, layer normalization reduces the dependence of a layer's output on the specific scaling of its input. This enables the use of higher learning rates and helps mitigate the exploding/vanishing gradient issues common in very deep neural networks.
- ✗
To replace the need for weight initialization techniques.
Why it's wrong here
Proper weight initialization (like Xavier or Kaiming initialization) is still required even with layer normalization. While normalization makes the network more robust to initialization, it cannot compensate for poor initial weight distributions, which can still lead to training failure or slow convergence in deep models.
- ✗
To compress the model to fit on smaller GPU devices.
Why it's wrong here
Layer normalization does not compress the model. Model compression is achieved through techniques such as pruning, quantization, or knowledge distillation. Layer normalization is a computational step that occurs during forward and backward passes, and it consumes additional memory rather than reducing the memory footprint of weights.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.