NCP-GENL LLM Architecture Practice Question
A team is training a large language model and wants to reduce the memory used by the optimizer without changing the model architecture. They are using Adam and notice that optimizer state consumes more GPU memory than the model weights. Which technique should they apply to reduce optimizer memory while keeping the model architecture unchanged?
⚠ Common exam trap
The trap here is assuming that larger batches or architectural tweaks solve optimizer memory, when the moments themselves are the cost.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Switch to a memory-efficient optimizer such as 8-bit Adam or Adafactor that stores reduced-precision or factored optimizer state.
Adam maintains two full-precision moment estimates for every trainable parameter, so optimizer state can be several times the size of the model weights. Memory-efficient optimizers such as 8-bit Adam or Adafactor reduce this footprint by quantizing or factoring the state, directly lowering the dominant memory cost without changing the model architecture.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Replace the feed-forward activation with a gated variant to reduce activation memory.
Why it's wrong here
Changing the activation can affect activation memory and compute, but it modifies the architecture and does not address the optimizer's stored moments. The scenario explicitly requires keeping the architecture unchanged, so this does not meet the requirement.
- ✓
Switch to a memory-efficient optimizer such as 8-bit Adam or Adafactor that stores reduced-precision or factored optimizer state.
Why this is correct
8-bit Adam quantizes the first and second moment estimates to 8-bit, and Adafactor factors the second moment, both cutting optimizer memory substantially. These keep the model architecture and parameter count intact while lowering the dominant memory cost during training.
- ✗
Reduce the number of attention heads while keeping the hidden size constant.
Why it's wrong here
Altering the head count changes the architecture and can degrade model quality. It also does not shrink the optimizer state, which scales with the number of trainable parameters, not with how those parameters are organized into heads.
- ✗
Increase the batch size proportionally to the number of GPUs so optimizer states are shared across more tokens.
Why it's wrong here
Increasing batch size changes the optimization dynamics and memory profile but does not reduce the per-parameter optimizer state. Adam still maintains two moments per parameter, so memory per parameter is unchanged regardless of how many tokens are processed per step.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.