NCP-GENL Fine-Tuning Practice Question
A team is fine-tuning a 70B-parameter LLM with NVIDIA NeMo on a multi-node cluster and wants to reduce the memory footprint per GPU without changing the model architecture. They are already using mixed precision and a reasonable micro-batch size. Which two techniques should they apply? (Choose two.)
⚠ Common exam trap
Watch out — candidates often confuse batch size reduction with memory reduction, when per-GPU memory during forward and backward passes is largely determined by model states and activations rather than the global batch size.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable ZeRO-style optimizer state partitioning across data-parallel ranks.
Partitioning optimizer states across data-parallel ranks and recomputing activations during the backward pass both reduce per-GPU memory without changing the model architecture. The first shrinks optimizer memory, and the second shrinks activation memory, which together address the dominant contributors for a 70B model. Changing attention heads, switching optimizers, or shrinking the global batch do not achieve the same effect.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Enable ZeRO-style optimizer state partitioning across data-parallel ranks.
Why this is correct
ZeRO-style partitioning shards optimizer states, and optionally gradients and parameters, across data-parallel ranks. Each GPU then holds only a fraction of the optimizer memory, which substantially reduces per-GPU footprint for a 70B model. This directly addresses the memory constraint without altering the architecture and is a standard technique in large-scale NeMo training.
- ✗
Switch the optimizer to SGD with momentum and remove weight decay.
Why it's wrong here
SGD with momentum uses less optimizer memory than Adam, but it changes the optimization dynamics and typically requires extensive retuning for LLM fine-tuning. It is not the standard memory-reduction technique for a 70B model and does not address activation memory. The scenario asks for techniques that reduce footprint without architectural changes, and this is a training-dynamics change rather than a memory strategy.
- ✓
Activate activation recomputation so intermediate activations are discarded and recomputed during the backward pass.
Why this is correct
Activation recomputation trades additional compute for lower activation memory by storing only a subset of activations and recomputing the rest during backpropagation. For a 70B model, activations can be a large share of memory, so this technique meaningfully reduces the per-GPU footprint while leaving the architecture unchanged. It is widely used alongside sharding strategies.
- ✗
Reduce the global batch size by a factor of eight and keep the micro-batch size the same.
Why it's wrong here
Reducing the global batch size lowers the number of micro-batches accumulated per step, but it does not reduce the memory required for the model, optimizer states, or activations at any given moment. The per-GPU footprint during forward and backward passes is largely unchanged. This would also alter convergence behavior and is not a memory-reduction technique.
- ✗
Increase the number of attention heads to spread the computation across more GPUs.
Why it's wrong here
Changing the number of attention heads modifies the model architecture and would invalidate the pre-trained weights. The scenario explicitly requires no architecture change, so this option is disqualified on that basis alone. Even if it were allowed, it would not reduce memory per GPU in the way the team needs.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.