NCP-GENL Fine-Tuning Practice Question
An enterprise team is preparing a supervised fine-tuning job in NVIDIA NeMo for a 20B LLM. They want to reduce GPU memory consumption during training without changing the model architecture or the dataset. Which two configuration changes should they apply? (Choose two.)
⚠ Common exam trap
The trap here is treating data-level changes like truncation or augmentation as memory optimizations, when the constraints require configuration-level changes that leave architecture and dataset intact.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable gradient checkpointing to recompute activations during the backward pass.
Gradient checkpointing and bfloat16 mixed precision both reduce memory during training without modifying model architecture or dataset. Checkpointing lowers stored activation memory by recomputation, while bfloat16 halves tensor memory and is supported natively in NeMo. Architectural changes, data expansion, and sequence truncation either violate the constraints or fail to address memory consumption per training step.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Expand the training dataset with additional synthetic examples.
Why it's wrong here
Adding synthetic data changes the dataset and increases training time without reducing memory consumption per step. The scenario requires keeping the dataset unchanged and lowering GPU memory. Data augmentation is a quality or coverage strategy, not a memory optimization, so it does not address the stated objective.
- ✗
Reduce the maximum sequence length by truncating all samples to 64 tokens.
Why it's wrong here
Truncating sequences lowers memory but alters the dataset content and may remove target responses, violating the constraint that the dataset remain unchanged. It also risks degrading model quality on longer inputs. The scenario asks for configuration changes that reduce memory without touching data or architecture, which truncation does not satisfy.
- ✓
Enable gradient checkpointing to recompute activations during the backward pass.
Why this is correct
Gradient checkpointing stores only selected activations and recomputes the rest during backpropagation, trading additional compute for substantially lower activation memory. It does not alter the model architecture or dataset, so it fits the stated constraints. This makes it an effective lever for fitting larger models or batch sizes into the same GPU memory budget during NeMo fine-tuning runs.
- ✗
Increase the number of attention heads in each transformer block.
Why it's wrong here
Adding attention heads changes the model architecture and increases parameter count and memory usage. The scenario explicitly forbids architectural changes and seeks memory reduction, so this option works against both requirements. Attention head configuration is a design-time decision, not a runtime memory optimization for an existing fine-tuning job.
- ✓
Use mixed precision with bfloat16 for forward and backward passes.
Why this is correct
bfloat16 halves the memory footprint of activations, gradients, and optimizer-related tensors compared with float32 while retaining a wide dynamic range suitable for training. It does not change model architecture or data, satisfying the constraints. In NeMo, enabling bfloat16 mixed precision is a standard way to reduce memory and increase throughput on supported NVIDIA GPUs.
About these practice questions
One of 352 original NCP-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.