NCP-GENL Fine-Tuning Practice Question
A team is fine-tuning a 70B parameter model with NVIDIA NeMo using LoRA on eight H100 GPUs. They want to reduce GPU memory usage during training while preserving the base model's pretrained knowledge. Which two configuration changes should they apply? (Choose two.)
⚠ Common exam trap
The trap here is treating bfloat16 as a stability-only choice and forgetting that float32 doubles memory, which conflicts with the explicit memory-reduction objective.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Enable LoRA by setting adapter dimensions and target modules so only adapter weights receive gradients.
LoRA reduces trainable parameters and optimizer state while freezing the base model, and gradient checkpointing reduces activation memory by recomputing activations. Together they lower peak GPU memory during fine-tuning. Increasing micro batch size, using float32, or disabling tensor parallelism all increase memory usage and work against the stated goal.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the micro batch size to improve GPU utilization and reduce the number of optimizer steps.
Why it's wrong here
Increasing micro batch size raises activation memory per GPU, which is the opposite of reducing memory usage. While it can improve utilization, it often triggers out-of-memory errors on large models. The goal here is memory reduction, so larger micro batches are counterproductive unless paired with other memory-saving techniques.
- ✓
Enable LoRA by setting adapter dimensions and target modules so only adapter weights receive gradients.
Why this is correct
LoRA freezes the base model and trains low-rank adapter matrices, which drastically reduces the number of trainable parameters and optimizer state. This lowers memory usage during fine-tuning while preserving the pretrained weights. In NeMo, this is configured through the PEFT section with adapter dimensions and target modules, making it a direct memory-saving measure.
- ✗
Disable tensor parallelism so each GPU holds a full copy of the model weights.
Why it's wrong here
Disabling tensor parallelism would require each GPU to store the full model, which is impossible for a 70B parameter model on a single H100. Tensor parallelism is precisely what distributes model weights across GPUs. Removing it would increase per-GPU memory and likely cause immediate out-of-memory failures.
- ✗
Switch from bfloat16 to float32 precision to improve numerical stability during training.
Why it's wrong here
Using float32 doubles the memory footprint of weights, activations, and optimizer states compared to bfloat16. While it can improve numerical stability, it directly contradicts the objective of reducing GPU memory usage. On H100 GPUs, bfloat16 is typically preferred for large model fine-tuning to save memory and increase throughput.
- ✓
Enable gradient checkpointing to recompute activations during the backward pass instead of storing them.
Why this is correct
Gradient checkpointing trades compute for memory by storing only a subset of activations and recomputing the rest during backpropagation. This significantly reduces activation memory, which is often the largest consumer during fine-tuning of large models. It is a standard NeMo option and complements LoRA by further lowering peak memory.
About these practice questions
This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.