NCP-GENL Fine-Tuning Practice Question
Which THREE factors significantly influence the memory consumption during LLM fine-tuning? (Choose three)
⚠ Common exam trap
Candidates mistakenly select dataset size or sequence length as primary hardware memory drivers, overlooking how optimizer states, trainable parameters, and activation maps dictate fine-tuning VRAM consumption.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
The precision of the optimizer states
Memory consumption is driven by the model parameters, the optimizer states (which are often larger than the model weights), and the activations generated during the forward pass. Efficient management of these components is critical for scaling to larger models. By optimizing these three areas, practitioners can successfully fit large models into limited VRAM, ensuring stability throughout the fine-tuning process.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
The precision of the optimizer states
Why this is correct
The optimizer (e.g., AdamW) stores states for every trainable parameter. If using 32-bit precision, these states can occupy 8 bytes per parameter. Reducing this precision or utilizing paged techniques is essential for saving VRAM, as optimizer states are one of the largest contributors to memory exhaustion during the training cycle.
- ✓
The number of trainable parameters
Why this is correct
Every trainable parameter requires memory for its current value, its gradient, and associated optimizer states. Methods like LoRA reduce this by only training a small subset of parameters. Reducing the number of parameters directly correlates to lower memory usage, which is fundamental to successful fine-tuning on limited hardware.
- ✓
The activation maps stored for the backward pass
Why this is correct
Activations are stored during the forward pass to compute gradients during the backward pass. For long context windows or large batch sizes, the memory required to hold these activations can exceed the available GPU VRAM. Techniques like gradient checkpointing are used to mitigate this by recomputing activations instead of storing them.
- ✗
The number of CPU cores available for data loading
Why it's wrong here
CPU cores impact data pre-processing speed, not GPU VRAM consumption. While fast CPU processing ensures the GPU is constantly fed data, it does not alleviate the memory constraints on the GPU itself. Bottlenecks in data loading cause under-utilization of the GPU, but they do not cause OOM errors during training.
- ✗
The version of the operating system kernel
Why it's wrong here
The kernel version does not directly impact the VRAM footprint of an LLM. While drivers and CUDA versions are critical for compatibility and performance, they do not dictate how much memory a specific model architecture consumes. Fine-tuning memory is strictly determined by model configuration, optimizer settings, and the batch size.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.