NCP-GENL Fine-Tuning Practice Question
Exhibit
Config: { "gradient_accumulation_steps": 16, "per_device_train_batch_size": 1 }Refer to the exhibit. What is the effective batch size for this fine-tuning job?
⚠ Common exam trap
Candidates frequently add or subtract batch size parameters instead of multiplying per-device batch size by gradient accumulation steps to find the effective batch size.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
16
The effective batch size is calculated by multiplying the per-device batch size by the gradient accumulation steps. In this case, 1 * 16 equals 16. Gradient accumulation allows the model to simulate a larger batch size by running multiple forward and backward passes before performing a single optimizer step, which helps in achieving more stable gradient updates without increasing the immediate memory overhead for activations.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
1
Why it's wrong here
A batch size of 1 is only the per-device size. Because gradient accumulation is set to 16, the effective batch size is significantly larger. The gradient is accumulated over 16 steps before updating the model weights, which is a standard practice to improve training stability when memory is limited.
- ✗
15
Why it's wrong here
The math for effective batch size is multiplication, not addition or subtraction. Adding 1 and 16 incorrectly assumes an offset in the accumulation process. The effective batch size is defined as the product of the per-device size and the accumulation steps, which is critical for understanding the training dynamics.
- ✓
16
Why this is correct
The effective batch size is the product of the per_device_train_batch_size (1) and the gradient_accumulation_steps (16). This results in an effective batch size of 16, which helps in stabilizing training by providing a more representative gradient estimate over multiple mini-batches before weights are actually updated in the optimizer.
- ✗
17
Why it's wrong here
This result stems from incorrectly adding the two parameters. In deep learning frameworks, these parameters are multiplied to determine how many total samples contribute to each weight update. Calculating this value incorrectly can lead to misleading assumptions about the training throughput and the convergence behavior of the model during fine-tuning.
About these practice questions
Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.