NCA-GENL Experimentation Practice Question
Exhibit
log_output: [Rank 0] Model shards loaded: 8/8. Memory usage per GPU: 78GB. Peak VRAM limit: 80GB. Training step 500: Latency 450ms. Step 501: Latency 460ms. Step 502: Latency 890ms.
Refer to the exhibit. An engineer is monitoring a large model training job. Based on the sudden latency spike at step 502, what is the most likely cause during the experimentation phase?
⚠ Common exam trap
Examinees often attribute sudden latency spikes to complex network bottlenecks or gradient explosions, overlooking routine systems maintenance tasks like periodic model checkpointing.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
A periodic checkpointing operation triggered at step 502.
The sudden latency jump suggests a periodic operation like checkpointing, data logging, or a hardware-level thermal throttling event. In LLM training, frequent checkpoints or massive synchronization steps are primary suspects for sudden, brief stalls. Identifying these spikes early allows developers to tune checkpoint frequency or optimize I/O paths, ensuring that experiments maintain consistent performance and avoid unnecessary overhead during training.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
The model has reached a local minimum in the loss function.
Why it's wrong here
Reaching a local minimum affects the convergence trajectory of the model weights but does not manifest as a sudden, massive increase in step latency. Latency spikes are typically related to system-level operations or resource contention rather than the mathematical optimization state of the neural network during training.
- ✓
A periodic checkpointing operation triggered at step 502.
Why this is correct
Periodic I/O operations such as saving model weights to disk or synchronizing distributed state often cause transient latency spikes. Because the latency doubled at step 502, it is highly indicative of a blocking I/O operation or a synchronization barrier that is occurring at regular training intervals.
- ✗
The learning rate scheduler reduced the step size.
Why it's wrong here
Learning rate schedulers modify the hyperparameters for the optimizer but do not introduce overhead that would cause a major spike in training step latency. The computational complexity of weight updates remains constant regardless of the learning rate value currently being applied to the model parameters.
- ✗
The model architecture was automatically reconfigured.
Why it's wrong here
Automatic structural reconfiguration of a model during training is not a standard behavior for LLM training frameworks. Changes to the model architecture must be explicitly defined and triggered by the researcher; such a process would likely cause a program restart or crash rather than a latency spike.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.