Courseiva
Fine-Tuning →hardMultiple Choice

NCP-GENL Fine-Tuning Practice Question

An engineer is fine-tuning a 13B parameter model with NVIDIA NeMo using tensor parallelism across four GPUs. After resuming from a checkpoint, training loss spikes and then diverges. The checkpoint was saved with a different tensor parallel size than the current run. What is the most likely cause of the divergence?

⚠ Common exam trap

The trap here is overlooking that checkpoint compatibility in NeMo depends not only on model architecture but also on the parallelism configuration used when the checkpoint was written.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

The optimizer state and model shards were partitioned differently, so the restored weights do not match the current tensor parallel layout.

Tensor parallelism partitions model weights across GPUs, and the partition layout depends on the tensor parallel size. Loading a checkpoint saved with a different tensor parallel degree without resharding places weights incorrectly, causing loss spikes and divergence. Data order, scheduler restart, and gradient accumulation changes do not explain the immediate failure tied to a changed parallel layout.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    The optimizer state and model shards were partitioned differently, so the restored weights do not match the current tensor parallel layout.

    Why this is correct

    Tensor parallelism splits weight matrices across GPUs, and the sharding layout depends on tensor parallel size. A checkpoint saved with one tensor parallel degree cannot be directly loaded into a run with a different degree without resharding. The mismatched partitioning causes incorrect weight placement, leading to loss spikes and divergence when training resumes.

  • ✗

    The learning rate scheduler restarted from step zero and applied a large learning rate to all parameters.

    Why it's wrong here

    A scheduler restart could cause instability, but NeMo typically saves and restores scheduler state with the checkpoint. Even if the scheduler restarted, the effect would be gradual rather than the immediate divergence seen here. The stronger signal is the change in tensor parallel size, which directly affects how weights are partitioned.

  • ✗

    Gradient accumulation steps were reduced, effectively increasing the global batch size beyond the original configuration.

    Why it's wrong here

    Changing gradient accumulation alters the effective batch size and could affect convergence, but it would not produce the abrupt divergence described. The scenario explicitly notes a different tensor parallel size between checkpoint and current run, which is a direct cause of weight layout mismatch. Batch size changes are unrelated to sharding compatibility.

  • ✗

    The data loader random seed was not preserved, causing the model to see samples in a different order.

    Why it's wrong here

    A different data order alone would not cause immediate divergence after resuming; it would only change the sequence of batches. Loss might fluctuate slightly, but a consistent spike and divergence points to a structural mismatch in model or optimizer state. Data ordering is a secondary concern compared to incorrect weight sharding.

About these practice questions

Courseiva writes every NCP-GENL question from scratch — 352 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.