Courseiva
Fine-Tuning →mediumMultiple Choice

NCP-GENL Fine-Tuning Practice Question

A team is fine-tuning a 13B-parameter Llama model with NVIDIA NeMo on a node of eight A100 80GB GPUs. They want the optimizer state to be partitioned across data-parallel ranks so that per-GPU memory drops, while keeping the model replicas synchronized. Which distributed strategy should they select in the NeMo training configuration?

⚠ Common exam trap

The trap here is assuming that any multi-GPU parallelism mode reduces optimizer memory, when only data-parallel optimizer sharding actually partitions the Adam state.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

ZeRO Stage 1 via the distributed optimizer in Megatron Core

Sharding optimizer state across data-parallel ranks is exactly what the distributed optimizer in Megatron Core does, providing ZeRO Stage 1 behavior with full model replicas per rank. Tensor and pipeline parallelism change how parameters and layers are split, and recomputation targets activations. Only optimizer-state partitioning reduces the Adam moments and master weights each GPU must hold without adding model-parallel communication.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Tensor Parallelism (TP) with tensor_model_parallel_size set to 8

    Why it's wrong here

    Tensor parallelism shards individual weight matrices and their activations across GPUs inside each layer, which reduces per-GPU parameter and activation memory but requires heavy all-reduce communication on every layer. It does not partition the optimizer state across data-parallel ranks, and setting it to 8 for a 13B model creates needless communication overhead while leaving the optimizer state replicated.

  • ✗

    Pipeline Parallelism (PP) with pipeline_model_parallel_size set to 8

    Why it's wrong here

    Pipeline parallelism splits the transformer layers into stages spread across GPUs, which lowers the parameter memory held by each rank but does not shard the optimizer state. It also introduces pipeline bubbles and requires micro-batch scheduling, so it addresses a different memory pressure than the optimizer state the team wants to partition.

  • ✗

    Activation recomputation with recompute_granularity set to full

    Why it's wrong here

    Activation recomputation trades extra compute during the backward pass for lower stored activation memory; it does not touch the optimizer state at all. Enabling it would not partition Adam moments or master weights across data-parallel ranks, so the memory reduction the team is targeting would not be achieved through this setting alone.

  • ✓

    ZeRO Stage 1 via the distributed optimizer in Megatron Core

    Why this is correct

    The distributed optimizer in Megatron Core implements ZeRO Stage 1 semantics by sharding the optimizer states, such as Adam first and second moments plus the master weights, across the data-parallel ranks. Each rank keeps a full model replica for the forward and backward pass, so no extra model-parallel communication is introduced, and per-GPU memory falls in proportion to the data-parallel size.

About these practice questions

This NCP-GENL question is part of Courseiva's 352-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCP-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCP-GENL exam.