NCA-GENL Experimentation Practice Question
A researcher is running an ablation study in which they vary the number of attention heads in a NeMo Megatron GPT model while holding parameter count, dataset, and learning rate fixed. After the first run, they change tensor parallel size and pipeline parallel size to fit larger variants on the available GPUs. A colleague argues this invalidates the comparison. Which statement best explains the scientific concern?
⚠ Common exam trap
The trap here is assuming that any configuration change between ablation runs is harmless, when parallelism settings can silently introduce confounds into the comparison.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Different parallelism configurations change numerical reduction order and kernel selection, introducing variance unrelated to the attention-head variable.
An ablation study aims to attribute observed differences to one variable. When tensor and pipeline parallel sizes change between runs, the reduction order, kernel selection, and communication patterns change too, introducing numerical and convergence variance that is unrelated to attention head count. That makes the comparison confounded, even though optimizer sharding and micro-batching can be handled without altering the effective update or global batch size.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Attention head count is not a valid ablation variable because it is determined by hidden size and cannot be varied independently.
Why it's wrong here
In Megatron-style architectures, hidden size must be divisible by the number of attention heads, but the head count is still a configurable design choice that can be varied within that constraint. Treating it as non-independent is incorrect. The real issue in this scenario is that changing parallelism between runs introduces confounds, not that the variable itself is invalid.
- ✓
Different parallelism configurations change numerical reduction order and kernel selection, introducing variance unrelated to the attention-head variable.
Why this is correct
Tensor and pipeline parallelism alter how partial sums are reduced across GPUs and which kernels execute, producing small numerical differences and sometimes different convergence behavior. In an ablation isolating attention heads, those parallelism-induced effects become confounds. The colleague is right because the observed result could reflect parallelism rather than the variable under study, weakening the causal claim.
- ✗
Changing parallelism alters the optimizer state sharding, so the effective learning rate changes per parameter.
Why it's wrong here
Parallelism does redistribute optimizer state across ranks, but NeMo Framework is designed so the mathematical update applied to each parameter remains equivalent regardless of sharding. The learning rate is not silently rescaled per parameter by tensor or pipeline parallel size. While optimizer state placement changes, this is not the primary reason the colleague's concern about the ablation comparison is scientifically valid.
- ✗
Pipeline parallelism forces different micro-batch sizes, which always changes the global batch size and therefore the loss landscape.
Why it's wrong here
Pipeline parallelism introduces micro-batches to keep stages busy, but the global batch size can be preserved by adjusting gradient accumulation steps. The claim that it always changes global batch size is inaccurate. Even if batch size shifted, the more fundamental confound in this ablation is numerical and kernel variation from parallelism, not an unavoidable batch-size change.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.