NCA-GENL Experimentation Practice Question
A team is designing a controlled experiment to measure whether increasing LoRA rank improves instruction-following accuracy on a held-out benchmark. They want the comparison to be scientifically valid. Which experimental design choice best supports a valid conclusion?
⚠ Common exam trap
The trap here is believing that changing multiple factors at once is more efficient, when it actually destroys the ability to attribute the result to LoRA rank.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Hold the base model, dataset, prompt template, and evaluation harness fixed while varying only LoRA rank.
A scientifically valid comparison changes one independent variable while holding everything else constant. Here, LoRA rank is the factor under test, so the base model, training data, prompt template, and evaluation harness must remain identical across runs. Using a shared held-out benchmark and reporting aggregate accuracy across seeds ensures observed differences reflect rank rather than confounding factors or evaluation noise.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Change LoRA rank and the base model simultaneously so the experiment covers more configurations in one run.
Why it's wrong here
Varying two independent factors at once confounds the result: any accuracy change could be due to rank, the base model, or their interaction. A valid comparison requires isolating the variable of interest. This design would produce an observation but not a causal conclusion about LoRA rank, which is what the team explicitly wants to measure.
- ✗
Report only the best accuracy observed across several random seeds for each rank.
Why it's wrong here
Selecting the maximum across seeds inflates the reported metric and hides run-to-run variance, making the comparison statistically misleading. It also rewards lucky runs rather than stable improvements. A defensible experiment reports a central tendency such as mean accuracy with variance across seeds, evaluated on a shared held-out set.
- ✗
Evaluate each rank on a freshly sampled held-out set to reduce any dataset bias.
Why it's wrong here
Using a different held-out set for each rank introduces evaluation variance that is unrelated to the rank change. Differences in accuracy could reflect which examples landed in each sample rather than the model change. A valid comparison requires a common, fixed evaluation set so scores are directly comparable across configurations.
- ✓
Hold the base model, dataset, prompt template, and evaluation harness fixed while varying only LoRA rank.
Why this is correct
Controlling all other factors isolates LoRA rank as the single independent variable, so any accuracy difference can be attributed to it. Fixed data, prompts, and evaluation harness also ensure the held-out benchmark measures the same capability across runs. This is the core requirement for a valid controlled comparison in LLM experimentation.
About these practice questions
This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.