NCA-GENL Experimentation Practice Question
A team is running an A/B experiment comparing two prompt templates for a customer-facing LLM assistant. After one week, template A shows a 2% higher task-completion rate with a p-value of 0.04. The team lead wants to declare A the winner immediately. Which consideration is most important before making that decision?
⚠ Common exam trap
The trap here is treating a p-value just below 0.05 as a definitive result without questioning whether the data-collection process allowed the team to stop at a favorable moment.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Whether the experiment was stopped early or the sample size was fixed in advance, because repeatedly checking and stopping inflates false-positive rates.
A p-value of 0.04 is fragile if the team monitored results continuously and stopped at the first significant reading, because optional stopping inflates the false-positive rate. The critical check is whether the sample size and stopping rule were fixed in advance or whether a sequential correction applies. Test direction, author identity, and sampling temperature are secondary concerns.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Whether the two prompt templates were written by different engineers, since authorship affects output quality.
Why it's wrong here
Authorship is not a statistical confound in a randomized A/B test because assignment to templates is random with respect to who wrote them. The templates are fixed treatments, not varying conditions. This concern distracts from the actual threat to validity, which is the analysis procedure.
- ✓
Whether the experiment was stopped early or the sample size was fixed in advance, because repeatedly checking and stopping inflates false-positive rates.
Why this is correct
Peeking at results and stopping as soon as significance appears inflates the Type I error rate well above the nominal threshold. A p-value of 0.04 from an optional-stopping procedure may not reflect a true 4% false-positive probability. Before declaring a winner, the team must confirm the analysis plan was pre-registered or apply a sequential-testing correction.
- ✗
Whether the assistant uses a temperature greater than zero, since sampling randomness invalidates A/B tests.
Why it's wrong here
Stochastic generation adds variance but does not invalidate randomized experiments; it is absorbed into the error term and handled by adequate sample size. Many production LLM A/B tests run with nonzero temperature successfully. This option misidentifies a manageable variance source as a fatal flaw.
- ✗
Whether the p-value was computed with a one-tailed or two-tailed test, since one-tailed tests are always invalid.
Why it's wrong here
One-tailed tests are valid when the direction of the effect is hypothesized in advance; they are not inherently invalid. The choice affects interpretation but is not the most important concern here. This option raises a methodological detail while missing the larger risk of acting on an early peek at accumulating data.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.