NCA-GENL Experimentation Practice Question
A research team is designing an experiment to measure how prompt phrasing affects the factuality of an LLM in a retrieval-augmented question-answering pipeline. Which two design choices are necessary to attribute observed factuality differences to the prompt rather than to other pipeline components? (Choose two.)
⚠ Common exam trap
The trap here is optimizing each variant's surrounding pipeline to make it look best, which destroys the control needed to attribute results to the prompt.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Hold the retrieved context, model checkpoint, and decoding parameters constant across prompt variants
Attributing factuality differences to prompt phrasing requires isolating the prompt as the only changed variable. Holding retrieval, model, and decoding settings constant removes competing explanations, while a shared evaluation set and rubric make scores comparable. Altering the retrieval index, temperature, or datasets per variant introduces confounds that make any observed improvement impossible to credit to the prompt itself.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Hold the retrieved context, model checkpoint, and decoding parameters constant across prompt variants
Why this is correct
If the retrieved passages, model weights, or sampling settings change between prompt variants, any factuality difference could come from those factors rather than the prompt. Fixing them makes the prompt the only manipulated variable, which is the definition of a controlled experiment. This isolation is what allows the team to draw a causal conclusion about phrasing.
- ✗
Evaluate each prompt variant on a different dataset tailored to its strengths
Why it's wrong here
Tailoring datasets to each variant guarantees that scores reflect dataset differences rather than prompt quality. A variant evaluated on questions aligned with its phrasing will look better regardless of true factuality. Controlled experiments require a common evaluation set so that every condition is measured under identical conditions.
- ✗
Increase the temperature for variants that produce shorter answers to equalize response length
Why it's wrong here
Adjusting temperature per variant deliberately changes the decoding distribution, which directly affects factuality and variability. It also confounds the comparison because temperature is a known driver of hallucination. Equalizing length by altering sampling is a manipulation, not a control, and it undermines the goal of attributing differences to prompt phrasing alone.
- ✓
Use a fixed, representative evaluation set of questions with reference answers scored by the same rubric
Why this is correct
A constant evaluation set and consistent scoring rubric ensure that every prompt variant faces identical questions and is judged by identical criteria. Without this, a variant could appear more factual because it happened to receive easier questions or a more lenient judge. Shared evaluation data and rubric turn raw outputs into comparable measurements.
- ✗
Allow the retrieval index to be rebuilt with different embedding models for each prompt variant
Why it's wrong here
Changing the embedding model changes which passages are retrieved, so differences in factuality could stem from better or worse evidence rather than from prompt phrasing. This introduces a confounding variable. To isolate the prompt's effect, the retrieval component must remain identical, including the index, embeddings, and retrieval parameters.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.