NCA-GENL Experimentation Practice Question
An ML engineer is using NVIDIA NeMo to evaluate a retrieval-augmented generation pipeline. They want to measure whether adding a reranker improves answer faithfulness, but they must ensure the experiment is reproducible and comparable across runs. Which practice best supports a valid comparison between the pipeline with and without the reranker?
⚠ Common exam trap
The trap here is thinking that a larger dataset or more compute makes an experiment more valid, when the real requirement is holding all non-tested variables constant across conditions.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Use the same evaluation dataset, same base LLM checkpoint, and same decoding parameters for both pipeline variants, changing only the reranker component.
To attribute a change in faithfulness to the reranker, every other element of the RAG pipeline must be held constant. Using the same evaluation set, base checkpoint, and decoding parameters for both variants ensures the only difference is the presence or absence of the reranker. This controlled design yields a clean, reproducible comparison and avoids confounding variables that would invalidate the result.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Increase the temperature of the LLM in the reranker variant to encourage more diverse answers and better faithfulness.
Why it's wrong here
Changing temperature alongside the reranker introduces a second independent variable. Higher temperature generally reduces determinism and can lower faithfulness, so the experiment would no longer isolate the reranker's effect. Decoding parameters must remain fixed across conditions for the comparison to be interpretable.
- ✗
Evaluate the reranker variant on a larger and more diverse dataset than the baseline so the results are more statistically significant.
Why it's wrong here
Using different datasets for the two variants confounds the comparison. Any observed improvement could be due to easier or different questions rather than the reranker. Statistical significance does not rescue a confounded design; both variants must be evaluated on the same held-out set for the comparison to be valid.
- ✗
Run the baseline and reranker variants on different GPU models to test robustness across hardware configurations.
Why it's wrong here
Different hardware can change numerical precision, kernel behavior, and even tokenization throughput, adding another confound. The goal is to isolate the reranker, not to test hardware robustness. Hardware variation should be controlled or tested separately, not mixed into a component-level A/B comparison.
- ✓
Use the same evaluation dataset, same base LLM checkpoint, and same decoding parameters for both pipeline variants, changing only the reranker component.
Why this is correct
A valid A/B comparison requires that only the component under test—the reranker—differs between conditions. Keeping the evaluation dataset, base checkpoint, and decoding parameters identical ensures that any change in faithfulness scores is attributable to the reranker rather than to data, model, or sampling differences. This is the controlled-variable principle applied to RAG pipeline experimentation.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.