Courseiva
Experimentation →mediumMultiple Select

NCA-GENL Experimentation Practice Question

An engineer is designing an experiment to measure how prompt phrasing affects the output quality of a deployed LLM. They will test several prompt templates against a fixed evaluation set and want the comparison to be valid. Which two practices are required? (Choose two.)

⚠ Common exam trap

The trap here is assuming that varying other factors like temperature or dataset adds useful coverage, when in a controlled prompt comparison every factor except the prompt must stay constant.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Hold the model, decoding parameters, and evaluation dataset constant across all prompt templates.

A valid prompt experiment isolates the prompt template as the only variable. Keeping the model, decoding parameters, and evaluation dataset fixed means any quality difference can be attributed to phrasing, and scoring all outputs with the same rubric ensures the measurement itself does not shift. Changing datasets, retraining, or varying temperature introduces confounds that make the comparison meaningless.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✗

    Retrain the model after each prompt template so it adapts to the new phrasing.

    Why it's wrong here

    Retraining between templates changes the model itself, so the experiment would compare different models rather than different prompts. Prompt engineering experiments are meant to evaluate inference-time behavior of a fixed model. Retraining also adds cost and time and destroys the controlled comparison the engineer is trying to achieve.

  • ✗

    Use a different evaluation dataset for each prompt template so the results are not correlated.

    Why it's wrong here

    Varying the evaluation data across templates confounds the prompt effect with dataset difficulty differences. Scores would no longer be comparable because each template faces different questions. A fixed, shared evaluation set is what makes the quality numbers directly comparable across prompt variants.

  • ✗

    Increase the temperature for each successive template to explore more diverse outputs.

    Why it's wrong here

    Changing temperature between templates alters the sampling distribution, so quality differences could stem from randomness rather than phrasing. Decoding parameters must stay fixed for a prompt comparison. Higher temperature would also add variance that obscures the effect the experiment is designed to measure.

  • ✓

    Hold the model, decoding parameters, and evaluation dataset constant across all prompt templates.

    Why this is correct

    Controlling model, decoding settings, and evaluation data ensures the only difference between runs is the prompt template, which is the variable under study. If any of these drift, an observed quality change could be caused by the model or sampling rather than the phrasing, invalidating the comparison the experiment is meant to support.

  • ✓

    Score every template's outputs with the same rubric and scoring procedure.

    Why this is correct

    Applying one consistent rubric and scoring process to all outputs keeps the measurement instrument constant, so differences reflect the prompts rather than changes in how quality was judged. If scoring criteria shift between templates, the comparison becomes subjective and any apparent improvement may come from more lenient grading.

About these practice questions

Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.