Courseiva
Experimentation →hardMultiple Select

NCA-GENL Experimentation Practice Question

When designing an experiment to evaluate the performance of an LLM on a downstream classification task, which THREE factors should be controlled to ensure the results are comparable across different model sizes?

⚠ Common exam trap

Candidates often overlook inference-time parameters like temperature and top-p, assuming that only the dataset and model weights need to remain identical during comparative evaluations.

Answer choices

Why each option matters

Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.

Correct answer & explanation

✓

Decoding strategy (temperature, top-p)

To conduct valid cross-model comparisons, one must neutralize external variables. Inference-time settings like decoding parameters, the specific prompt template, and the evaluation dataset must remain constant. If these vary, the results reflect differences in the evaluation environment rather than the intrinsic capabilities of the models being tested, rendering the experimentation data inconclusive for determining which model size is truly optimal for the specific classification application.

Answer analysis

Option-by-option breakdown

For each option: why learners choose it and why it is or isn't the right answer here.

  • ✓

    Decoding strategy (temperature, top-p)

    Why this is correct

    The decoding strategy directly influences the stochastic nature of the output. If one model uses greedy search and another uses high-temperature sampling, the performance differences are skewed by the sampling method. Consistency here is essential to isolate the model's architectural capacity from its generation behavior during the experiment.

  • ✓

    The prompt template used

    Why this is correct

    Prompts are the interface between the user and the model. Different templates can lead to drastically different performance levels due to varying sensitivities in the model's pre-training. Standardizing the prompt ensures that every model size is evaluated under identical conditions, providing a fair basis for comparison across your experiments.

  • ✗

    The learning rate used during fine-tuning

    Why it's wrong here

    Learning rates are typically tuned per model size because different architectures have unique optimal convergence rates. Enforcing the exact same learning rate across vastly different model sizes would be detrimental, as the smaller model might underfit while the larger one might diverge, leading to an unfair experimental comparison.

  • ✓

    The validation/test dataset

    Why this is correct

    Using a consistent evaluation dataset is the baseline for any comparative experiment. If models are tested on different data splits, you cannot determine if a performance difference is due to the model's capabilities or the difficulty level of the test samples. This is critical for scientific integrity in evaluation.

  • ✗

    The hardware GPU generation (e.g., A100 vs H100)

    Why it's wrong here

    Hardware generation influences throughput and latency, but it does not change the model's generative logic or classification accuracy. As long as the precision is consistent, the model's output should remain identical regardless of the underlying GPU architecture, making this factor unnecessary to control for accuracy metrics.

About these practice questions

This NCA-GENL question is part of Courseiva's 367-question bank — original exam-style content with full explanations and wrong-answer analysis, never real exam questions or exam dumps. Learn why practice questions differ from exam dumps →

How Courseiva writes practice questions · Editorial policy

JA

Written and reviewed by Johnson Ajibi, MSc IT Security

Senior Network & Security Engineer · founder of Courseiva

Last reviewed September 2026 · checked against the official NVIDIA exam blueprint

This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.