NCA-GENL Experimentation Practice Question
An AI engineer is conducting an experiment to compare two different fine-tuning approaches for a large language model using NVIDIA NeMo: full fine-tuning versus parameter-efficient fine-tuning (PEFT) with LoRA. The engineer wants to determine which approach yields better performance on a downstream question-answering task while minimizing computational cost. Which metric should the engineer prioritize to evaluate the trade-off between performance and cost?
⚠ Common exam trap
The trap here is focusing solely on parameter count or training speed, which are incomplete because they ignore either the performance or the cost side of the trade-off.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Validation accuracy per GPU-hour
Validation accuracy per GPU-hour effectively combines the two objectives: it measures the model's performance on the question-answering task (validation accuracy) and normalizes it by the computational cost (GPU-hours). This allows the engineer to compare full fine-tuning and LoRA on a level playing field, identifying which method provides the best accuracy for the resources invested. It directly supports the goal of minimizing cost while maximizing performance.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✗
Training loss convergence rate
Why it's wrong here
Training loss convergence rate indicates how quickly the model optimizes the training objective, but it does not directly measure downstream task performance or computational cost. A faster convergence does not necessarily mean better question-answering accuracy, and it ignores the cost dimension. The engineer needs a metric that captures both performance and cost, which convergence rate alone cannot provide.
- ✓
Validation accuracy per GPU-hour
Why this is correct
Validation accuracy per GPU-hour quantifies the trade-off between model performance and computational expense. It measures how much accuracy is gained for each unit of GPU time, directly addressing the goal of minimizing cost while maximizing performance. This metric allows the engineer to compare full fine-tuning and LoRA by showing which approach delivers better accuracy more efficiently.
- ✗
Number of trainable parameters
Why it's wrong here
The number of trainable parameters is a measure of model complexity and can indicate potential memory savings with PEFT, but it does not reflect actual performance on the question-answering task. A method with fewer parameters might underperform, and this metric alone does not account for computational cost during training or inference. It is an incomplete proxy for the performance-cost trade-off.
- ✗
Inference latency on a CPU
Why it's wrong here
Inference latency on a CPU is relevant for deployment but does not capture training cost or the accuracy of the fine-tuned model on the downstream task. The experiment focuses on fine-tuning approaches, so training efficiency and validation performance are more critical. CPU inference latency may be misleading if the model is deployed on GPUs, and it ignores the performance dimension entirely.
About these practice questions
Courseiva writes every NCA-GENL question from scratch — 367 in total, each with an explanation and a wrong-answer breakdown. None are copied from real exams or dumps. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.