NCA-GENL Experimentation Practice Question
An engineer is evaluating different prompting strategies (Zero-shot, Few-shot, Chain-of-Thought) for an RAG pipeline. Which TWO metrics are most effective for quantifying the quality of the generative output during this experimentation?
⚠ Common exam trap
Candidates often choose general performance metrics like BLEU or ROUGE instead of RAG-specific metrics like Faithfulness and Answer Relevance, which specifically measure grounding against the retrieved context.
Answer choices
Why each option matters
Answer the question above first, then reveal the full breakdown to understand why each option is right or wrong.
Correct answer & explanation
✓
Faithfulness score
Quantitative evaluation is essential to move beyond subjective intuition in LLM experiments. Faithfulness (grounding in context) and Answer Relevance (usefulness to the query) provide distinct, measurable dimensions of performance. By measuring these, developers can iterate on prompts with empirical data, ensuring that changes to the prompt template actually improve system utility rather than just altering the verbosity or style of the generated response.
Answer analysis
Option-by-option breakdown
For each option: why learners choose it and why it is or isn't the right answer here.
- ✓
Faithfulness score
Why this is correct
Faithfulness measures whether the generated response is derived strictly from the retrieved context. This is critical in RAG experimentation to ensure the model does not hallucinate information outside the provided documents, maintaining accuracy and reliability for business applications where fact-based responses are required for user trust.
- ✗
Average token generation speed
Why it's wrong here
While important for production latency, token generation speed does not measure the quality or accuracy of the generative output. This is an operational performance metric rather than an experimentation metric for model behavior or prompt effectiveness, and focusing on it ignores the goal of improving response quality.
- ✓
Answer Relevance score
Why this is correct
Answer relevance evaluates how well the generated output directly addresses the user's prompt. During experimentation, identifying which prompting strategy yields the most focused and helpful responses allows researchers to prune ineffective approaches and optimize for user satisfaction without sacrificing the core utility of the RAG system.
- ✗
Number of model parameters
Why it's wrong here
The number of parameters is a static architectural property of the model, not a metric for evaluating prompting strategy effectiveness. It does not fluctuate during prompt engineering experiments and provides no feedback on how well the model is interpreting the retrieved context or user intent.
- ✗
Total GPU memory consumption
Why it's wrong here
GPU memory consumption is a hardware and infrastructure constraint. While vital for system stability, it is irrelevant for evaluating the quality or logic of the text generated by the model. Relying on this during prompt experimentation fails to inform the developer about the semantic correctness of the output.
About these practice questions
One of 367 original NCA-GENL practice questions on Courseiva, each with a full explanation and wrong-answer analysis — not exam dumps or protected exam content. Learn why practice questions differ from exam dumps →
JA
Written and reviewed by Johnson Ajibi, MSc IT Security
Senior Network & Security Engineer · founder of Courseiva
Last reviewed September 2026 · checked against the official NVIDIA exam blueprint
This NCA-GENL practice question is part of Courseiva's free NVIDIA certification practice question bank. Courseiva provides original exam-style practice questions with explanations, topic-based practice, mock exams, readiness tracking, and study analytics to help learners prepare for the NCA-GENL exam.